Audio Encoding With Layered Noise Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding methods face challenges in managing noise suppression effectively across different bit rates, leading to quality loss at low bit rates or distortion at high bit rates, especially when encoding speech and music signals.
Innovation Solution
Applying different amounts of noise suppression to generate multiple target signals, where the primary coded data is produced using the highest noise suppression for low bit rate speech and enhancement data is generated using progressively lower noise suppression levels for higher bit rates, allowing optimal noise management for each coding layer.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If high noise suppression is applied to improve low bit rate speech quality, then speech encoding quality is improved, but distortion is introduced to music and audio signals
Solution Approach 1:
The audio signal is segmented into different categories (speech, music, audio) and different noise suppression levels are applied to each segment. The system divides the encoding process into multiple layers with different noise suppression characteristics, allowing speech to receive high noise suppression while music and audio signals receive lower or no noise suppression.
Solution Approach 2:
The noise suppression level is made dynamic rather than fixed. The system adapts the noise suppression amount based on the signal type and bit rate conditions. Enhancement layers dynamically adjust the noise suppression characteristics to match the specific encoding requirements of each layer.
2Object-generated harmful factors
If no noise suppression is applied to preserve music and audio signal quality, then signal distortion is avoided, but low bit rate speech encoding quality deteriorates
Solution Approach 1:
The encoding process is segmented into a core encoder and multiple enhancement layers. The core encoder applies high noise suppression optimized for speech, while enhancement layers progressively reduce noise suppression to preserve music and audio quality at higher bit rates.
Solution Approach 2:
Noise suppression is applied preliminarily in the core encoding stage to ensure speech quality at low bit rates. Subsequent enhancement layers then progressively refine the signal with reduced noise suppression, allowing the system to start with aggressive noise suppression and gradually recover signal fidelity.
3Device complexity
If a single noise suppression level is used for all bit rates, then device complexity is reduced, but encoding quality is lost at specific bit rates
Solution Approach 1:
The coding structure is segmented into a core encoder and discrete enhancement layers. Each layer can be independently configured with appropriate noise suppression levels, allowing the system to maintain simple operation at low bit rates while providing quality options at higher bit rates through selective layer activation.
4Adaptability or versatility
If embedded variable rate coding with multiple layers is implemented to achieve flexible bit rates, then adaptability is improved, but device complexity increases due to multiple encoders
Solution Approach 1:
The enhancement layers are nested within the core encoder structure, forming an embedded variable rate coding system. Each enhancement layer builds upon the previous layer, allowing flexible bit rate adaptation by selectively including or excluding layers while maintaining a compact nested architecture rather than independent parallel encoders.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
For taking account of different requirements on encoded audio data having different bit rates, at least two different amounts of pre-processing are applied to an audio signal to obtain at least two different target signals. A first one of the target signals is then encoded to obtain primary coded data. At least a second one of the at least two different target signals is moreover used for generating enhancement data for the primary coded data.