Audio Codec Companding with Spectral Extension for Lower Coding Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio codecs introduce noticeable distortion in the form of coding noise due to lossy compression techniques, particularly evident during low-intensity segments of audio signals, which can manifest as pre-echo artifacts, and current solutions like filters introduce phase distortion or reduce frequency resolution.
Innovation Solution
A companding technique is employed to process audio signals by dividing them into short time segments, calculating and applying wideband gains in the frequency domain to amplify low-intensity segments and attenuate high-intensity segments during compression, and inversely doing so during expansion to restore the original dynamic range, effectively reducing quantization noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If lossy data compression techniques are used to reduce storage or data rate requirements, then the data rate is reduced, but the fidelity of source content deteriorates and coding noise is introduced
Solution Approach 1:
The audio signal is divided into multiple frames, and each frame is further divided into frequency subbands using a filter bank. This segmentation allows different quantization strategies to be applied to different frequency regions, enabling better quality at lower bitrates by preserving important frequency components while compressing less critical ones.
Solution Approach 2:
Different quantization precision is applied to different frequency subbands based on their perceptual importance. High-frequency subbands use coarser quantization while low-frequency subbands use finer quantization, optimizing the trade-off between bitrate and perceived audio quality.
2Object-affected harmful factors
If coding noise is shaped in the frequency domain over long frames, then the noise becomes less audible through masking effects, but pre-echo distortion occurs during low intensity parts of the frame
Solution Approach 1:
Each frame is divided into multiple frequency subbands using a filter bank. This allows independent noise shaping control for each subband, enabling the system to reduce pre-echo in transient regions while maintaining noise masking benefits in stationary regions.
Solution Approach 2:
The noise shaping filter characteristics are made adaptive and time-varying within each frame. The system dynamically adjusts the noise shaping parameters based on the local signal characteristics, allowing it to respond to transients and prevent pre-echo while maintaining effective noise masking during steady-state portions.
3Object-generated harmful factors
If filters are used to avoid pre-echo artifacts, then pre-echo is reduced, but phase distortion and temporal smearing are introduced
Solution Approach 1:
Different noise shaping strategies are applied to different frequency subbands. By operating in the frequency domain with subband-specific parameters, the system can reduce pre-echo in transient regions without applying broad spectral filtering that would cause phase distortion and temporal smearing across the entire signal.
4Object-generated harmful factors
If smaller transform windows are used to reduce pre-echo, then transient response is improved, but frequency resolution is significantly reduced
Solution Approach 1:
The signal is segmented into frequency subbands using a filter bank with longer transform windows. This segmentation allows the system to achieve good frequency resolution in each subband while using adaptive noise shaping to handle transients, avoiding the need to reduce the overall window size and lose frequency resolution.
Solution Approach 2:
The problem is solved by moving from a single-dimensional time-domain approach to a two-dimensional frequency-subband domain. By transforming the problem into the frequency domain and applying subband-specific processing, the system can maintain long window lengths for frequency resolution while addressing transient issues through frequency-selective noise shaping.
Data Source
AI summary
Embodiments are directed to a companding method and system for reducing coding noise in an audio codec. A compression process reduces an original dynamic range of an initial audio signal through a compression process that divides the initial audio signal into a plurality of segments using a defined window shape, calculates a wideband gain in the frequency domain using a non-energy based average of frequency domain samples of the initial audio signal, and applies individual gain values to amplify segments of relatively low intensity and attenuate segments of relatively high intensity. The compressed audio signal is then expanded back to the substantially the original dynamic range that applies inverse gain values to amplify segments of relatively high intensity and attenuating segments of relatively low intensity. A QMF filterbank is used to analyze the initial audio signal to obtain a frequency domain representation.


