Audio Decoder Gain Scaling for Low-Bitrate CELP Mismatch
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Code Excited Linear Prediction (CELP) speech coders face challenges in maintaining high-quality speech and audio reproduction at low data rates, especially for music and non-speech signals, due to model mismatch, leading to degraded audio quality.
Innovation Solution
A method and apparatus for generating an enhancement layer within an audio coding system that scales the core layer output audio using a set of gain values to minimize error between the input and coded signals, transmitting the optimal gain and error values as part of the enhancement layer, and using these to improve audio quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If CELP speech coding is used to compress audio signals, then data rate is reduced, but audio quality deteriorates for music and non-speech signals due to model mismatch
Solution Approach 1:
The audio signal is divided into two separate layers: a core layer coded with CELP for basic speech representation, and an enhancement layer that adds additional detail. This segmentation allows each layer to be optimized for its specific function, maintaining low overall data rate while improving audio quality for music and non-speech signals.
Solution Approach 2:
The invention changes the coding parameters by introducing an enhancement layer with different coding characteristics than the core layer. The enhancement layer uses parameters specifically designed to complement the core layer, addressing the model mismatch problem for music and non-speech signals while maintaining compression efficiency.
2Productivity
If CELP model is applied to music and non-speech signals, then compression is achieved, but model mismatch causes severely degraded audio quality
Solution Approach 1:
By segmenting the coding into core and enhancement layers, the system achieves compression through the core layer while the enhancement layer specifically addresses audio quality for music and non-speech signals, resolving the model mismatch issue without sacrificing compression efficiency.
Solution Approach 2:
The audio coding system uses a composite structure combining two different coding approaches: CELP for the core layer and a complementary enhancement layer. This composite structure leverages the strengths of each approach to achieve both compression efficiency and high audio quality for diverse signal types.
3Manufacturing precision
If core layer only is used, then data rate is low, but audio quality is insufficient; if enhancement layer is added, then audio quality improves, but data rate increases
Solution Approach 1:
The enhancement layer applies partial action by selectively adding detail only where needed to improve audio quality, rather than fully re-coding the entire signal. This approach improves audio quality for music and non-speech signals while minimizing the additional data rate requirement.
Data Source
AI summary
A set of peaks in a reconstructed audio vector Ŝ of a received audio signal is detected and a scaling mask ψ(Ŝ) based on the detected set of peaks is generated. A gain vector g* is generated based on at least the scaling mask and an index j representative of the gain vector. The reconstructed audio signal is scaled with the gain vector to produce a scaled reconstructed audio signal. A distortion is generated based on the audio signal and the scaled reconstructed audio signal. The index of the gain vector based on the generated distortion is output.


