Adaptive Spectral Code Book Audio Encoder
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding technologies, such as CELP, face challenges in maintaining high-quality audio synthesis at low bit rates, particularly in representing noise-like sounds, leading to artifacts like buzzing distortion and poor representation of complex noise-like signals.
Innovation Solution
The proposed method involves encoding and decoding audio signals using an adaptive spectral code book and a fixed spectral code book in the frequency domain, with the encoder transforming time domain signal segments into frequency domain representations to control spectral distribution, allowing for efficient encoding of noise-like sounds even at low bit rates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If CELP encoding is used to reduce bandwidth consumption, then transmission bandwidth and power consumption are reduced, but the quality of noise-like sounds deteriorates with buzzing distortion artifacts
Solution Approach 1:
The excitation signal is segmented into two distinct components: a periodic component for voiced segments and a noise component for unvoiced segments. This segmentation allows each component to be optimized independently, with the noise component specifically designed to represent noise-like sounds without the buzzing distortion artifacts that plague traditional CELP encoders at low bitrates.
Solution Approach 2:
The invention introduces a noise excitation dimension that is separate from the traditional periodic excitation dimension. By adding this new dimension for noise representation, the encoder can capture noise-like sounds (fricatives, background noise) effectively without being constrained by the pulse-based excitation model that causes buzzing artifacts.
2Manufacturing precision
If pulse-based excitation is used in CELP, then voiced segments are well represented, but noise-like segments fail to capture spectral energy distribution at low bitrates
Solution Approach 1:
The encoder dynamically switches between periodic excitation for voiced segments and noise excitation for unvoiced segments. This dynamic adaptation allows the system to maintain high representation accuracy for voiced segments while simultaneously achieving good representation of noise-like segments, overcoming the limitation of fixed pulse-based excitation.
Solution Approach 2:
The invention changes the excitation signal parameters based on the segment type: using periodic pulse parameters for voiced segments and noise parameters for unvoiced segments. This parameter change allows the encoder to adapt its representation strategy to match the characteristics of the input signal, achieving high accuracy across different sound types.
3Quantity of substance
If low bit rate encoding is applied, then bandwidth consumption is reduced, but the synthesized speech quality deteriorates with sparseness artifacts
Solution Approach 1:
The encoder performs preliminary classification of signal segments to identify voiced versus unvoiced portions before encoding. This preliminary action allows the system to apply the appropriate excitation model (periodic or noise) in advance, preventing the sparseness artifacts that would otherwise occur when pulse-based excitation is inappropriately applied to noise-like segments at low bitrates.
Data Source
Figure 1
Figure 2~3
Figure 4
AI summary
The present invention relates to a frequency domain based method of encoding and decoding an audio signal, wherein an adaptive spectral code book is updated with synthesized frequency domain representations of a time domain signal segment. A frequency analysis is performed of a received time domain signal segment in order to obtain a frequency domain representation, and the adaptive spectral code book is searched for a first approximation of the frequency domain representation. A fixed spectral code book is searched for an approximation of the residual frequency representation. A synthesized frequency domain representation may be generated from the two approximations.