Hierarchical Audio Encoding Using CELP Parameter Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional scalable speech encoding methods face inefficiencies in encoding residual signals, leading to loss of speech components and noise inclusion, especially when using CELP schemes in enhancement layers, which compromises the quality of decoded signals.
Innovation Solution
A speech encoding apparatus that employs a CELP scheme for both initial and secondary encoding stages, using parameters like quantized LSP, adaptive excitation lag, and excitation gains to encode the speech signal hierarchically, avoiding the use of residual signals directly in the enhancement layer.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional scalable encoding uses residual signals as input for enhancement layer encoding, then hierarchical encoding structure is achieved, but speech components are lost and noise components are included leading to poor encoding efficiency
Solution Approach 1:
The patent segments the speech signal encoding into multiple independent layers, where each layer encodes the original speech signal separately using CELP. This allows the enhancement layer to receive clean speech signal input rather than degraded residual signals, maintaining speech component integrity while achieving hierarchical structure through layered parameter encoding.
Solution Approach 2:
The patent performs preliminary encoding of the speech signal in the base layer to extract parameters such as LSP coefficients and excitation signals. These parameters are then used in the enhancement layer to guide the encoding process, allowing the enhancement layer to efficiently encode differences or refinements without losing speech components.
2Adaptability or versatility
If CELP scheme is applied to residual signals in enhancement layer, then hierarchical encoding is achieved, but encoding efficiency deteriorates due to loss of speech components in residual signals
Solution Approach 1:
The patent introduces parameters from the base layer encoding as intermediaries that bridge the base layer and enhancement layer. These parameters (LSP coefficients, excitation signals) serve as mediators that carry essential speech information to the enhancement layer, enabling it to perform CELP encoding on reconstructed speech signals rather than on degraded residual signals.
Solution Approach 2:
The patent creates a copy of the speech signal representation through parameter encoding in the base layer. The enhancement layer uses these copied parameters to reconstruct or refine the speech signal, effectively working with a clean copy rather than with residual signals that have lost speech components.
3Quantity of substance
If scalable encoding uses residual signals for enhancement layer, then bit rate reduction is achieved, but decoded signal quality deteriorates due to noise inclusion
Solution Approach 1:
The patent segments the encoding task into multiple layers where each layer independently processes speech parameters rather than residual signals. This segmentation allows progressive refinement of speech quality at different bit rates while maintaining signal integrity, as each layer works with clean speech representations rather than accumulating noise from residual signal processing.
Solution Approach 2:
The patent changes the fundamental parameter being encoded from residual signals to speech signal parameters (LSP coefficients, excitation signals). This parameter transformation enables the enhancement layer to improve decoded signal quality by encoding refined versions of speech parameters rather than noise-containing residuals, achieving better quality at higher bit rates.
Data Source
AI summary
There is disclosed an audio encoding device capable of realizing effective encoding while using audio encoding of the CELP method in an extended layer when hierarchically encoding an audio signal. In this device, a first encoding section (115) subjects an input signal (S11) to audio encoding processing of the CELP method and outputs the obtained first encoded information (S12) to a parameter decoding section (120). The parameter decoding section (120) acquires a first quantization LSP code (L1), a first adaptive excitation lag code (A1), and the like from the first encoded information (S12), obtains a first parameter group (S13) from these codes, and outputs it to a second encoding section (130). The second encoding section (130) subjects the input signal (S11) to a second encoding processing by using the first parameter group (S13) and obtains second encoded information (S14). A multiplexing section (154) multiplexes the first encoded information (S12) with the second encoded information (S14) and outputs them via a transmission path N to a decoding apparatus (150).


