Speech Encoder Shape Vector Quantization Order
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Transform encoding techniques for speech signals, particularly when quantizing gain and shape vectors in order, result in increased quantization distortion, especially at lower bit rates, leading to quality deterioration for signals with strong tonality like vowels.
Innovation Solution
An encoding apparatus and method that divides transform coefficients into subbands, encodes shape vectors, calculates target gains, and forms a gain vector to encode spectral shapes accurately, minimizing distortion by prioritizing shape vector encoding before gain vector encoding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If gain and shape vectors are quantized in order (gain first, then shape), then encoding efficiency is improved, but quantization distortion increases especially for signals with strong tonality
Solution Approach 1:
The patent applies preliminary action by performing shape vector quantization before gain vector quantization, reversing the conventional order. This ensures that the spectral shape is accurately determined first, and then the gain is applied. The shape vector encoding section quantizes the shape vector to generate shape vector code, and subsequently the gain vector encoding section quantizes the gain vector based on the already-quantized shape, thereby preventing gain quantization distortion from affecting shape accuracy.
Solution Approach 2:
The patent inverts the conventional quantization order by quantizing the shape vector before the gain vector. Instead of the traditional gain-first approach, this invention processes shape information first and then applies gain information, which fundamentally changes the quantization sequence to preserve spectral shape accuracy while maintaining encoding efficiency.
2Productivity
If lower bit rates are used to efficiently utilize radio wave resources, then transmission efficiency is improved, but speech quality deteriorates due to increased quantization distortion
Solution Approach 1:
By performing shape vector quantization before gain vector quantization, the patent ensures that the more critical spectral shape information is preserved with higher accuracy even at lower bit rates. This preliminary action on shape encoding protects against quality deterioration while maintaining the benefits of low bit rate transmission.
Solution Approach 2:
The patent changes the encoding parameter sequence by prioritizing shape vector parameters over gain vector parameters in the quantization process. This parameter reordering ensures that bits are allocated more effectively to preserve spectral characteristics, thereby maintaining speech quality at lower bit rates.
3Quantity of substance
If gain vector is quantized first to reduce data量, then transmission data volume is reduced, but spectral shape representation accuracy decreases
Solution Approach 1:
The patent performs shape vector encoding as a preliminary step before gain vector encoding. This ensures that the spectral shape is accurately captured and represented first, and then the gain information is added. This sequence maintains high spectral shape representation accuracy while still achieving data volume reduction through efficient quantization of both vectors.
Solution Approach 2:
The patent inverts the conventional encoding sequence by encoding shape vectors before gain vectors. This reversal ensures that spectral shape information, which is critical for quality, is preserved with higher fidelity, while still achieving compression through systematic quantization of both parameters.
Data Source
AI summary
An encoding apparatus includes a first layer encoder that encodes an input signal, a first layer decoder that decodes the first layer encoded data, a weighting filter that filters a first layer error signal to acquire a weighted first layer error signal, a first layer error transform coefficient calculator that transforms the weighted first layer error signal into a frequency domain, and a second layer encoder that encodes the first layer error transform coefficient. The second layer encoder includes a first shape vector encoder that refers the first layer error transform coefficient to generate a first shape vector and first shape encoded information. A target gain calculator calculates a target gain using the first layer error transform coefficient and the first shape vector, a gain vector generator generates a gain vector, and a gain vector encoder encodes the gain vector to acquire gain encoded information.


