Two-Stage Vector Quantization for Low-Bitrate Audio Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech and audio coding technologies face challenges in reducing memory and computational complexity while maintaining coding efficiency, particularly for audio signals with voiced, unvoiced, generic, transition, or comfort noise parts, where lattice-based codebooks are not efficient at lower bitrates.
Innovation Solution
The approach involves a two-stage quantization process where the input vector is first quantized using a codebook, and then the vector components are split into groups based on energy values, with swapping performed to meet specific energy criteria, followed by a second quantization stage using lattice structures selected based on the initial codevector index.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If lattice-based codebooks are used for quantization, then memory requirements and computational complexity are reduced, but quantization efficiency deteriorates for lower bitrates and certain coding modes
Solution Approach 1:
The codebook is segmented into multiple subsets, each optimized for specific coding modes (voiced, unvoiced, generic, transition, CNG). This allows the system to select the most appropriate subset for each signal type, maintaining high quantization efficiency while keeping each individual subset compact and computationally efficient.
Solution Approach 2:
The codebook structure is made dynamic through mode-dependent selection. Different coding modes (voiced, unvoiced, generic, transition, CNG) can select from different codebook subsets, allowing the system to adapt its quantization strategy to the specific characteristics of the signal being encoded, thereby optimizing efficiency for each mode while maintaining overall complexity control.
2Device complexity
If lattice-based codebooks are used for quantization, then memory requirements are reduced, but quantization quality deteriorates for certain coding modes
Solution Approach 1:
The codebook is divided into multiple specialized subsets, each tailored for specific coding modes. This segmentation allows each subset to be optimized for its specific purpose, ensuring high quantization quality for voiced, unvoiced, generic, transition, and CNG parts independently, while the overall memory requirement remains controlled due to the structured organization.
Solution Approach 2:
Different regions or subsets of the codebook are designed with different qualities and characteristics appropriate for specific coding modes. For example, certain subsets may have higher resolution for voiced segments while others are optimized for unvoiced or comfort noise generation, ensuring locally optimized quality for each signal type without compromising overall memory efficiency.
3Device complexity
If a single codebook is used for all coding modes, then device complexity is reduced, but coding efficiency deteriorates for different signal types
Solution Approach 1:
The unified codebook is segmented into multiple mode-specific subsets, allowing the system to maintain a relatively simple overall structure while providing specialized optimization for each coding mode (voiced, unvoiced, generic, transition, CNG). This segmentation enables mode-dependent selection that improves coding efficiency for each signal type without requiring completely separate codebooks for each mode.
Solution Approach 2:
The codebook structure is designed to be universal and multi-functional, serving all coding modes through a single organized structure. Different subsets within the codebook can be selected based on the coding mode, allowing one codebook structure to fulfill multiple functions and optimize performance across all signal types without requiring separate dedicated codebooks for each mode.
Data Source
Figure 1a~1b
Figure 2~3
Figure 4a~4d
AI summary
It is inter alia disclosed to determine a first quantized representation of an input vector, and to determine a second quantized representation of the input vector based on a codebook depending on the first quantized representation