Adaptive Speech Coding Bit Allocation for Dominant Frequency Bands
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing bit allocation schemes in speech/audio coding do not consider input signal characteristics, leading to inefficient bit allocation and limited improvement in sound quality.
Innovation Solution
A speech/audio coding apparatus that identifies dominant frequency bands and adaptively determines group widths based on input signal characteristics, allocating bits using both energy and norm variance to prioritize perceptually important groups and subbands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional bit allocation schemes are used without considering input signal characteristics, then the encoding process is simple, but bit allocation efficiency is poor and sound quality improvement is limited
Solution Approach 1:
The patent applies dynamics by making the bit allocation scheme adaptive to input signal characteristics. The system dynamically adjusts bit allocation based on detected signal features (transient/stationary frames, spectral norm distribution), transforming a static allocation method into a dynamic one that responds to signal content, thereby improving both encoding efficiency and sound quality
Solution Approach 2:
The patent changes parameters by introducing signal characteristic detection (transient/stationary identification, spectral norm estimation) that modifies the bit allocation parameters based on input signal properties. This parameter adaptation allows the system to optimize bit distribution according to actual signal needs, resolving the contradiction between simple processing and high quality output
2Manufacturing precision
If more quantization bits are allocated to perceptually important bands, then sound quality improves, but the complexity of bit allocation increases
Solution Approach 1:
The patent applies local quality by differentiating bit allocation across different frequency bands based on their perceptual importance and signal characteristics. Instead of uniform allocation, the system identifies specific bands requiring more bits (those with higher spectral norms or perceptual significance) and allocates resources locally to those regions, improving sound quality where it matters most without unnecessarily increasing overall complexity
Solution Approach 2:
The patent uses preliminary action by detecting signal characteristics (transient/stationary frames, spectral envelope) before performing bit allocation. This pre-analysis allows the system to prepare appropriate allocation strategies in advance, reducing the complexity of the actual allocation process while ensuring optimal quality outcomes
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided are a voice audio encoding device, voice audio decoding device, voice audio encoding method, and voice audio decoding method that efficiently perform bit distribution and improve sound quality. Dominant frequency band identification unit (301) identifies a dominant frequency band having a norm factor value that is the maximum value within the spectrum of an input voice audio signal. Dominant group determination units (302-1 to 302-N) and non-dominant group determination unit (303) group all sub-bands into a dominant group that contains the dominant frequency band and a non-dominant group that contains no dominant frequency band. Group bit distribution unit (308) distributes bits to each group on the basis of the energy and norm variance of each group. Sub-band bit distribution unit (309) redistributes the bits that have been distributed to each group to each sub-band in accordance with the ratio of the norm to the energy of the groups.