Voice Audio Encoding with Group-Based Subband Bit Allocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing bit allocation schemes in speech/audio coding do not consider input signal characteristics, leading to inefficient bit allocation and limited improvement in sound quality.
Innovation Solution
A speech/audio coding apparatus and method that transforms an input signal into a frequency domain, estimates energy envelopes for subbands, groups them, and allocates bits on a group-by-group basis, with a second allocation step distributing bits to subbands based on energy and norm variance to enhance perceptual importance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If existing bit allocation schemes are used without considering input signal characteristics, then the coding process is simple, but bit allocation efficiency is poor and sound quality improvement is limited
Solution Approach 1:
The frequency spectrum is divided into multiple subbands, and each subband is further divided into groups based on energy envelope characteristics. This segmentation allows different bit allocation strategies to be applied to different groups, improving bit allocation efficiency while maintaining manageable complexity through systematic organization.
Solution Approach 2:
The energy envelope is estimated and quantized before bit allocation occurs. This preliminary analysis of signal characteristics enables the subsequent bit allocation to be optimized according to actual signal needs, rather than using uniform allocation, thereby improving efficiency without excessive complexity.
2Manufacturing precision
If uniform bit allocation is used across all subbands, then the coding process is simple, but perceptual importance of different subbands is not considered leading to poor sound quality
Solution Approach 1:
Different bit allocation strategies are applied to different subband groups based on their local energy characteristics and perceptual importance. Groups with higher energy or greater perceptual significance receive more bits, while less important groups receive fewer bits, optimizing sound quality without requiring complex global optimization.
Solution Approach 2:
The bit allocation is dynamically adjusted based on energy envelope parameters and norm variance of each group. By changing the allocation parameters according to signal characteristics rather than using fixed uniform allocation, the system achieves better sound quality with controlled complexity.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided are a voice audio encoding device, voice audio decoding device, voice audio encoding method, and voice audio decoding method that efficiently perform bit distribution and improve sound quality. Dominant frequency band identification unit (301) identifies a dominant frequency band having a norm factor value that is the maximum value within the spectrum of an input voice audio signal. Dominant group determination units (302-1 to 302-N) and non-dominant group determination unit (303) group all sub-bands into a dominant group that contains the dominant frequency band and a non-dominant group that contains no dominant frequency band. Group bit distribution unit (308) distributes bits to each group on the basis of the energy and norm variance of each group. Sub-band bit distribution unit (309) redistributes the bits that have been distributed to each group to each sub-band in accordance with the ratio of the norm to the energy of the groups.