Audio Encoding Bandwidth Extension via Energy Offset Quantization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding technologies face inefficiencies in low-bitrate coding, particularly for speech signals with high-frequency components, and struggle to find effective Bandwidth Extension (BWE) encoding schemes that work for all types of audio signals.
Innovation Solution
The method involves transforming audio signals into a domain, selecting energy offsets for high-frequency bands based on energy measures from lower frequency bands, and using these offsets to encode and decode audio signals efficiently, allowing for scalar quantization of spectrum envelopes in subbands, thereby optimizing bit-rate and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If CELP coding is used for low-bitrate speech encoding, then speech quality is improved, but high-frequency components are not modeled well
Solution Approach 1:
The audio signal is divided into two separate bands: low band (0-6.4 kHz) encoded by CELP and high band (6.4-11.6 kHz) encoded by transform coding. This segmentation allows each coding method to specialize in its optimal frequency range, with CELP handling speech-quality requirements and transform coding preserving high-frequency components.
Solution Approach 2:
Different coding schemes are applied to different frequency regions based on their specific requirements. The low band uses CELP for speech quality while the high band uses transform coding for frequency accuracy. Additionally, the high band is further divided into two sub-bands with different quantization strategies to optimize local quality in each region.
2Reliability
If transform coding is used for general audio encoding, then moderate and high-bitrate quality is improved, but low-bitrate speech efficiency decreases
Solution Approach 1:
The frequency spectrum is segmented into low band and high band, with each band using the most appropriate coding method. The low band uses CELP for speech efficiency while the high band uses transform coding for audio quality, achieving both goals simultaneously through band-specific optimization.
3Productivity
If Bandwidth Extension is implemented to reduce bit-rate, then transmission efficiency is improved, but encoding complexity increases
Solution Approach 1:
The high band is divided into two sub-bands (first high band above low band, and second high band between low band and first high band). Each sub-band uses different energy reference strategies, simplifying the overall encoding process by breaking down the complex BWE task into manageable segments with specialized handling.
Solution Approach 2:
The patent uses different energy reference parameters for different sub-bands. The first sub-band uses energy from a first reference band in the low band, while the second sub-band uses energy from a second reference band. This parameter differentiation simplifies the encoding model and reduces complexity compared to a uniform approach.
Data Source
Figure 1~2A
Figure 2B~3A
Figure 3B~13
AI summary
A method for encoding of an audio signal comprises performing (214) of a transform of the audio signal. An energy offset is selected (216) for each of the first subbands. An energy measure of a first reference band within a low band of an encoding of a synthesis signal is obtained (212). The first high band is encoded (220) by providing quantization indices representing a respective scalar quantization of a spectrum envelope in the first subbands of the first high band relative to the energy measure of the first reference band by use of the selected energy offset. An encoder apparatus comprises means for carrying out the steps of the method. Corresponding decoder methods and apparatuses are also described.