Audio Encoder Bandwidth Extension for Low-Bitrate High-Band Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice encoding technologies face challenges in achieving high-quality encoding of voice signals with wide frequency bands at low bit rates, leading to noise generation and subjective quality reduction due to the replication of low-band spectra in high-band regions.
Innovation Solution
The proposed solution involves separating tonal and non-tonal components of the voice signal for individual encoding, using noise addition and bandwidth extension techniques to accurately reproduce energy levels, thereby reducing bit rate while maintaining high-quality encoding and decoding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of substance
If the low-band spectrum is replicated as the high-band spectrum without modification, then the encoding bit rate is reduced, but noise resembling a ringing bell is generated and subjective quality deteriorates
Solution Approach 1:
The patent applies local quality by differentiating between tonal and non-tonal components within the spectrum. Non-tonal components are replicated from low-band to high-band, while tonal components are processed differently (transposed or synthesized). This selective processing based on local spectral characteristics prevents the ringing bell noise that occurs when all components are uniformly replicated, while still achieving bit rate reduction.
Solution Approach 2:
The spectrum is segmented into tonal and non-tonal components using tone detection. This segmentation allows different processing strategies to be applied to different parts of the spectrum: non-tonal components are replicated for efficiency, while tonal components are handled separately to maintain quality. This resolves the contradiction by applying appropriate processing to each segment.
2Device complexity
If all components including both tonal and non-tonal components are used for dynamic range definition, then the encoding process is simplified, but the quality is not optimized because tonal components require different handling
Solution Approach 1:
The encoding process is segmented into tone detection, tonal component extraction, non-tonal component processing, and high-band synthesis. This segmentation increases processing complexity but enables optimized handling of different spectral components, ultimately improving encoding quality by applying appropriate processing to each component type rather than treating all components uniformly.
Solution Approach 2:
Different quality processing is applied to different components: non-tonal components use simple replication while tonal components use transposition or synthesis. This local quality approach optimizes the overall encoding quality by ensuring each component type is processed according to its characteristics, resolving the contradiction between simplicity and quality.
3Productivity
If the low-band spectrum is replicated at high band, then encoding efficiency is improved, but the energy levels and tonal accuracy in the high-band region are not accurately reproduced
Solution Approach 1:
The non-tonal spectral envelope from the low-band is copied to the high-band region. This copying provides a efficient basis for high-band reconstruction while maintaining the characteristic noise-like properties of non-tonal components. Combined with tonal component processing, this achieves both encoding efficiency and accurate energy level reproduction.
Solution Approach 2:
The tonal component extraction and processing acts as an intermediary between the low-band replication and the final high-band output. By separating and processing tonal components through transposition or synthesis, the system accurately reproduces energy levels and tonal characteristics in the high-band region while maintaining overall encoding efficiency.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An encoder according to the present disclosure includes a first encoding unit that generates a first encoded signal in which a low-band signal having a frequency lower than or equal to a predetermined frequency from a voice or audio input signal is encoded, and a low-band decoded signal; a second encoding unit that encodes, on the basis of the low-band decoded signal, a high-band signal having a band higher than that of the low-band signal to generate a high-band encoded signal; and a first multiplexing unit that multiplexes the first encoded signal and the high-band encoded signal to generate and output an encoded signal. The second encoding unit calculates an energy ratio between a high-band noise component, which is a noise component of the high-band signal, and a high-band non-tonal component of a high-band decoded signal generated from the low-band decoded signal and outputs the calculated ratio as the high-band encoded signal.