Comfort Noise Generation for Low Bit-Rate Audio Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies face challenges in efficiently coding noisy speech at low bit-rates, leading to artifacts and decreased audio quality due to the inability to effectively separate and code speech and background noise using single-source models.
Innovation Solution
A decoder and encoder system that estimates the noise level and spectral shape of background noise, generates artificial comfort noise, and combines it with the decoded audio signal to mask coding artifacts and enhance quality, while also optionally using noise reduction schemes to improve signal-to-noise ratio.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If single-source speech coding models are used for noisy speech at low bit-rates, then bit-rate is reduced, but audio quality deteriorates due to artifacts and inability to separate speech from background noise
Solution Approach 1:
The patent segments the audio signal into two separate sources: speech signal and background noise. The encoder separately codes speech parameters and noise parameters, allowing independent optimization of each source. This segmentation enables the decoder to reconstruct both components and combine them, maintaining audio quality while operating at low bit-rates by only transmitting essential parameters for each source rather than full audio data.
Solution Approach 2:
The patent introduces a dual-mode coding structure that acts as an intermediary between single-source coding and full-bandwidth transmission. The system uses voice activity detection to determine whether to activate speech-only mode or speech-plus-noise mode, providing a flexible intermediate solution that adapts to signal conditions and maintains quality across varying bit-rate requirements.
2Reliability
If noise reduction techniques are applied to enhance speech intelligibility, then signal-to-noise ratio improves, but naturalness deteriorates due to distortion and musical noise artifacts
Solution Approach 1:
The patent extracts the background noise component from the mixed audio signal and codes it separately from the speech signal. By identifying and separating noise parameters through noise estimation algorithms, the system can represent noise independently, allowing the decoder to reconstruct the original noise characteristics without applying aggressive noise reduction that would distort speech or create artifacts.
Solution Approach 2:
The patent changes the representation parameters by coding noise characteristics (such as spectral shape, level, and temporal properties) separately from speech parameters. This parameter-based approach allows flexible control over noise reconstruction, enabling the system to maintain natural noise characteristics while enhancing speech intelligibility through selective noise management rather than aggressive filtering.
3Productivity
If coarse quantization is used for coding parameters at low bit-rates, then transmission efficiency improves, but perceptual quality worsens due to time fluctuations and annoying artifacts
Solution Approach 1:
The patent segments the parameter coding into speech-specific parameters and noise-specific parameters, each with their own quantization schemes. This allows optimization of quantization precision for each parameter type based on its perceptual importance and variability, improving overall efficiency while maintaining quality. Critical speech parameters can use finer quantization while less critical noise parameters use coarser quantization.
Solution Approach 2:
The patent employs adaptive quantization strategies where the precision of parameter coding is dynamically adjusted based on signal conditions, bit-rate availability, and perceptual importance. By changing quantization parameters adaptively and using entropy coding techniques, the system achieves efficient compression while minimizing perceptible artifacts from quantization.
Data Source
AI summary
The invention provides a decoder being configured for processing an encoded audio bitstream, wherein the decoder includes: a bitstream decoder configured to derive a decoded audio signal from the bitstream, wherein the decoded audio signal includes at least one decoded frame; a noise estimation device configured to produce a noise estimation signal containing an estimation of the level and/or the spectral shape of a noise in the decoded audio signal; a comfort noise generating device configured to derive a comfort noise signal from the noise estimation signal; and a combiner configured to combine the decoded frame of the decoded audio signal and the comfort noise signal in order to obtain an audio output signal.


