Time-Domain Decoder Spectral Masking to Reduce Quantization Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
State-of-the-art conversational codecs struggle to render music with good quality at low bitrates, as they are optimized for speech and modifications to the bitstream can break interoperability.
Innovation Solution
A device and method that converts time-domain excitation into frequency-domain excitation, applies a weighting mask to retrieve spectral information lost in quantization noise, and modifies the frequency-domain excitation to increase spectral dynamics, all without adding coding delay.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech-model based codecs are used to render music, then speech quality is maintained at low bitrate, but music rendering quality deteriorates
Solution Approach 1:
The codec separates speech and music rendering paths by detecting signal type (voiced/unvoiced speech vs. music) and applying different processing modes. For music, it uses frequency-domain post-processing independent from the speech-optimized encoder, allowing music quality improvement without affecting speech codec compatibility
Solution Approach 2:
A frequency-domain post-processing stage acts as an intermediary between the speech-optimized encoder and the final audio output. This intermediate processing layer specifically targets music signals and applies spectral corrections without modifying the original speech-optimized bitstream, thus preserving speech quality while improving music rendering
2Manufacturing precision
If codec bitstream is modified to improve music rendering, then music quality improves, but interoperability is broken
Solution Approach 1:
The invention extracts the music improvement functionality from the core speech codec by implementing it as a separate post-processing stage. The standardized speech codec bitstream remains unchanged and interoperable, while the extracted post-processing layer handles music-specific optimizations independently
Solution Approach 2:
The system dynamically adapts its behavior based on signal type detection. For speech signals, it processes through the standard codec path maintaining interoperability. For music signals, it activates additional frequency-domain post-processing. This dynamic mode switching allows the same device to maintain codec compatibility while improving music rendering quality
3Manufacturing precision
If frequency-domain post-processing is applied to reduce quantization noise, then spectral dynamics improve, but processing delay increases
Solution Approach 1:
The system performs preliminary spectral analysis and mask generation using historical frame data before applying the final frequency-domain correction. By pre-computing the weighting mask based on past frames' spectral characteristics, the real-time processing requirement is reduced and delay is minimized
Solution Approach 2:
Instead of applying full spectral processing to all frequency components, the invention selectively applies corrections only to specific frequency bands where quantization noise is most prominent. This partial action approach reduces processing complexity and delay while maintaining effective noise reduction in the critical frequency ranges
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
The present disclosure relates to a device and method for reducing quantization noise in a signal contained in a time-domain excitation decoded by a time-domain decoder. The decoded time-domain excitation is converted into a frequency-domain excitation. A weighting mask is produced for retrieving spectral information lost in the quantization noise. The frequency-domain excitation is modified to increase spectral dynamics by application of the weighting mask. The modified frequency-domain excitation is converted into a modified time-domain excitation. The method and device can be used for improving music content rendering of linear-prediction (LP) based codecs. Optionally, a synthesis of the decoded time-domain excitation may be classified into one of a first set of excitation categories and a second set of excitation categories, the second set including INACTIVE or UNVOICED categories, the first set including an OTHER category.