Time-Domain Decoder Spectral Masking to Reduce Quantization Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

State-of-the-art conversational codecs struggle to render music with good quality at low bitrates, as they are optimized for speech and modifications to the bitstream can break interoperability.

Innovation Solution

A device and method that converts time-domain excitation into frequency-domain excitation, applies a weighting mask to retrieve spectral information lost in quantization noise, and modifies the frequency-domain excitation to increase spectral dynamics, all without adding coding delay.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speech-model based codecs are used to render music, then speech quality is maintained at low bitrate, but music rendering quality deteriorates

Engineering Contradiction:
Improvespeech qualityVSAvoidmusic rendering quality
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The codec separates speech and music rendering paths by detecting signal type (voiced/unvoiced speech vs. music) and applying different processing modes. For music, it uses frequency-domain post-processing independent from the speech-optimized encoder, allowing music quality improvement without affecting speech codec compatibility

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A frequency-domain post-processing stage acts as an intermediary between the speech-optimized encoder and the final audio output. This intermediate processing layer specifically targets music signals and applies spectral corrections without modifying the original speech-optimized bitstream, thus preserving speech quality while improving music rendering

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If codec bitstream is modified to improve music rendering, then music quality improves, but interoperability is broken

Engineering Contradiction:
Improvemusic rendering qualityVSAvoidinteroperability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The invention extracts the music improvement functionality from the core speech codec by implementing it as a separate post-processing stage. The standardized speech codec bitstream remains unchanged and interoperable, while the extracted post-processing layer handles music-specific optimizations independently

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system dynamically adapts its behavior based on signal type detection. For speech signals, it processes through the standard codec path maintaining interoperability. For music signals, it activates additional frequency-domain post-processing. This dynamic mode switching allows the same device to maintain codec compatibility while improving music rendering quality

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If frequency-domain post-processing is applied to reduce quantization noise, then spectral dynamics improve, but processing delay increases

Engineering Contradiction:
Improvespectral dynamicsVSAvoidprocessing delay
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary spectral analysis and mask generation using historical frame data before applying the final frequency-domain correction. By pre-computing the weighting mask based on past frames' spectral characteristics, the real-time processing requirement is reduced and delay is minimized

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of applying full spectral processing to all frequency components, the invention selectively applies corrections only to specific frequency bands where quantization noise is most prominent. This partial action approach reduces processing complexity and delay while maintaining effective noise reduction in the critical frequency ranges

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP4246516B1Device and method for reducing quantization noise in a time-domain decoder
Publication Date: 2025.07.23 VOICEAGE EVS LLC
  • EP4246516B1 patent drawingFigure 1
  • EP4246516B1 patent drawingFigure 2A
  • EP4246516B1 patent drawingFigure 2B

AI summary

The present disclosure relates to a device and method for reducing quantization noise in a signal contained in a time-domain excitation decoded by a time-domain decoder. The decoded time-domain excitation is converted into a frequency-domain excitation. A weighting mask is produced for retrieving spectral information lost in the quantization noise. The frequency-domain excitation is modified to increase spectral dynamics by application of the weighting mask. The modified frequency-domain excitation is converted into a modified time-domain excitation. The method and device can be used for improving music content rendering of linear-prediction (LP) based codecs. Optionally, a synthesis of the decoded time-domain excitation may be classified into one of a first set of excitation categories and a second set of excitation categories, the second set including INACTIVE or UNVOICED categories, the first set including an OTHER category.