CELP Decoder Excitation Classification for Low-Bitrate Audio Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech codecs face suboptimal quality at low bit rates for sound signals other than clean speech due to lack of multimodal decoding, and modifying encoders is difficult without breaking interoperability.

Innovation Solution

A device and method that classifies time-domain excitation into categories, converts it to frequency-domain, modifies it based on category, and synthesizes it back to time-domain to improve quality, maintaining interoperability by using a classifier, converter, modifier, and synthesis filter.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a standardized bitstream is used for interoperability, then codec compatibility is maintained, but the ability to modify the encoder to improve quality is lost

Engineering Contradiction:
Improvecodec interoperabilityVSAvoidencoder modification capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system separates the encoder and decoder modification capabilities. The encoder remains standardized and unmodified to maintain interoperability, while only the decoder is modified to perform classification and frequency-domain processing. This segmentation allows quality improvement without breaking compatibility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing stage in the decoder that classifies decoded excitation into categories (voiced, unvoiced, onset) and applies category-specific modifications in the frequency domain. This intermediary layer enables adaptive quality improvement while maintaining compatibility with standardized encoders.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If a single decoding mode is used for all sound categories, then decoder complexity is reduced, but quality for non-clean speech at low bit rates deteriorates

Engineering Contradiction:
Improvedecoder structureVSAvoidspeech quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The decoder dynamically adapts its processing based on the classified category of the decoded excitation. Different frequency-domain modification strategies are applied depending on whether the signal is voiced, unvoiced, or onset, allowing optimal quality for each category while using a single unified decoder structure.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes processing parameters based on classification results. For voiced categories, different frequency modifications are applied compared to unvoiced or onset categories. This parameter adaptation enables high quality across diverse sound types without requiring multiple separate decoders.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP2774145B1Improving non-speech content for low rate CELP decoder
Publication Date: 2020.06.17 VOICEAGE EVS LLC
  • EP2774145B1 patent drawingFigure 1
  • EP2774145B1 patent drawingFigure 2
  • EP2774145B1 patent drawingFigure 3

AI summary

A method and device for modifying a synthesis of a time-domain excitation decoded by a time-domain decoder, wherein the synthesis of the decoded time- domain excitation is classified into one of a number of categories. The decoded time-domain excitation is converted into a frequency-domain excitation, and the frequency-domain excitation is modified as a function of the category in which the synthesis of the decoded time-domain excitation is classified. The modified frequency-domain excitation is converted into a modified time-domain excitation, and a synthesis filter is supplied with the modified time-domain excitation to produce a modified synthesis of the decoded time-domain excitation.