CELP Decoder Excitation Modification for Low-Bitrate Non-Speech Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech codecs struggle to maintain high speech quality at low bit rates, especially for sound signals different from clean speech, due to suboptimal coding modes and the difficulty of modifying standardized bitstreams without breaking interoperability.

Innovation Solution

A device and method for modifying the synthesis of time-domain excitation decoded by a time-domain decoder, using a multimodal decoding approach that categorizes the signal as inactive speech, active voiced, active unvoiced, or generic audio, and applies specific frequency domain modifications, such as normalization and noise addition, to enhance perceived quality while maintaining interoperability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a standardized speech codec is deployed without multi-modal coding, then interoperability is maintained, but speech quality deteriorates for non-clean speech signals at low bit rates

Engineering Contradiction:
ImproveinteroperabilityVSAvoidspeech quality
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent segments the speech signal into different categories (voiced, unvoiced, onset) and applies different coding modes to each category. This segmentation allows the decoder to improve quality for specific signal types without requiring encoder modifications, thus maintaining interoperability while enhancing speech quality for non-clean speech signals.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of modifying the encoder to implement multi-modal coding (which would break interoperability), the patent inverts the approach by implementing the multi-modal decoding capability solely in the decoder. This inversion allows quality improvement without affecting the standardized encoder, resolving the contradiction between interoperability and speech quality.

Inventive Principle:
Principle #13The other way round (Inversion)

2Manufacturing precision

If multi-modal coding is implemented in the encoder, then speech quality improves for different signal categories, but interoperability is broken due to standardized bitstream requirements

Engineering Contradiction:
Improvespeech qualityVSAvoidinteroperability
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent inverts the conventional approach by implementing multi-modal coding capability only in the decoder rather than the encoder. This allows the system to achieve improved speech quality through category-specific decoding modes while maintaining full interoperability with standardized encoders that produce conventional bitstreams.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The decoder performs self-service by autonomously categorizing and processing different signal types using its internal multi-modal decoding capabilities, without requiring any modifications to the encoder or changes to the standardized bitstream format. This enables quality improvement while preserving interoperability.

Inventive Principle:
Principle #25Self-service

3Manufacturing precision

If noise compensation is applied to all signal types, then perceived quality improves, but processing complexity increases

Engineering Contradiction:
Improveperceived qualityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies noise compensation locally and selectively based on signal category rather than uniformly to all signals. Different coding modes with appropriate noise compensation are applied only to specific categories (voiced, unvoiced, onset), which improves perceived quality while avoiding unnecessary processing complexity for signals that don't require enhancement.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3709298B1Improving non-speech content for low rate CELP decoder
Publication Date: 2024.11.20 VOICEAGE EVS LLC
  • EP3709298B1 patent drawingFigure 1
  • EP3709298B1 patent drawingFigure 2
  • EP3709298B1 patent drawingFigure 3

AI summary

A method and device for modifying a synthesis of a time-domain excitation decoded by a time-domain decoder, wherein the synthesis of the decoded time-domain excitation is classified into one of a number of categories. The decoded time-domain excitation is converted into a frequency-domain excitation, and the frequency-domain excitation is modified as a function of the category in which the synthesis of the decoded time-domain excitation is classified. The modified frequency-domain excitation is converted into a modified time-domain excitation, and a synthesis filter is supplied with the modified time-domain excitation to produce a modified synthesis of the decoded time-domain excitation.