CELP Decoder Excitation Modification for Low-Bitrate Non-Speech Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech codecs struggle to maintain high speech quality at low bit rates, especially for sound signals different from clean speech, due to suboptimal coding modes and the difficulty of modifying standardized bitstreams without breaking interoperability.
Innovation Solution
A device and method for modifying the synthesis of time-domain excitation decoded by a time-domain decoder, using a multimodal decoding approach that categorizes the signal as inactive speech, active voiced, active unvoiced, or generic audio, and applies specific frequency domain modifications, such as normalization and noise addition, to enhance perceived quality while maintaining interoperability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a standardized speech codec is deployed without multi-modal coding, then interoperability is maintained, but speech quality deteriorates for non-clean speech signals at low bit rates
Solution Approach 1:
The patent segments the speech signal into different categories (voiced, unvoiced, onset) and applies different coding modes to each category. This segmentation allows the decoder to improve quality for specific signal types without requiring encoder modifications, thus maintaining interoperability while enhancing speech quality for non-clean speech signals.
Solution Approach 2:
Instead of modifying the encoder to implement multi-modal coding (which would break interoperability), the patent inverts the approach by implementing the multi-modal decoding capability solely in the decoder. This inversion allows quality improvement without affecting the standardized encoder, resolving the contradiction between interoperability and speech quality.
2Manufacturing precision
If multi-modal coding is implemented in the encoder, then speech quality improves for different signal categories, but interoperability is broken due to standardized bitstream requirements
Solution Approach 1:
The patent inverts the conventional approach by implementing multi-modal coding capability only in the decoder rather than the encoder. This allows the system to achieve improved speech quality through category-specific decoding modes while maintaining full interoperability with standardized encoders that produce conventional bitstreams.
Solution Approach 2:
The decoder performs self-service by autonomously categorizing and processing different signal types using its internal multi-modal decoding capabilities, without requiring any modifications to the encoder or changes to the standardized bitstream format. This enables quality improvement while preserving interoperability.
3Manufacturing precision
If noise compensation is applied to all signal types, then perceived quality improves, but processing complexity increases
Solution Approach 1:
The patent applies noise compensation locally and selectively based on signal category rather than uniformly to all signals. Different coding modes with appropriate noise compensation are applied only to specific categories (voiced, unvoiced, onset), which improves perceived quality while avoiding unnecessary processing complexity for signals that don't require enhancement.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method and device for modifying a synthesis of a time-domain excitation decoded by a time-domain decoder, wherein the synthesis of the decoded time-domain excitation is classified into one of a number of categories. The decoded time-domain excitation is converted into a frequency-domain excitation, and the frequency-domain excitation is modified as a function of the category in which the synthesis of the decoded time-domain excitation is classified. The modified frequency-domain excitation is converted into a modified time-domain excitation, and a synthesis filter is supplied with the modified time-domain excitation to produce a modified synthesis of the decoded time-domain excitation.