CELP Decoder Excitation Classification for Low-Bitrate Audio Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech codecs face suboptimal quality at low bit rates for sound signals other than clean speech due to lack of multimodal decoding, and modifying encoders is difficult without breaking interoperability.
Innovation Solution
A device and method that classifies time-domain excitation into categories, converts it to frequency-domain, modifies it based on category, and synthesizes it back to time-domain to improve quality, maintaining interoperability by using a classifier, converter, modifier, and synthesis filter.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a standardized bitstream is used for interoperability, then codec compatibility is maintained, but the ability to modify the encoder to improve quality is lost
Solution Approach 1:
The system separates the encoder and decoder modification capabilities. The encoder remains standardized and unmodified to maintain interoperability, while only the decoder is modified to perform classification and frequency-domain processing. This segmentation allows quality improvement without breaking compatibility.
Solution Approach 2:
The patent introduces an intermediary processing stage in the decoder that classifies decoded excitation into categories (voiced, unvoiced, onset) and applies category-specific modifications in the frequency domain. This intermediary layer enables adaptive quality improvement while maintaining compatibility with standardized encoders.
2Device complexity
If a single decoding mode is used for all sound categories, then decoder complexity is reduced, but quality for non-clean speech at low bit rates deteriorates
Solution Approach 1:
The decoder dynamically adapts its processing based on the classified category of the decoded excitation. Different frequency-domain modification strategies are applied depending on whether the signal is voiced, unvoiced, or onset, allowing optimal quality for each category while using a single unified decoder structure.
Solution Approach 2:
The system changes processing parameters based on classification results. For voiced categories, different frequency modifications are applied compared to unvoiced or onset categories. This parameter adaptation enables high quality across diverse sound types without requiring multiple separate decoders.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method and device for modifying a synthesis of a time-domain excitation decoded by a time-domain decoder, wherein the synthesis of the decoded time- domain excitation is classified into one of a number of categories. The decoded time-domain excitation is converted into a frequency-domain excitation, and the frequency-domain excitation is modified as a function of the category in which the synthesis of the decoded time-domain excitation is classified. The modified frequency-domain excitation is converted into a modified time-domain excitation, and a synthesis filter is supplied with the modified time-domain excitation to produce a modified synthesis of the decoded time-domain excitation.