Time-Warping Frames in Wideband Vocoder for Packet Jitter
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing time-warping methods in vocoders do not effectively address the asynchronous arrival of vocoder packets in packet-switched networks, leading to suboptimal quality and increased computational load, especially when time-warping is performed outside the vocoder.
Innovation Solution
The method involves time-warping speech frames by manipulating the residual low band signal before synthesis and the high band signal after synthesis in a Fourth Generation Vocoder (4GV) wideband vocoder, using Code-Excited Linear Prediction (CELP) and Noise-Excited Linear Prediction (NELP) techniques, allowing for expansion or compression of speech frames to mitigate delay jitter in packet-switched networks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If time-warping is performed outside the vocoder, then flexibility in handling asynchronous packets is improved, but quality of warped frames deteriorates and computational load increases
Solution Approach 1:
The vocoder is divided into distinct functional modules: analysis filterbank, LPC analysis, residual calculation, time-warping engine, and synthesis filterbank. This segmentation allows the time-warping operation to be integrated at the optimal point in the signal processing chain, maintaining quality while providing flexibility for asynchronous packet handling.
Solution Approach 2:
The time-warping operation is performed on the residual signal before synthesis, preparing the signal in advance for the correct temporal alignment. This preliminary action ensures that when the speech is synthesized, it is already time-warped to the correct duration, avoiding the need for post-synthesis time-warping that would degrade quality.
2Adaptability or versatility
If time-warping is performed outside the vocoder, then flexibility in handling asynchronous packets is improved, but computational load increases
Solution Approach 1:
The computational task of time-warping is segmented and integrated into the existing vocoder architecture, sharing computational resources with other vocoder functions. This avoids duplicating the entire vocoder pipeline outside the vocoder, reducing overall computational load while maintaining flexibility.
Solution Approach 2:
The time-warping engine is designed as a multi-functional component that can operate on residual signals from different vocoder modes (CELP, NELP, silence coding) and packet types. This universal design reduces the need for separate processing paths, lowering computational overhead while handling diverse asynchronous packet scenarios.
3Productivity
If split-band technique is used to encode lower and upper bands separately, then efficiency in bandwidth utilization is improved, but complexity of time-warping operation increases
Solution Approach 1:
The speech signal is segmented into lower band (0-3.5 kHz) and upper band (3.5-7 kHz) components that are processed separately through the vocoder pipeline. Each band undergoes independent time-warping based on its own residual signal, allowing efficient bandwidth utilization while managing complexity through modular processing.
Solution Approach 2:
Different time-warping parameters and operations are applied to different frequency bands according to their specific characteristics. The lower band and upper band can have different pitch periods and time-warping factors, optimizing quality for each band while maintaining overall system efficiency.
Data Source
AI summary
A method of communicating speech comprising time-warping a residual low band speech signal to an expanded or compressed version of the residual low band speech signal, time-warping a high band speech signal to an expanded or compressed version of the high band speech signal, and merging the time-warped low band and high band speech signals to give an entire time-warped speech signal. In the low band, the residual low band speech signal is synthesized after time-warping of the residual low band signal while in the high band, an unwarped high band signal is synthesized before time-warping of the high band speech signal. The method may further comprise classifying speech segments and encoding the speech segments. The encoding of the speech segments may be one of code-excited linear prediction, noise-excited linear prediction or ⅛ frame (silence) coding.


