Vocoder Residual Time-Warping for Packet Network Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing vocoder technologies face challenges in efficiently time-warping speech frames to address asynchronous packet arrival in packet-switched networks, leading to suboptimal quality and increased computational load when performed outside the vocoder.
Innovation Solution
The method involves time-warping speech frames by manipulating the residual speech signal within the vocoder, using techniques such as code-excited linear prediction, noise-excited linear prediction, and prototype pitch period encoding, which allows for expansion or compression of speech segments by estimating and adjusting pitch periods, and applying gains to different parts of the signal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If time-warping is performed outside the vocoder, then implementation flexibility is improved, but quality of warped frames deteriorates and computational load increases
Solution Approach 1:
The patent introduces the residual signal as an intermediary element that carries timing information. By manipulating this intermediary (the residual) rather than the full speech signal, the system achieves time-warping with reduced computational load while maintaining quality. The residual acts as a mediator between the original signal and the warped output.
Solution Approach 2:
The patent extracts the essential timing information from the full speech signal by working only with the residual component. This extraction allows time-warping to be performed on a simplified representation, reducing computational complexity while preserving the quality of the warped frames when reconstructed through the vocoder.
2Adaptability or versatility
If time-warping is performed outside the vocoder, then implementation flexibility is improved, but computational load increases
Solution Approach 1:
The patent extracts only the necessary residual information required for time-warping, discarding redundant components. This extraction principle reduces the data volume that needs processing, thereby lowering computational load while maintaining implementation flexibility.
Solution Approach 2:
By using the residual signal as an intermediary, the patent performs time-warping operations on a simplified representation rather than the full speech signal. This intermediary approach significantly reduces computational requirements while preserving the ability to achieve desired time-warping effects.
3Measurement precision
If pitch periods are overlapped during compression, then synchronization accuracy is improved, but processing complexity increases
Solution Approach 1:
The patent applies preliminary weighting to pitch periods before overlapping them. This preliminary action prepares the signal segments in advance, ensuring that when they are overlapped, the synchronization is accurate and artifacts are minimized. The weighting is calculated and applied beforehand, simplifying the overall processing.
Solution Approach 2:
The patent applies different weighting factors to different parts of the pitch periods based on their local characteristics. This local quality approach ensures that each region contributes appropriately to the overlapped output, improving synchronization accuracy while managing complexity through localized rather than global processing.
Data Source
AI summary
In one embodiment, the present invention comprises a vocoder having at least one input and at least one output, an encoder comprising a filter having at least one input operably connected to the input of the vocoder and at least one output, a decoder comprising a synthesizer having at least one input operably connected to the at least one output of the encoder, and at least one output operably connected to the at least one output of the vocoder, wherein the encoder comprises a memory and the encoder is adapted to execute instructions stored in the memory comprising classifying speech segments and encoding speech segments, and the decoder comprises a memory and the decoder is adapted to execute instructions stored in the memory comprising time-warping a residual speech signal to an expanded or compressed version of the residual speech signal.


