Vocoder Phase Matching Time-Warping Frames
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In packet-switched systems, de-jitter buffers can cause erasures between consecutive frames, leading to phase discontinuities and artifacts in voice decoders due to synchronization issues between the encoder and decoder.
Innovation Solution
The method involves phase matching and time-warping of speech frames to maintain signal continuity, where phase matching adjusts the frame's sample count to match the encoder and decoder phases, and time-warping expands or compresses frames to fill gaps caused by erasures, using techniques like overlap-add and interpolation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If de-jitter buffer is used to store and deliver frames in sequence, then frame delivery order is improved, but phase discontinuities and artifacts are introduced due to erasures between consecutive frames
Solution Approach 1:
The patent applies preliminary action by performing phase matching and time-warping operations on frames before they are delivered to the decoder. The de-jitter buffer pre-processes frames by adjusting their phase and timing to anticipate potential synchronization issues, rather than correcting them after artifacts appear. This prevents phase discontinuities before they degrade voice quality.
Solution Approach 2:
The patent changes temporal parameters of frames by modifying the number of speech samples in each frame through time-warping operations. By adjusting frame duration and sample count dynamically, the system adapts to varying network conditions and maintains phase continuity despite erasures, resolving the contradiction between reliable frame delivery and artifact prevention.
2Reliability
If erasures are inserted between consecutive frames, then packet loss handling is improved, but encoder and decoder fall out of sync causing voice quality degradation
Solution Approach 1:
The patent implements feedback mechanisms where the decoder monitors phase continuity and provides information about synchronization status. When erasures cause phase drift, the system uses feedback from previous frame phases to adjust subsequent frame processing, maintaining encoder-decoder sync despite packet losses. This closed-loop approach ensures voice quality is preserved even when erasures occur.
Solution Approach 2:
The system performs preliminary phase matching on incoming frames before decoding, using phase information from previously decoded frames to pre-adjust the timing and sample count of current frames. This anticipatory adjustment prevents synchronization loss before it occurs, allowing the system to handle erasures reliably while maintaining voice quality.
3Stability of the object's composition
If frame sample count is adjusted for phase matching, then phase continuity is improved, but frame duration may become inconsistent
Solution Approach 1:
The patent dynamically changes the temporal parameters of frames by adjusting sample count and duration based on phase requirements. Time-warping operations modify frame duration continuously to maintain phase continuity, accepting variable frame lengths as necessary to preserve the stability of phase composition across consecutive frames.
Solution Approach 2:
The system transitions from static frame processing to dynamic frame adjustment, where frame duration and sample count are flexible parameters that adapt to phase continuity requirements. This dynamic approach allows the system to maintain stable phase composition while accommodating necessary variations in frame timing and length.
Data Source
AI summary
In one embodiment, the present invention comprises a vocoder having at least one input and at least one output, an encoder comprising a filter having at least one input operably connected to the input of the vocoder and at least one output, a decoder comprising a synthesizer having at least one input operably connected to the at least one output of the encoder, and at least one output operably connected to the at least one output of the vocoder, wherein the decoder comprises a memory and the decoder is adapted to execute instructions stored in the memory comprising phase matching and time-warping a speech frame.


