Speech Coding Pitch Gain Adaptation for Packet Loss
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Parametric speech coding methods face challenges in maintaining speech coding quality during both good and bad channel conditions, particularly due to error propagation caused by packet loss, which is exacerbated by high pitch gains in voiced speech.
Innovation Solution
The method involves classifying speech frames into different classes based on pitch characteristics, limiting or reducing the pitch gain for the first subframe within a frame, and adjusting the code-excitation codebook size or adding an extra excitation component to compensate for the reduced energy, while maintaining regular CELP or analysis-by-synthesis approaches for other subframes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high pitch gain is used in voiced speech coding, then long-term prediction accuracy is improved, but error propagation during packet loss is exacerbated
Solution Approach 1:
The patent applies dynamics by making the pitch gain adaptive rather than fixed. The pitch gain is dynamically adjusted based on the subframe index within a frame, reducing it for later subframes where error propagation is more problematic, while maintaining higher values for earlier subframes where long-term prediction is most beneficial. This dynamic adjustment resolves the contradiction between prediction accuracy and error propagation resistance.
Solution Approach 2:
The patent segments the speech frame into multiple subframes and applies different pitch gain values to different subframes. By dividing the frame into subframes and treating them differently based on their position, the system can optimize pitch gain for each segment, reducing it for later subframes to minimize error propagation while maintaining it for earlier subframes to preserve prediction accuracy.
2Reliability
If pitch gain is reduced for later subframes, then error propagation is minimized, but coding quality may deteriorate
Solution Approach 1:
The patent applies local quality by making the pitch gain characteristic specific to each subframe's position within the frame. Instead of using a uniform pitch gain for the entire frame, the system assigns different pitch gain values to different subframes based on their local characteristics and error propagation risks. This allows optimization of both error propagation resistance and coding quality for each local region.
Solution Approach 2:
The patent changes the pitch gain parameter across different subframes within a frame. By varying this key parameter based on subframe position, the system can reduce pitch gain for later subframes to minimize error propagation while maintaining higher values for earlier subframes to preserve coding quality, thus resolving the contradiction between reliability and precision.
3Manufacturing precision
If code-excitation codebook size is increased to compensate for reduced pitch gain energy, then coding quality is maintained, but bit rate increases
Solution Approach 1:
The patent applies partial action by selectively increasing the code-excitation codebook size only for subframes where pitch gain is reduced, rather than uniformly increasing it for all subframes. This partial approach allows quality compensation where needed while avoiding unnecessary bit rate increase in subframes where full pitch gain is maintained, thus resolving the contradiction between quality and bit rate.
Data Source
AI summary
A speech coding method of reducing error propagation due to voice packet loss, is achieved by limiting or reducing a pitch gain only for the first subframe or the first two subframes within a speech frame, the excitation of a next frame is obtained according to the reduced or limited pitch gain value of the first subframe, and the next frame is encoded according to the obtained excitation. The method is used for a voiced speech class.


