Speech Coding Pitch Gain Adaptation for Packet Loss
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech coding methods, such as CELP, face significant error propagation due to packet loss during voice transmission, particularly for voiced speech, where a high pitch gain can lead to prolonged errors if the phase relationship between excitation components is disrupted, and simply setting the pitch gain to zero to mitigate this issue compromises quality or requires higher bit rates.
Innovation Solution
The proposed solution involves classifying speech signals into different categories and limiting the pitch gain for the first pitch cycle of each frame, using a larger coded excitation codebook or adding an extra stage of code-excitation for certain classes to reduce error propagation while maintaining significant long-term prediction contributions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the pitch gain is set to zero to mitigate error propagation from packet loss, then error propagation is reduced, but speech quality deteriorates or higher bit rates are required
Solution Approach 1:
The patent applies different pitch gain values to different parts of the speech frame. Specifically, the pitch gain is set to zero for the first pitch cycle (to prevent error propagation after packet loss) while maintaining normal pitch gain values for subsequent pitch cycles (to preserve speech quality). This local differentiation allows the system to simultaneously achieve error propagation resistance and maintain speech quality without requiring higher bit rates.
2Productivity
If the pitch gain is kept high to maintain long-term prediction contributions, then coding efficiency is improved, but error propagation duration increases
Solution Approach 1:
The patent segments the speech frame into different pitch cycles and applies different pitch gain values to each segment. The first pitch cycle has pitch gain set to zero to limit error propagation, while subsequent pitch cycles maintain normal pitch gain values to preserve coding efficiency. This segmentation allows the system to achieve both reduced error propagation duration and maintained coding efficiency.
3Manufacturing precision
If a larger coded excitation codebook is used to compensate for reduced pitch gain, then speech quality is maintained, but device complexity increases
Solution Approach 1:
The patent dynamically adjusts the pitch gain value based on the position within the speech frame and packet loss conditions. The system switches between different pitch gain values (zero for the first pitch cycle, normal values for subsequent cycles) based on real-time conditions. This dynamic adaptation allows the system to maintain speech quality without requiring a permanently larger codebook, thus avoiding increased device complexity.
Data Source
AI summary
A speech coding method of significantly reducing error propagation due to voice packet loss, while still greatly profiting from a pitch prediction or Long-Term Prediction (LTP), is achieved by limiting or reducing a pitch gain only for the first subframe or the first two subframes within a speech frame. The method is used for a speech class decided by a classification algorithm; the classification algorithm is designed, depending on at least one pitch cycle length compared to one subframe size. Speech coding quality loss due to the pitch gain reduction is compensated by increasing a coded excitation codebook size or adding one more stage of excitation only for the first subframe or the first two subframes within the speech frame.


