VoIP Fill Frame Generation Using Speech Parameter History Registers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice over Internet Protocol (VoIP) systems experience voice quality degradation due to lost and late packets, which are not addressed by the 'fire and forget' User Datagram Protocol (UDP) protocol, leading to gaps, delays, and garbled speech.
Innovation Solution
A method and apparatus for generating fill frames in VoIP applications using a speech synthesizer, which determines frame loss and utilizes Line Spectral Frequencies (LSF), Voicing Cutoff (VCUT), pitch, and Root Mean Squared (RMS) gain parameters to create a speech signal, shifting and interpolating values from history registers to maintain continuous audio when packets are lost.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If UDP protocol is used for VoIP packet transmission, then transmission speed and simplicity are improved, but packet loss and reliability deteriorate
Solution Approach 1:
The patent introduces fill frame generation as an intermediary mechanism between packet loss and audio output. When packets are lost, the system generates synthetic fill frames using historical speech parameters (LSF, VCUT, pitch, RMS) to bridge the gaps, thus mediating the impact of unreliable UDP transmission on audio quality without changing the underlying transmission protocol.
Solution Approach 2:
The system changes the parameters of the audio signal by using historical speech parameters (Line Spectral Frequencies, Voicing Cutoff, pitch, RMS gain) to synthesize new fill frames. This parameter transformation allows the system to maintain audio continuity despite packet loss, adapting the speech characteristics to match the expected pattern during lost frames.
2Reliability
If packet loss mitigation is implemented, then voice quality is improved, but system complexity increases
Solution Approach 1:
The system performs preliminary action by continuously maintaining history registers that store recent speech parameters (LSF, VCUT, pitch, RMS) from received packets. This pre-computed historical data is ready to be used immediately when packet loss is detected, eliminating the need for complex real-time analysis during the loss event and reducing overall system complexity.
Solution Approach 2:
The patent uses copying by creating fill frames that replicate the characteristics of previous valid frames. Instead of generating entirely new speech content, the system copies and adapts historical speech parameters to produce fill frames that match the expected speech pattern, simplifying the mitigation process while maintaining voice quality.
3Stability of the object's composition
If fill frame generation is implemented, then voice continuity is improved, but processing time increases
Solution Approach 1:
The system applies partial action by generating fill frames only when and where packet loss occurs, rather than continuously processing all audio frames. The fill frame generation is triggered selectively based on the frame loss flag, reducing unnecessary processing time while maintaining voice continuity only when needed to address actual packet loss events.
Data Source
AI summary
A method and apparatus that generates fill frames for Voice over Internet Protocol (VoIP) applications in a communication device is disclosed. The method may include determining if there is a lost frame in a received communication, wherein if it is determined that there is a lost frame, setting a frame loss flag and storing the frame loss flag in the frame loss history register, shifting a loss history register, a line spectral frequency (LSF) history register, a voicing cutoff (VCUT) history register, a pitch history register, and a root mean squared (RMS) gain history register, wherein the loss history register, the LSF history register, the VCUT history register, the pitch history register, and the RMS history register include at least three registers, the three registers being a newest, a middle and an oldest registers, reading the frame loss flag into a newest loss history register, determining contents of the middle register of each of the LSF history register, the VCUT history register, the pitch history register, and the RMS history register, and sending the contents of the middle registers to a synthesizer to generate an output speech signal.


