Speech Bitstream Decoding for Stable Redundant Frame Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing redundancy coding algorithms in VoIP systems often result in signal instability due to low bit rate encoding, leading to poor quality speech/audio signals, especially under conditions of high packet loss and delay jitter.
Innovation Solution
A speech/audio bitstream decoding method that acquires and post-processes decoding parameters from current and adjacent frames to recover stable speech/audio signals, using techniques such as adaptive weighting and attenuation of parameters like spectral pairs and adaptive codebook gains, to improve signal quality during frame transitions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If redundancy coding algorithm encodes speech/audio frame information at lower bit rate, then packet loss compensation capability is improved, but signal stability deteriorates
Solution Approach 1:
The patent implements dynamic switching between normal decoding mode and redundancy decoding mode based on packet loss detection. The decoder adapts its operation by selecting different decoding paths dynamically, transitioning from standard high-quality decoding to redundancy-based reconstruction when packet loss is detected, thereby maintaining signal stability while providing packet loss compensation.
Solution Approach 2:
The patent changes the decoding parameters and processing path based on the detected packet loss condition. When packet loss is detected, the system switches to using redundancy information with different decoding parameters, adjusting the reconstruction process to maintain signal quality and stability under degraded conditions.
2Stability of the object's composition
If jitter buffer buffers lost speech/audio frames, then delay jitter is smoothed, but transmission delay increases
Solution Approach 1:
The patent embeds redundancy information in advance within the bitstream structure, preparing compensation data before packet loss occurs. This preliminary action allows the decoder to reconstruct lost frames using pre-included redundancy data without needing to buffer additional frames, thus smoothing delay jitter without significantly increasing transmission delay.
3Adaptability or versatility
If multiple routers are used in VoIP transmission, then network routing flexibility is improved, but transmission delay jitter increases
Solution Approach 1:
The patent incorporates redundancy information as a cushioning mechanism against the instability caused by multi-router transmission. The embedded redundancy data provides a safety buffer that compensates for the delay jitter and packet loss introduced by dynamic routing through multiple routers, maintaining transmission stability despite network flexibility.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
Embodiments of the present invention disclose a speech/audio bitstream decoding method and apparatus. The speech/audio bitstream decoding method may include: acquiring a speech/audio decoding parameter of a current speech/audio frame, where the foregoing current speech/audio frame is a redundant decoded frame or a speech/audio frame previous to the foregoing current speech/audio frame is a redundant decoded frame; performing post processing on the speech/audio decoding parameter of the foregoing current speech/audio frame according to speech/audio parameters of X speech/audio frames, to obtain a post-processed speech/audio decoding parameter of the foregoing current speech/audio frame, where the foregoing X speech/audio frames include M speech/audio frames previous to the foregoing current speech/audio frame and/or N speech/audio frames next to the foregoing current speech/audio frame, and M and N are positive integers; and recovering a speech/audio signal of the foregoing current speech/audio frame by using the post-processed speech/audio decoding parameter of the foregoing current speech/audio frame. The technical solutions of the present invention help improve quality of an output speech/audio signal.