Speech Bitstream Decoding for Stable Redundant Frame Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing redundancy coding algorithms in VoIP systems often result in signal instability due to low bit rate encoding, leading to poor quality speech/audio signals, especially under conditions of high packet loss and delay jitter.

Innovation Solution

A speech/audio bitstream decoding method that acquires and post-processes decoding parameters from current and adjacent frames to recover stable speech/audio signals, using techniques such as adaptive weighting and attenuation of parameters like spectral pairs and adaptive codebook gains, to improve signal quality during frame transitions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If redundancy coding algorithm encodes speech/audio frame information at lower bit rate, then packet loss compensation capability is improved, but signal stability deteriorates

Engineering Contradiction:
Improvepacket loss compensation capabilityVSAvoidsignal stability
Core Design Contradiction:
ReliabilityVSStability of the object's composition

Solution Approach 1:

The patent implements dynamic switching between normal decoding mode and redundancy decoding mode based on packet loss detection. The decoder adapts its operation by selecting different decoding paths dynamically, transitioning from standard high-quality decoding to redundancy-based reconstruction when packet loss is detected, thereby maintaining signal stability while providing packet loss compensation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the decoding parameters and processing path based on the detected packet loss condition. When packet loss is detected, the system switches to using redundancy information with different decoding parameters, adjusting the reconstruction process to maintain signal quality and stability under degraded conditions.

Inventive Principle:
Principle #35Parameter changes

2Stability of the object's composition

If jitter buffer buffers lost speech/audio frames, then delay jitter is smoothed, but transmission delay increases

Engineering Contradiction:
Improvedelay jitter smoothingVSAvoidtransmission delay
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The patent embeds redundancy information in advance within the bitstream structure, preparing compensation data before packet loss occurs. This preliminary action allows the decoder to reconstruct lost frames using pre-included redundancy data without needing to buffer additional frames, thus smoothing delay jitter without significantly increasing transmission delay.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple routers are used in VoIP transmission, then network routing flexibility is improved, but transmission delay jitter increases

Engineering Contradiction:
Improvenetwork routing flexibilityVSAvoidtransmission delay stability
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent incorporates redundancy information as a cushioning mechanism against the instability caused by multi-router transmission. The embedded redundancy data provides a safety buffer that compensates for the delay jitter and packet loss introduced by dynamic routing through multiple routers, maintaining transmission stability despite network flexibility.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentEP3121812B1Voice frequency code stream decoding method and device
Publication Date: 2020.03.11 HUAWEI TECH CO LTD
  • EP3121812B1 patent drawingFigure 1
  • EP3121812B1 patent drawingFigure 2
  • EP3121812B1 patent drawingFigure 3~4

AI summary

Embodiments of the present invention disclose a speech/audio bitstream decoding method and apparatus. The speech/audio bitstream decoding method may include: acquiring a speech/audio decoding parameter of a current speech/audio frame, where the foregoing current speech/audio frame is a redundant decoded frame or a speech/audio frame previous to the foregoing current speech/audio frame is a redundant decoded frame; performing post processing on the speech/audio decoding parameter of the foregoing current speech/audio frame according to speech/audio parameters of X speech/audio frames, to obtain a post-processed speech/audio decoding parameter of the foregoing current speech/audio frame, where the foregoing X speech/audio frames include M speech/audio frames previous to the foregoing current speech/audio frame and/or N speech/audio frames next to the foregoing current speech/audio frame, and M and N are positive integers; and recovering a speech/audio signal of the foregoing current speech/audio frame by using the post-processed speech/audio decoding parameter of the foregoing current speech/audio frame. The technical solutions of the present invention help improve quality of an output speech/audio signal.