Speech Frame Loss Compensation Using Non-Cyclic Pulse Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech decoding methods for frame loss concealment often produce perceptually annoying effects, such as loud beep sounds or degraded articulation, due to the repeated use of preceding frames with high amplitude regions or background noise, leading to unnatural and noisy decoded speech.
Innovation Solution
A speech decoding apparatus that detects non-periodic pulse waveforms in lost frames and suppresses them using noise signals, maintaining the characteristics of the preceding frame while substituting noise for the excitation signal in specific regions, thereby preventing the generation of annoying sounds and ensuring natural decoded speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Stability of the object's composition
If the excitation signal from the preceding frame is repeatedly used for frame loss concealment, then the continuity of speech signal is maintained, but loud beep sounds and perceptually annoying effects are generated in regions with high amplitude or non-periodic pulse waveforms
Solution Approach 1:
The patent applies different processing to different regions of the excitation signal based on their characteristics. Regions identified as having non-periodic pulse waveforms or high amplitude are selectively modified using noise signals, while other regions maintain the original repeated excitation signal. This local differentiation resolves the contradiction by preserving continuity in stable regions while suppressing harmful artifacts in problematic regions.
Solution Approach 2:
The patent introduces noise signals as an intermediary element to mediate between the repeated excitation signal and the final output. By adding noise to specific regions of the excitation signal, the patent creates a transition that prevents the generation of loud beep sounds while maintaining overall signal continuity. The noise acts as a buffer that smooths out the harmful artifacts.
2Object-generated harmful factors
If noise signal is added to excitation signal from noise codebook for unvoiced frame concealment, then perceptually strong annoying effects are prevented, but articulation of decoded speech degrades and noticeable noise is generated in the entire frame
Solution Approach 1:
The patent applies noise addition selectively only to regions where it is beneficial - specifically to unvoiced frame regions and regions without strong periodicity. By limiting noise application to these specific local regions rather than applying it globally, the patent prevents annoying effects where needed while preserving articulation quality in regions where the original excitation signal remains intact.
Solution Approach 2:
The patent applies noise signal processing partially rather than completely - only to specific regions of the excitation signal that require it. This partial application prevents over-processing that would degrade overall articulation, while still providing sufficient noise suppression in the regions where it is most needed.
3Reliability
If parameters of the immediately preceding frame are repeatedly used for voiced frame concealment, then frame loss concealment is achieved, but decoded speech with perceptually strong annoying effects is produced when the preceding frame contains plosive consonants or background noise
Solution Approach 1:
The patent performs preliminary analysis of the excitation signal to identify regions with non-periodic pulse waveforms or high amplitude before applying the repeated excitation signal. By detecting these problematic regions in advance and pre-processing them with noise signals, the patent prevents the generation of harmful artifacts while maintaining reliable frame loss concealment functionality.
Solution Approach 2:
The patent introduces noise signals as an intermediary modification to the repeatedly used excitation signal. This intermediary processing step prevents the direct transmission of harmful characteristics from plosive consonants or background noise in the preceding frame, while still maintaining the overall structure and reliability of frame loss concealment.
Data Source
AI summary
An audio decoding device performs frame loss compensation capable of obtaining a decoded audio which is natural for ears with little noise. The audio decoding device includes a non-cyclic pulse waveform detection unit for detecting a non-cyclic pulse waveform section in a n−1-th frame, which is repeatedly used with a pitch cycle in the n-th frame upon compensation of loss of the n-th frame. The audio coding device also includes a non-cyclic pulse waveform suppression unit for suppressing a non-cyclic pulse waveform by replacing an audio source signal existing in the non-cyclic pulse waveform section in the n−1-th frame by a noise signal. The audio coding device further includes a synthesis filter for using a linear prediction coefficient decoded by an LPC decoding unit to perform synthesis by a synthesis filter by using the audio source signal of the n−1-th frame from the non-cyclic pulse waveform suppression unit as a drive audio source, thereby obtaining the decoded audio signal of the n-th frame.


