Speech Codec Frame Erasure Concealment with Signal Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital speech encoding techniques face challenges in maintaining speech quality due to frame erasures caused by channel errors in wireless systems and packet losses in voice over packet networks, particularly in wideband speech applications where the adaptive codebook's sensitivity to frame loss affects the pitch predictor, leading to prolonged convergence of the synthesized signal.
Innovation Solution
A method for improving frame erasure concealment and decoder recovery by determining and transmitting concealment/recovery parameters, allowing the decoder to conduct erasure frame concealment and recovery, thereby reducing the impact of frame losses and accelerating the convergence of the synthesized signal to the intended one.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If frame erasure concealment is performed using traditional methods, then the decoder can continue operating after frame loss, but the synthesized signal diverges from the intended signal and convergence is prolonged
Solution Approach 1:
The encoder performs preliminary classification of the speech signal type (voiced, unvoiced, transient) before transmission and sends this classification information to the decoder. When frame erasure occurs, the decoder uses this pre-transmitted classification to immediately select the appropriate concealment strategy, avoiding signal divergence and accelerating convergence without requiring complex real-time analysis at the decoder.
2Measurement precision
If the adaptive codebook content is updated continuously, then the pitch predictor maintains accuracy during normal operation, but frame loss causes desynchronization between encoder and decoder adaptive codebooks
Solution Approach 1:
The system implements feedback by transmitting speech signal type classification from encoder to decoder. This feedback allows the decoder to mirror the encoder's understanding of the signal characteristics, ensuring both sides use the same classification for adaptive codebook management during concealment, thereby maintaining synchronization without sacrificing pitch prediction accuracy.
3Device complexity
If frame erasure concealment uses simple repetition, then the implementation is computationally simple, but the speech quality degrades significantly
Solution Approach 1:
The system applies local quality by tailoring the concealment strategy to the specific type of speech signal segment that was lost. Voiced segments use periodic extension, unvoiced segments use noise generation, and transients use different handling. This localized approach matching concealment method to signal type maintains high speech quality without requiring complex universal algorithms.
4Ease of operation
If the codec processes all frame types uniformly, then the processing is straightforward and consistent, but the convergence after frame loss is prolonged
Solution Approach 1:
The speech signal is segmented into different types (voiced, unvoiced, transient) based on classification. Each segment type has a dedicated concealment strategy optimized for its characteristics. This segmentation allows the system to move from uniform processing to type-specific processing, significantly reducing convergence time after frame loss while maintaining operational clarity through structured classification.
Data Source
AI summary
The present invention relates to a method and device for improving concealment of frame erasure caused by frames of an encoded sound signal erased during transmission from an encoder (106) to a decoder (110), and for accelerating recovery of the decoder after non erased frames of the encoded sound signal have been received. For that purpose, concealment/recovery parameters are determined in the encoder or decoder. When determined in the encoder (106), the concealment/recovery parameters are transmitted to the decoder (110). In the decoder, erasure frame concealment and decoder recovery is conducted in response to the concealment/recovery parameters. The concealment/recovery parameters may be selected from the group consisting of: a signal classification parameter, an energy information parameter and a phase information parameter. The determination of the concealment/recovery parameters comprises classifying the successive frames of the encoded sound signal as unvoiced, unvoiced transition, voiced transition, voiced, or onset, and this classification is determined on the basis of at least a part of the following parameters: a normalized correlation parameter, a spectral tilt parameter, a signal-to-noise ratio parameter, a pitch stability parameter, a relative frame energy parameter, and a zero crossing parameter.


