Adaptive Codebook Concealment With Fractional Pitch Resynchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech codecs like G.718 and G.729.1 suffer from inaccurate pitch lag reconstruction during frame loss, leading to significant differences between reconstructed and actual pitch lags, and fail to account for the reliability of pitch information and pulse resynchronization in concealed frames.
Innovation Solution
The proposed solution weights pitch lags based on their reliability, using past adaptive codebook gains and time of reception, and improves pulse resynchronization by considering non-integer pitch cycles and fractional pitch lag, allowing for more accurate pitch prediction and frame reconstruction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If simple repetition based concealment is used, then device complexity is reduced, but pitch reconstruction accuracy deteriorates
Solution Approach 1:
The patent applies preliminary action by pre-calculating and storing pitch lag values in a buffer during normal operation (before frame loss occurs). When frame loss happens, these pre-stored pitch lag values are immediately used for reconstruction without requiring complex real-time calculations, thus resolving the contradiction between simplicity and accuracy.
Solution Approach 2:
The patent changes the parameter approach by using multiple pitch lag values (N values) instead of a single repeated value. It introduces a buffer storing recent pitch lag values and selects appropriate values based on frame loss duration and voice activity detection, allowing accurate pitch reconstruction while maintaining manageable complexity through parameter variation rather than algorithmic complexity.
2Measurement precision
If pitch extrapolation is applied, then pitch reconstruction accuracy is improved, but reliability of pitch information deteriorates due to accumulated errors
Solution Approach 1:
The patent implements feedback by continuously monitoring voice activity detection (VAD) results and frame loss conditions to dynamically adjust pitch lag selection. When voice activity is detected, it uses recent pitch lag values from the buffer; when silence is detected, it switches to repetition mode. This feedback mechanism ensures reliability by adapting to actual speech conditions while maintaining accuracy through appropriate pitch lag selection.
Solution Approach 2:
The patent applies dynamics by making the pitch reconstruction approach adaptive rather than static. It dynamically switches between different pitch lag values from the buffer based on VAD results and frame loss characteristics. This dynamic adaptation allows the system to maintain reliability across varying speech conditions while achieving accurate pitch reconstruction through context-aware parameter selection.
3Manufacturing precision
If pulse resynchronization is performed, then speech quality is improved, but device complexity increases
Solution Approach 1:
The patent extracts and applies fractional pitch lag values separately from the integer pitch lag values. By taking out the fractional part and applying it as an offset to the resynchronized pulses, it achieves improved speech quality through precise pitch alignment without requiring complete redesign of the resynchronization algorithm, thus managing complexity while enhancing quality.
4Measurement precision
If fractional pitch lag is considered, then pitch reconstruction accuracy is improved, but computational complexity increases
Solution Approach 1:
The patent uses a simple buffering approach for storing fractional pitch lag values rather than implementing complex fractional pitch analysis algorithms. The fractional values are captured during normal operation and stored in the existing pitch lag buffer, then applied as simple offsets during reconstruction. This disposable-like approach (storing and reusing values) achieves high precision without significant computational complexity.
Data Source
AI summary
An apparatus for reconstructing a frame including a speech signal as a reconstructed frame is provided, the apparatus including a determination unit and a frame reconstructor being configured to reconstruct the reconstructed frame, such that the reconstructed frame completely or partially includes the first reconstructed pitch cycle, such that the reconstructed frame completely or partially includes a second reconstructed pitch cycle, and such that the number of samples of the first reconstructed pitch cycle differs from a number of samples of the second reconstructed pitch cycle.


