Speech Decoding Device Frame Loss Compensation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech codecs, particularly those using CELP with an adaptive codebook as a main layer, face challenges in decoding current frames when a preceding frame is lost, leading to quality degradation due to low correlation with the preceding frame and distortion in the decoded speech signal.
Innovation Solution
A speech decoding apparatus that receives frame loss information and uses encoded pitch pulse information from the lost frame to generate a compensated excitation signal by learning a pitch pulse waveform, allowing for effective concealment even if a preceding frame is lost, thereby improving speech encoding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a predictive type of encoding method using past encoded information is used as a main layer, then encoding efficiency is improved, but decoding cannot be performed correctly when a preceding frame is lost
Solution Approach 1:
The patent segments the excitation signal into two independent parts: a decoded excitation signal generated from current frame encoded information, and a pitch pulse signal generated from lost frame pitch pulse information. This segmentation allows the current frame to be decoded independently without relying on preceding frame data, resolving the contradiction between encoding efficiency and decoding reliability.
Solution Approach 2:
The patent introduces a pitch pulse learning section as an intermediary that generates pitch pulse waveforms based on stored pitch pulse information from lost frames. This intermediary component enables the system to recover pitch pulse characteristics without needing the actual preceding frame data, maintaining both encoding efficiency and decoding correctness.
2Ease of operation
If conventional concealment processing is used in speech onset portions, then processing simplicity is maintained, but quality degradation occurs due to low correlativity with preceding frame signal
Solution Approach 1:
The patent performs preliminary action by storing pitch pulse information (position, amplitude, waveform) from each frame before potential loss occurs. When a frame is lost, this pre-stored information is immediately available for accurate reconstruction, eliminating the need for simple but ineffective conventional concealment processing and significantly improving speech quality in onset portions.
3Manufacturing precision
If encoded information for concealment processing is transmitted together with current frame information, then lost frame concealment quality is improved, but information transmission load increases
Solution Approach 1:
The patent extracts only the essential pitch pulse information (position, amplitude, and waveform characteristics) from the lost frame and stores it for future concealment processing. This selective extraction provides high-quality concealment capability while minimizing the amount of additional information that needs to be transmitted, thus reducing the information transmission load compared to transmitting full concealment processing information.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided is an audio decoding device capable of suppressing an information amount for a lost frame compensation process and encoding efficiency. In this device, a decoded sound source generation unit (203) generates a lost frame decoded sound source signal; a pitch pulse information decoding unit (204) decodes the pitch pulse position information and the pitch pulse amplitude information; a pitch pulse waveform learning unit (205) learns a pitch pulse learning waveform in the past frame in advance from the lost frame; a convolution unit (206) amplitude-adjusts the pitch pulse learning waveform according to the pitch pulse amplitude information, and convolutes the pitch pulse waveform into a time axis which has been amplitude-adjusted according to the pitch pulse position information; a sound source signal correction unit (207) adds or replaces the pitch pulse waveform convoluted into the time axis to the lost frame decoded sound source signal.