Speech Error Concealment Using Local Quality Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech transmission methods over networks face issues with errors and delays, leading to degradation in acoustic quality due to missing speech signal frames, which necessitate the use of substitute frames to ensure continuous output.
Innovation Solution
The method employs linear prediction analysis and synthesis filters to produce substitute speech signal frames, using noise signals for frames without voice and fundamental frequency signals for frames with voice, with adaptive scaling to maintain signal energy and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Duration of action of stationary object
If substitute speech signal frames are used to ensure continuous output, then the continuity of speech signal output is improved, but the acoustic quality of the speech signal deteriorates
Solution Approach 1:
The patent applies local quality by differentiating the error concealment approach based on the local characteristics of the lost frame. The system determines whether the lost frame contained voiced or unvoiced speech and applies different synthesis methods accordingly - using fundamental frequency signals for voiced segments and noise signals for unvoiced segments. This localized adaptation maintains acoustic quality by matching the statistical properties of the original speech in each specific region.
Solution Approach 2:
The patent employs parameter changes by dynamically adjusting the spectral and temporal parameters of the substitute speech signal based on the characteristics of surrounding frames. The system modifies parameters such as fundamental frequency, spectral envelope, and noise variance to match the local speech characteristics. This allows the substitute frames to blend seamlessly with the original speech, maintaining perceptual quality while ensuring continuous output.
2Device complexity
If simple substitute frames are used to maintain continuous output, then the complexity of the error concealment system is reduced, but the perceptive quality of the output speech signal deteriorates
Solution Approach 1:
The patent uses parameter changes to adapt the substitute speech signal characteristics dynamically. By analyzing parameters such as autocorrelation values, energy levels, and spectral properties of surrounding frames, the system adjusts the synthesis parameters (fundamental frequency, spectral envelope, noise variance) to match the local speech characteristics. This approach maintains high perceptive quality without requiring complex neural networks or iterative optimization algorithms.
Solution Approach 2:
The system applies local quality analysis by examining the specific characteristics of the frame preceding the lost frame and the frame following it. Based on this local analysis, the system determines whether to use voiced or unvoiced synthesis methods and adjusts the parameters accordingly. This localized approach provides high quality error concealment with relatively simple processing compared to global optimization methods.
3Speed
If delayed speech signal frames are discarded to maintain real-time output, then the real-time performance of speech transmission is improved, but the loss of speech information increases
Solution Approach 1:
The patent applies preliminary action by preparing substitute speech signal frames in advance before the actual lost frames are needed for output. The error concealment process uses the preceding and following frames to generate appropriate substitute content that is ready for immediate insertion. This preliminary preparation ensures that no output delay occurs while minimizing information loss through intelligent synthesis based on local speech characteristics.
Data Source
AI summary
The invention relates to a method for outputting a speech signal. Speech signal frames are received and are used in a predetermined sequence in order to produce a speech signal to be output. If one speech signal frame to be received is not received, then a substitute speech signal frame is used in its place, which is produced as a function of a previously received speech signal frame. According to the invention, in the situation in which the previously received speech signal frame has a voiceless speech signal, the substitute speech signal frame is produced by means of a noise signal.


