Adaptive Formant Sharpening for Speech Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech coding technologies face challenges in maintaining speech quality, especially at low signal-to-noise ratios and in bandwidth extension, due to limitations in formant sharpening techniques and noise modeling in the source-system speech model.
Innovation Solution
The implementation of adaptive formant sharpening techniques, where the formant-sharpening factor is varied based on the long-term signal-to-noise ratio, and the application of specific filters to the fixed codebook vector to emphasize formant regions, thereby improving speech reconstruction quality in both clean and noisy conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If formant sharpening is applied to improve speech reconstruction quality, then speech quality is improved in clean conditions, but artifacts are introduced in noisy conditions and bandwidth extension
Solution Approach 1:
The patent applies dynamic adaptation of the formant sharpening filter based on noise conditions. The filter is selectively applied or adapted in magnitude and frequency based on the measured noise level, allowing the system to optimize speech reconstruction quality while avoiding artifacts in noisy conditions. This is achieved through noise level detection and conditional filter application.
Solution Approach 2:
The patent changes the parameters of the formant sharpening filter based on noise conditions. The filter magnitude and frequency characteristics are adjusted according to the measured noise level, allowing the system to maintain speech quality in clean conditions while reducing artifact introduction in noisy conditions. This parameter adaptation resolves the contradiction between quality improvement and artifact prevention.
2Manufacturing precision
If pitch and formant sharpening are applied to low band excitation for bandwidth extension, then perceptual quality of low-band synthesis is improved, but audible artifacts are more likely to occur in high band synthesis
Solution Approach 1:
The patent applies dynamic control of formant sharpening in bandwidth extension based on noise level detection. The filter is selectively applied or adapted in magnitude and frequency based on the measured noise level, allowing the system to optimize low-band synthesis quality while avoiding artifact propagation to high band synthesis. This conditional adaptation resolves the contradiction between quality improvement and artifact prevention in bandwidth extension.
3Measurement precision
If fixed codebook search is performed to capture aperiodic component, then excitation accuracy is improved, but computational complexity increases
Solution Approach 1:
The patent applies preliminary pitch sharpening to the excitation signal before performing the fixed codebook search. This pre-processing step enhances the periodic components in the excitation, allowing the fixed codebook search to more effectively capture the remaining aperiodic components. This preliminary action improves excitation accuracy while reducing the computational burden during the actual codebook search by providing a better-initialized signal.
Data Source
Figure 1
Figure 2
Figure 3A~3C
AI summary
A method of processing an audio signal includes determining an average signal-to-noise ratio for the audio signal over time. The method includes, based on the determined average signal-to-noise ratio, a formant-sharpening factor is determined. The method also includes applying a filter that is based on the determined formant-sharpening factor to a codebook vector that is based on information from the audio signal.