Speech Signal Noise Reduction via Harmonic Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current noise-reduction systems for speech signals, such as those used in automatic speech recognition and hearing aids, are ineffective due to the difficulty in accurately estimating non-stationary noise, leading to degraded performance under varying noise conditions.
Innovation Solution
A system and method that focuses on a subset of harmonics least corrupted by noise, disregards signal harmonics with low signal-to-noise ratios, and filters out amplitude modulations inconsistent with speech, using adaptive filters and processing modules to selectively extract and reconstruct speech information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If noise-reduction systems model and subtract noise from the signal, then noise mitigation is achieved, but the systems fail when noise is non-stationary or different from the model
Solution Approach 1:
Instead of modeling and subtracting noise as conventional systems do, this patent inverts the approach by directly modeling and reconstructing only the speech signal. The system extracts speech features (pitch, harmonics, amplitude modulation) from the noisy input and synthesizes a clean speech signal, completely disregarding the noise component rather than attempting to remove it.
Solution Approach 2:
The patent extracts only the essential speech-carrying components from the noisy signal - specifically the fundamental frequency, harmonic structure, and amplitude modulation characteristics. By taking out only these speech-relevant features and reconstructing the signal from them, the system eliminates noise without requiring noise modeling or subtraction.
2Adaptability or versatility
If training models are used to recognize noise-corrupted speech, then ASR systems can handle noisy input, but the models lack reliability when noise magnitude is too large or dynamic
Solution Approach 1:
The system extracts speech-specific features (pitch contour, harmonic structure, amplitude modulation patterns) that are invariant to noise. By focusing on these fundamental speech characteristics rather than training models to recognize noisy speech patterns, the system achieves reliable speech recognition without requiring extensive training data for various noise conditions.
3Measurement precision
If harmonic structure of speech is used to improve speech recognition, then speech features can be enhanced, but prior detection and tracking methods have been inadequate
Solution Approach 1:
The patent segments the speech signal analysis into distinct components: fundamental frequency detection, harmonic frequency identification, and amplitude modulation extraction. By processing each component separately through dedicated modules, the system achieves accurate harmonic structure analysis while maintaining manageable computational complexity.
Solution Approach 2:
The system dynamically tracks speech parameters (pitch, harmonics, amplitude modulation) over time using adaptive processing. The amplitude modulation detector continuously monitors and tracks the dynamic variations in harmonic amplitudes, allowing the system to adapt to changing speech characteristics in real-time while maintaining computational efficiency.
Data Source
Figure 1A
Figure 1B
Figure 2
AI summary
A system and method for processing a speech signal delivered in a noisy channel or with ambient noise that focuses on a subset of harmonics that are least corrupted by noise, that disregards the signal harmonics with low signal-to-noise ratio(s), and that disregards amplitude modulations inconsistent with speech.