Extended High-Frequency Speech Extraction from Masking Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing hearing aid technologies struggle to selectively amplify target speech while minimizing noise, especially when non-target speech is present, due to limited computational resources and the disregard of high-frequency audio signal contents above 6-8 kHz.
Innovation Solution
Implement computationally efficient audio filters that utilize elevated frequency content above a threshold frequency to detect target speech, generate filters, and apply them to enhance speech extraction, using methods like non-negative matrix factorization to separate noise from target speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional speech processing methods are used focusing on low frequencies below 6-8 kHz, then computational resources are saved, but speech perception in noisy environments deteriorates
Solution Approach 1:
The audio signal is segmented into high-frequency components above 6-8 kHz and processed separately using spectrogram analysis and non-negative matrix factorization. This allows the system to focus computational resources on the most informative frequency band for speech enhancement while maintaining efficiency.
Solution Approach 2:
The system performs preliminary analysis of the audio signal by computing spectrograms and identifying high-frequency content before applying the audio filter. This preliminary action enables the system to adapt the filter parameters based on the actual content of the audio signal, improving speech perception without excessive computational cost.
2Measurement precision
If high-frequency audio content above 6-8 kHz is included in speech processing, then speech extraction accuracy is improved, but device complexity increases
Solution Approach 1:
The system changes the parameter of frequency range inclusion by incorporating high-frequency components above 6-8 kHz into the speech processing. This parameter change enables better speech extraction accuracy by capturing additional spectral information that is crucial for speech perception in noisy environments.
Solution Approach 2:
The spectrogram serves as an intermediary representation that transforms the high-frequency audio content into a format that can be processed more efficiently. By converting the audio signal into a spectrogram and then applying non-negative matrix factorization, the system reduces the computational complexity of handling high-frequency data while maintaining extraction accuracy.
3Measurement precision
If non-negative matrix factorization is applied to separate noise from speech, then speech recognition is improved, but processing time increases
Solution Approach 1:
The system applies non-negative matrix factorization only to the high-frequency portion of the audio signal above 6-8 kHz, rather than processing the entire audio spectrum. This partial action reduces the amount of computation required while still capturing the most informative noise-speech distinctions, thereby reducing processing time while maintaining recognition accuracy.
Data Source
AI summary
Improved systems and methods are provided herein for extracting target speech from audio signals that can contain masking speech or other unwanted noise content. These systems and methods include detection of target speech in an input signal by detecting elevated frequency content in the signal above a threshold frequency. Portions of the signal determined to contain such elevated high frequency content are then used to generate audio filters to extract target speech from subsequently-obtained audio signals. This can include performing non-negative matrix factorization to determine a set of basis vectors to represent noise content in the spectral domain and then using the set of basis vectors to decompose subsequently-obtained audio signals into noise signals that can then be removed from the audio signals.


