Speech Recognition System Ego Noise Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems for robots face challenges in achieving high-accuracy speech recognition in environments with ego noise, which is generated by robot motors and has both diffuse and directional characteristics, leading to degraded audio signal quality and intelligibility.
Innovation Solution
A speech recognition system that includes a sound source separating and speech enhancing section, an ego noise predicting section, and a missing feature mask generating section to generate appropriate missing feature masks based on sound source separation and ego noise prediction, improving speech recognition accuracy by adjusting input data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If single-channel based noise reduction methods are used to reduce ego noise, then noise reduction is achieved, but intelligibility and quality of the audio signal are degraded
Solution Approach 1:
The patent segments the audio signal processing into multiple channels, using a microphone array to capture spatial information. By separating the noise reduction process into individual channel processing and then combining results, the system achieves noise reduction while preserving speech quality and intelligibility that single-channel methods cannot maintain.
Solution Approach 2:
The patent introduces an intermediary approach by using beamforming and spatial filtering as intermediate processing steps between raw microphone signals and final speech output. This intermediary spatial processing allows selective enhancement of speech from specific directions while suppressing ego noise, maintaining signal quality better than direct single-channel noise reduction.
2Productivity
If conventional speech recognition systems are used in environments with ego noise, then speech recognition is performed, but recognition accuracy is degraded due to ego noise contamination
Solution Approach 1:
The patent applies preliminary noise reduction and spatial filtering actions before the speech recognition process. By pre-processing the audio signals to remove ego noise contamination and enhance speech components using microphone array processing, the system ensures that the input to the speech recognition engine is already cleaned, thereby maintaining high recognition accuracy in noisy environments.
3Object-affected harmful factors
If directional noise model or diffuse background noise model is used, then noise suppression is performed, but the models do not entirely hold for ego-motion noise from near-field motors
Solution Approach 1:
The patent changes the parameters and assumptions of the noise model to accommodate near-field motor noise characteristics. Instead of assuming far-field directional or diffuse noise models, the system adapts the spatial filtering and beamforming parameters to handle the unique propagation characteristics of near-field noise, making the noise suppression effective for ego-motion noise specifically.
Data Source
AI summary
A speech recognition system and a speech recognizing method for high-accuracy speech recognition in the environment with ego noise are provided. A speech recognition system according to the present invention includes a sound source separating and speech enhancing section; an ego noise predicting section; and a missing feature mask generating section for generating missing feature masks using outputs of the sound source separating and speech enhancing section and the ego noise predicting section; an acoustic feature extracting section for extracting an acoustic feature of each sound source using an output for said each sound source of the sound source separating and speech enhancing section; and a speech recognizing section for performing speech recognition using outputs of the acoustic feature extracting section and the missing feature masks.


