Speech Recognition System Ego Noise Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems for robots face challenges in achieving high-accuracy speech recognition in environments with ego noise, which is generated by robot motors and has both diffuse and directional characteristics, leading to degraded audio signal quality and intelligibility.

Innovation Solution

A speech recognition system that includes a sound source separating and speech enhancing section, an ego noise predicting section, and a missing feature mask generating section to generate appropriate missing feature masks based on sound source separation and ego noise prediction, improving speech recognition accuracy by adjusting input data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If single-channel based noise reduction methods are used to reduce ego noise, then noise reduction is achieved, but intelligibility and quality of the audio signal are degraded

Engineering Contradiction:
Improveego noiseVSAvoidintelligibility and quality of audio signal
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent segments the audio signal processing into multiple channels, using a microphone array to capture spatial information. By separating the noise reduction process into individual channel processing and then combining results, the system achieves noise reduction while preserving speech quality and intelligibility that single-channel methods cannot maintain.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary approach by using beamforming and spatial filtering as intermediate processing steps between raw microphone signals and final speech output. This intermediary spatial processing allows selective enhancement of speech from specific directions while suppressing ego noise, maintaining signal quality better than direct single-channel noise reduction.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If conventional speech recognition systems are used in environments with ego noise, then speech recognition is performed, but recognition accuracy is degraded due to ego noise contamination

Engineering Contradiction:
Improvespeech recognition capabilityVSAvoidspeech recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary noise reduction and spatial filtering actions before the speech recognition process. By pre-processing the audio signals to remove ego noise contamination and enhance speech components using microphone array processing, the system ensures that the input to the speech recognition engine is already cleaned, thereby maintaining high recognition accuracy in noisy environments.

Inventive Principle:
Principle #10Preliminary action

3Object-affected harmful factors

If directional noise model or diffuse background noise model is used, then noise suppression is performed, but the models do not entirely hold for ego-motion noise from near-field motors

Engineering Contradiction:
Improvenoise suppressionVSAvoidmodel applicability to ego-motion noise
Core Design Contradiction:
Object-affected harmful factorsVSAdaptability or versatility

Solution Approach 1:

The patent changes the parameters and assumptions of the noise model to accommodate near-field motor noise characteristics. Instead of assuming far-field directional or diffuse noise models, the system adapts the spatial filtering and beamforming parameters to handle the unique propagation characteristics of near-field noise, making the noise suppression effective for ego-motion noise specifically.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8538751B2Speech recognition system and speech recognizing method
Publication Date: 2013.09.17 HONDA MOTOR CO LTD
  • US8538751B2 patent drawing
  • US8538751B2 patent drawing
  • US8538751B2 patent drawing

AI summary

A speech recognition system and a speech recognizing method for high-accuracy speech recognition in the environment with ego noise are provided. A speech recognition system according to the present invention includes a sound source separating and speech enhancing section; an ego noise predicting section; and a missing feature mask generating section for generating missing feature masks using outputs of the sound source separating and speech enhancing section and the ego noise predicting section; an acoustic feature extracting section for extracting an acoustic feature of each sound source using an output for said each sound source of the sound source separating and speech enhancing section; and a speech recognizing section for performing speech recognition using outputs of the acoustic feature extracting section and the missing feature masks.