Speech Detection System for Impulse Noise Removal

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems for devices with small form factors, such as cellular phones and PDAs, face limitations in throughput and accuracy due to push-to-speak configurations, which require users to initiate speech after a system indicator, introducing impulse noise and disrupting the natural flow of speech input.

Innovation Solution

A speech detection system that continuously listens for and identifies desired speech segments within an audio stream using pattern matching and signal processing techniques, incorporating modules for feature generation, time alignment, and determining desired speech segments, allowing for seamless integration with keypad inputs in multimodal interfaces.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If push-to-speak configuration is used to initiate speech recognition, then speech input can be controlled, but speech recognition accuracy decreases due to impulse noise from button pressing

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidoperation simplicity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent extracts and removes the harmful impulse noise component from the audio signal that is generated during button pressing. By separating the desired speech signal from the unwanted button-press noise, the system maintains speech recognition accuracy while preserving the ease of push-to-speak operation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary noise cancellation mechanism that mediates between the button press action and the speech recognition process. This intermediary component filters out the impulse noise generated by button pressing, allowing the speech recognition system to accurately process speech inputs without being disrupted by operational noise.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If push-to-speak configuration is used, then speech input can be initiated, but overall throughput decreases due to behavioral change requirements

Engineering Contradiction:
Improvetext input throughputVSAvoiduser behavior adaptation
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent inverts the traditional push-to-speak model by enabling continuous speech recognition without requiring users to adapt their natural speech behavior. Instead of requiring users to speak after a beep or upon button press, the system continuously processes speech inputs, allowing users to speak naturally at any time, thereby maintaining high throughput without behavioral adaptation requirements.

Inventive Principle:
Principle #13The other way round (Inversion)

3Ease of operation

If continuous speech listening is implemented, then natural speech input flow is enabled, but impulse noise from button presses affects recognition accuracy

Engineering Contradiction:
Improvenatural speech flowVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent converts the harmful impulse noise generated during continuous listening into a detectable signal characteristic. By analyzing the temporal and spectral properties of button-press noise, the system can identify and exclude these segments from speech recognition processing, thereby maintaining both continuous listening capability and high recognition accuracy.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS8645131B2Detecting segments of speech from an audio stream
Publication Date: 2014.02.04 RAO ASHWIN P
  • US8645131B2 patent drawing
  • US8645131B2 patent drawing
  • US8645131B2 patent drawing

AI summary

The disclosure describes a speech detection system for detecting one or more desired speech segments in an audio stream. The speech detection system includes an audio stream input and a speech detection technique. The speech detection technique may be performed in various ways, such as using pattern matching and/or signal processing. The pattern matching implementation may extract features representing types of sounds as in phrases, words, syllables, phonemes and so on. The signal processing implementation may extract spectrally-localized frequency-based features, amplitude-based features, and combinations of the frequency-based and amplitude-based features. Metrics may be obtained and used to determine a desired word in the audio stream. In addition, a keypad stream having keypad entries may be used in determining the desired word.