Speech Detection System for Impulse Noise Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems for devices with small form factors, such as cellular phones and PDAs, face limitations in throughput and accuracy due to push-to-speak configurations, which require users to initiate speech after a system indicator, introducing impulse noise and disrupting the natural flow of speech input.
Innovation Solution
A speech detection system that continuously listens for and identifies desired speech segments within an audio stream using pattern matching and signal processing techniques, incorporating modules for feature generation, time alignment, and determining desired speech segments, allowing for seamless integration with keypad inputs in multimodal interfaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If push-to-speak configuration is used to initiate speech recognition, then speech input can be controlled, but speech recognition accuracy decreases due to impulse noise from button pressing
Solution Approach 1:
The patent extracts and removes the harmful impulse noise component from the audio signal that is generated during button pressing. By separating the desired speech signal from the unwanted button-press noise, the system maintains speech recognition accuracy while preserving the ease of push-to-speak operation.
Solution Approach 2:
The patent introduces an intermediary noise cancellation mechanism that mediates between the button press action and the speech recognition process. This intermediary component filters out the impulse noise generated by button pressing, allowing the speech recognition system to accurately process speech inputs without being disrupted by operational noise.
2Productivity
If push-to-speak configuration is used, then speech input can be initiated, but overall throughput decreases due to behavioral change requirements
Solution Approach 1:
The patent inverts the traditional push-to-speak model by enabling continuous speech recognition without requiring users to adapt their natural speech behavior. Instead of requiring users to speak after a beep or upon button press, the system continuously processes speech inputs, allowing users to speak naturally at any time, thereby maintaining high throughput without behavioral adaptation requirements.
3Ease of operation
If continuous speech listening is implemented, then natural speech input flow is enabled, but impulse noise from button presses affects recognition accuracy
Solution Approach 1:
The patent converts the harmful impulse noise generated during continuous listening into a detectable signal characteristic. By analyzing the temporal and spectral properties of button-press noise, the system can identify and exclude these segments from speech recognition processing, thereby maintaining both continuous listening capability and high recognition accuracy.
Data Source
AI summary
The disclosure describes a speech detection system for detecting one or more desired speech segments in an audio stream. The speech detection system includes an audio stream input and a speech detection technique. The speech detection technique may be performed in various ways, such as using pattern matching and/or signal processing. The pattern matching implementation may extract features representing types of sounds as in phrases, words, syllables, phonemes and so on. The signal processing implementation may extract spectrally-localized frequency-based features, amplitude-based features, and combinations of the frequency-based and amplitude-based features. Metrics may be obtained and used to determine a desired word in the audio stream. In addition, a keypad stream having keypad entries may be used in determining the desired word.


