Speech Input Probability Filtering for Ambient Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech processing engines struggle to accurately identify and distinguish between input speech intended for processing and non-input speech in everyday environments, leading to inefficiencies and inaccuracies due to ambient noise and unintended audio signals.
Innovation Solution
The system employs a method to determine the probability that a portion of an audio signal represents input speech directed at a speech processing engine, using a combination of audio characteristics and sensor data from wearable devices to filter out non-input speech and present only relevant input to the engine.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the speech processing engine processes all audio input in everyday environments, then it can capture all potential speech signals, but it cannot distinguish between input speech and non-input speech, leading to processing inefficiencies and inaccuracies
Solution Approach 1:
The audio signal is segmented into multiple portions, and each portion is independently evaluated to determine whether it represents input speech. This segmentation allows the system to process only relevant segments, improving both accuracy and efficiency by avoiding unnecessary processing of non-input speech.
Solution Approach 2:
The system performs preliminary analysis on audio portions before full speech processing by evaluating characteristics such as voice activity detection, directionality, and probability scores. This preliminary action filters out non-input speech early in the process, preventing wasted computational resources on irrelevant audio segments.
2Productivity
If the system uses wake-up words to activate speech processing, then it can reduce unnecessary processing, but it introduces false positives where the wake-up word is spoken without intention of activating the engine
Solution Approach 1:
The system performs preliminary analysis on audio portions before full speech processing by evaluating characteristics such as voice activity detection, directionality, and probability scores. This preliminary action filters out non-input speech early in the process, preventing wasted computational resources on irrelevant audio segments.
Solution Approach 2:
The system continuously monitors audio input and adjusts processing based on real-time evaluation of speech characteristics. By providing feedback loops that assess voice activity, directionality, and probability scores, the system can dynamically respond to actual user intent rather than relying on static wake-up word triggers.
3Measurement precision
If multiple microphones are used to capture audio from different directions, then the system can better identify speech direction, but the complexity of the system increases
Solution Approach 1:
The system combines data from multiple microphones and sensor inputs to evaluate speech direction and characteristics. By merging information from multiple sources into a unified probability assessment, the system achieves improved direction detection without proportionally increasing overall system complexity.
Solution Approach 2:
The audio input module serves multiple functions: it captures audio signals, determines speech direction, evaluates voice activity, and generates probability scores for input speech identification. This multi-functionality reduces the need for separate dedicated components, managing system complexity while maintaining precision.
Data Source
AI summary
A method of presenting a signal to a speech processing engine is disclosed. According to an example of the method, an audio signal is received via a microphone. A portion of the audio signal is identified, and a probability is determined that the portion comprises speech directed by a user of the speech processing engine as input to the speech processing engine. In accordance with a determination that the probability exceeds a threshold, the portion of the audio signal is presented as input to the speech processing engine. In accordance with a determination that the probability does not exceed the threshold, the portion of the audio signal is not presented as input to the speech processing engine.


