Speech Input Probability Filtering for Ambient Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech processing engines struggle to accurately identify and distinguish between input speech intended for processing and non-input speech in everyday environments, leading to inefficiencies and inaccuracies due to ambient noise and unintended audio signals.

Innovation Solution

The system employs a method to determine the probability that a portion of an audio signal represents input speech directed at a speech processing engine, using a combination of audio characteristics and sensor data from wearable devices to filter out non-input speech and present only relevant input to the engine.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the speech processing engine processes all audio input in everyday environments, then it can capture all potential speech signals, but it cannot distinguish between input speech and non-input speech, leading to processing inefficiencies and inaccuracies

Engineering Contradiction:
Improvespeech identification accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The audio signal is segmented into multiple portions, and each portion is independently evaluated to determine whether it represents input speech. This segmentation allows the system to process only relevant segments, improving both accuracy and efficiency by avoiding unnecessary processing of non-input speech.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis on audio portions before full speech processing by evaluating characteristics such as voice activity detection, directionality, and probability scores. This preliminary action filters out non-input speech early in the process, preventing wasted computational resources on irrelevant audio segments.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If the system uses wake-up words to activate speech processing, then it can reduce unnecessary processing, but it introduces false positives where the wake-up word is spoken without intention of activating the engine

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidactivation accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary analysis on audio portions before full speech processing by evaluating characteristics such as voice activity detection, directionality, and probability scores. This preliminary action filters out non-input speech early in the process, preventing wasted computational resources on irrelevant audio segments.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously monitors audio input and adjusts processing based on real-time evaluation of speech characteristics. By providing feedback loops that assess voice activity, directionality, and probability scores, the system can dynamically respond to actual user intent rather than relying on static wake-up word triggers.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If multiple microphones are used to capture audio from different directions, then the system can better identify speech direction, but the complexity of the system increases

Engineering Contradiction:
Improvespeech direction detectionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system combines data from multiple microphones and sensor inputs to evaluate speech direction and characteristics. By merging information from multiple sources into a unified probability assessment, the system achieves improved direction detection without proportionally increasing overall system complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The audio input module serves multiple functions: it captures audio signals, determines speech direction, evaluates voice activity, and generates probability scores for input speech identification. This multi-functionality reduces the need for separate dedicated components, managing system complexity while maintaining precision.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250157471A1Determining input for speech processing engine
Publication Date: 2025.05.15 MAGIC LEAP INC
  • US20250157471A1 patent drawing
  • US20250157471A1 patent drawing
  • US20250157471A1 patent drawing

AI summary

A method of presenting a signal to a speech processing engine is disclosed. According to an example of the method, an audio signal is received via a microphone. A portion of the audio signal is identified, and a probability is determined that the portion comprises speech directed by a user of the speech processing engine as input to the speech processing engine. In accordance with a determination that the probability exceeds a threshold, the portion of the audio signal is presented as input to the speech processing engine. In accordance with a determination that the probability does not exceed the threshold, the portion of the audio signal is not presented as input to the speech processing engine.