Multimodal Utterance Detection Using Timing Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face challenges in accurately detecting utterances, especially in noisy environments, and existing solutions are not robust enough to handle various application scenarios, including those without manual start/stop buttons or always-listening modes.
Innovation Solution
A multimodal utterance detection method that uses knowledge from multiple modes, such as location, touch, and eye indications, to filter and identify desired speech segments from audio streams, incorporating timing markers and application-specific knowledge to distinguish speech from noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional single-mode utterance detection methods are used, then the system is simple to implement, but the detection accuracy deteriorates in noisy environments and diverse application scenarios
Solution Approach 1:
The patent combines multiple detection modes (audio, visual, tactile, location) into a unified utterance detection system. The system integrates signals from different sensors and processes them together to make detection decisions, thereby improving accuracy in noisy environments while managing complexity through a structured fusion approach.
Solution Approach 2:
The system is designed to handle multiple application scenarios universally, including scenarios with manual start/stop buttons, always-listening modes, and scenarios without any manual controls. The multimodal approach allows the same system to adapt to different interaction patterns and environmental conditions.
2Ease of operation
If manual start/stop buttons are required for speech input, then the system has simple control logic, but the user interface becomes rigid and cannot handle scenarios where users forget to synchronize speaking with buttons
Solution Approach 1:
The system uses automatic utterance detection based on multimodal signal analysis to identify speech segments without requiring manual start/stop buttons. The system serves itself by automatically detecting when speech occurs through analysis of audio, visual, and other modalities, eliminating the need for user synchronization with manual controls.
Solution Approach 2:
The system continuously monitors multiple signal modalities in advance, preparing detection data before speech actually occurs. This allows the system to reliably detect utterances even when users speak before pressing start buttons or continue speaking after pressing stop buttons, as the preliminary multimodal data is already available for analysis.
3Ease of operation
If always-listening mode is implemented without manual buttons, then the user interface is relaxed and convenient, but the system consumes more energy and processes more unnecessary audio data
Solution Approach 1:
The system performs partial processing by analyzing only relevant modalities or using lightweight detection algorithms for initial screening. Instead of continuously processing all audio data at full complexity, the system applies detection methods selectively based on the presence of triggering signals in multiple modalities, reducing unnecessary computation and energy consumption.
Data Source
AI summary
The disclosure describe a system and method for detecting one or more segments of desired speech utterances from an audio stream using timings of events from other modes that are correlated to the timings of the desired segments of speech. The redundant information from other modes results in a highly accurate and robust utterance detection.


