Multimodal Utterance Detection Using Timing Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in accurately detecting utterances, especially in noisy environments, and existing solutions are not robust enough to handle various application scenarios, including those without manual start/stop buttons or always-listening modes.

Innovation Solution

A multimodal utterance detection method that uses knowledge from multiple modes, such as location, touch, and eye indications, to filter and identify desired speech segments from audio streams, incorporating timing markers and application-specific knowledge to distinguish speech from noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional single-mode utterance detection methods are used, then the system is simple to implement, but the detection accuracy deteriorates in noisy environments and diverse application scenarios

Engineering Contradiction:
Improveutterance detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple detection modes (audio, visual, tactile, location) into a unified utterance detection system. The system integrates signals from different sensors and processes them together to make detection decisions, thereby improving accuracy in noisy environments while managing complexity through a structured fusion approach.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system is designed to handle multiple application scenarios universally, including scenarios with manual start/stop buttons, always-listening modes, and scenarios without any manual controls. The multimodal approach allows the same system to adapt to different interaction patterns and environmental conditions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If manual start/stop buttons are required for speech input, then the system has simple control logic, but the user interface becomes rigid and cannot handle scenarios where users forget to synchronize speaking with buttons

Engineering Contradiction:
Improveuser interface flexibilityVSAvoidutterance detection reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system uses automatic utterance detection based on multimodal signal analysis to identify speech segments without requiring manual start/stop buttons. The system serves itself by automatically detecting when speech occurs through analysis of audio, visual, and other modalities, eliminating the need for user synchronization with manual controls.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system continuously monitors multiple signal modalities in advance, preparing detection data before speech actually occurs. This allows the system to reliably detect utterances even when users speak before pressing start buttons or continue speaking after pressing stop buttons, as the preliminary multimodal data is already available for analysis.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If always-listening mode is implemented without manual buttons, then the user interface is relaxed and convenient, but the system consumes more energy and processes more unnecessary audio data

Engineering Contradiction:
Improveuser interface convenienceVSAvoidenergy consumption
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The system performs partial processing by analyzing only relevant modalities or using lightweight detection algorithms for initial screening. Instead of continuously processing all audio data at full complexity, the system applies detection methods selectively based on the presence of triggering signals in multiple modalities, reducing unnecessary computation and energy consumption.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9922640B2System and method for multimodal utterance detection
Publication Date: 2018.03.20 RAO ASHWIN P
  • US9922640B2 patent drawing
  • US9922640B2 patent drawing
  • US9922640B2 patent drawing

AI summary

The disclosure describe a system and method for detecting one or more segments of desired speech utterances from an audio stream using timings of events from other modes that are correlated to the timings of the desired segments of speech. The redundant information from other modes results in a highly accurate and robust utterance detection.