Voice Trigger Detection Using Audio-Visual Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice command-and-control systems rely solely on audio information, leading to high rates of false accepts and false rejects due to background noise and multiple speakers, which reduces their accuracy.

Innovation Solution

The integration of audio-based and vision-based cues, where an electronic device identifies and synchronizes audio and video signals to determine the utterance of a predefined trigger phrase, using a probabilistic model to assess the likelihood of the phrase being spoken.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice command-and-control systems rely solely on audio information to detect trigger phrases, then the system can operate with simpler architecture and lower cost, but the accuracy deteriorates due to background noise and multiple speakers causing high rates of false accepts and false rejects

Engineering Contradiction:
Improveaccuracy of trigger phrase detectionVSAvoidsystem architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines audio-based cues and vision-based cues into a unified trigger phrase detection system. The audio processing component analyzes audio signals for potential trigger phrases, while the video processing component analyzes video frames for corresponding visual cues. These two independent detection pathways are merged through a probabilistic model that computes the likelihood of a trigger phrase being spoken based on both audio and visual evidence, thereby improving reliability without requiring complete system redesign

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a probabilistic model as an intermediary component that bridges audio-based detection and vision-based detection. This mediator processes cues from both modalities, synchronizes them temporally and spatially, and computes a combined probability score. The probabilistic model acts as a decision-making intermediary that determines whether the trigger phrase was actually spoken, resolving conflicts between audio-only and visual-only detections and improving overall accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the system uses only audio processing, then the device complexity and power consumption remain low, but false accepts and false rejects increase due to background noise and multiple individuals speaking simultaneously

Engineering Contradiction:
Improvereduction of false accepts and false rejectsVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system employs periodic action by processing video frames at specific intervals rather than continuously. The video processing component analyzes a sequence of video frames to identify visual cues corresponding to trigger phrase utterance, but does so in a periodic manner where full video analysis is triggered only when audio-based cues suggest a potential trigger phrase. This periodic processing reduces power consumption while maintaining reliability by activating the more energy-intensive vision processing only when necessary

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system applies preliminary action by first processing the audio signal to identify audio-based cues indicating a possible trigger phrase utterance before activating full video processing. The audio processing acts as a preliminary filter that screens potential trigger events, and only when audio cues exceed a certain threshold does the system proceed to perform the more computationally intensive and power-consuming video analysis. This preliminary audio screening reduces overall power consumption while maintaining high reliability

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9916832B2Using combined audio and vision-based cues for voice command-and-control
Publication Date: 2018.03.13 SENSORY INC
  • US9916832B2 patent drawing
  • US9916832B2 patent drawing
  • US9916832B2 patent drawing

AI summary

Techniques for leveraging a combination of audio-based and vision-based cues for voice command-and-control are provided. In one embodiment, an electronic device can identify one or more audio-based cues in a received audio signal that pertain to a possible utterance of a predefined trigger phrase, and identify one or more vision-based cues in a received video signal that pertain to a possible utterance of the predefined trigger phrase. The electronic device can further determine a degree of synchronization or correspondence between the one or more audio-based cues and the one or more vision-based cues. The electronic device can then conclude, based on the one or more audio-based cues, the one or more vision-based cues, and the degree of synchronization or correspondence, whether the predefined trigger phrase was actually spoken.