Voice Trigger Detection Using Audio-Visual Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice command-and-control systems rely solely on audio information, leading to high rates of false accepts and false rejects due to background noise and multiple speakers, which reduces their accuracy.
Innovation Solution
The integration of audio-based and vision-based cues, where an electronic device identifies and synchronizes audio and video signals to determine the utterance of a predefined trigger phrase, using a probabilistic model to assess the likelihood of the phrase being spoken.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice command-and-control systems rely solely on audio information to detect trigger phrases, then the system can operate with simpler architecture and lower cost, but the accuracy deteriorates due to background noise and multiple speakers causing high rates of false accepts and false rejects
Solution Approach 1:
The patent combines audio-based cues and vision-based cues into a unified trigger phrase detection system. The audio processing component analyzes audio signals for potential trigger phrases, while the video processing component analyzes video frames for corresponding visual cues. These two independent detection pathways are merged through a probabilistic model that computes the likelihood of a trigger phrase being spoken based on both audio and visual evidence, thereby improving reliability without requiring complete system redesign
Solution Approach 2:
The patent introduces a probabilistic model as an intermediary component that bridges audio-based detection and vision-based detection. This mediator processes cues from both modalities, synchronizes them temporally and spatially, and computes a combined probability score. The probabilistic model acts as a decision-making intermediary that determines whether the trigger phrase was actually spoken, resolving conflicts between audio-only and visual-only detections and improving overall accuracy
2Reliability
If the system uses only audio processing, then the device complexity and power consumption remain low, but false accepts and false rejects increase due to background noise and multiple individuals speaking simultaneously
Solution Approach 1:
The system employs periodic action by processing video frames at specific intervals rather than continuously. The video processing component analyzes a sequence of video frames to identify visual cues corresponding to trigger phrase utterance, but does so in a periodic manner where full video analysis is triggered only when audio-based cues suggest a potential trigger phrase. This periodic processing reduces power consumption while maintaining reliability by activating the more energy-intensive vision processing only when necessary
Solution Approach 2:
The system applies preliminary action by first processing the audio signal to identify audio-based cues indicating a possible trigger phrase utterance before activating full video processing. The audio processing acts as a preliminary filter that screens potential trigger events, and only when audio cues exceed a certain threshold does the system proceed to perform the more computationally intensive and power-consuming video analysis. This preliminary audio screening reduces overall power consumption while maintaining high reliability
Data Source
AI summary
Techniques for leveraging a combination of audio-based and vision-based cues for voice command-and-control are provided. In one embodiment, an electronic device can identify one or more audio-based cues in a received audio signal that pertain to a possible utterance of a predefined trigger phrase, and identify one or more vision-based cues in a received video signal that pertain to a possible utterance of the predefined trigger phrase. The electronic device can further determine a degree of synchronization or correspondence between the one or more audio-based cues and the one or more vision-based cues. The electronic device can then conclude, based on the one or more audio-based cues, the one or more vision-based cues, and the degree of synchronization or correspondence, whether the predefined trigger phrase was actually spoken.


