Digital Assistant Gaze and Lip Detection for Voice Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional digital assistants face challenges in accurately detecting user attention, distinguishing between users' phrases, and maintaining context in voice-based interactions, leading to complex and adaptive user experiences.
Innovation Solution
A method for voice-based interactive communication that includes attention detection, speaker detection, speech sound detection with lip movement analysis, and response provision, using sensors like sound, image, and gaze detection to set the digital assistant into a listening mode, track the current speaker, and parse speech for context-aware feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If conventional digital assistants use predetermined catch-phrases to detect user attention, then the detection mechanism is simple, but the accuracy of detecting user attention and phrase start point deteriorates
Solution Approach 1:
The patent combines multiple detection modalities (audio signal analysis, lip movement detection via image sensor, and gaze detection) into a unified attention detection system. This merging of detection methods allows the digital assistant to accurately determine when a user is paying attention and when speech begins, resolving the contradiction between simple detection mechanisms and accurate attention detection.
Solution Approach 2:
The patent introduces lip movement detection as an intermediary signal between user intent and system response. By detecting lip movements through image sensors and correlating them with audio signals, the system gains an additional layer of information that improves the accuracy of detecting phrase start points and user attention without requiring complex predetermined catch-phrases.
2Measurement precision
If multiple sensors and detection methods are used to improve user attention and speech detection accuracy, then detection accuracy improves, but device complexity increases
Solution Approach 1:
The patent makes the image sensor serve multiple functions: it detects lip movements for speech synchronization, detects gaze direction for attention determination, and can potentially identify speaker identity. This multi-functionality allows the system to achieve high speech detection accuracy using a single versatile component rather than multiple specialized sensors, thereby reducing overall system complexity.
Solution Approach 2:
The patent divides the detection process into distinct functional modules: audio signal processing, image signal processing for lip movement detection, gaze detection for attention analysis, and context analysis for speaker identification. Each module handles a specific aspect of the detection task, making the overall complex system manageable and maintainable while achieving high accuracy through the coordinated operation of specialized sub-systems.
3Measurement precision
If the digital assistant continuously monitors speech and lip movements to accurately detect phrase endpoints, then endpoint detection accuracy improves, but processing time and energy consumption increase
Solution Approach 1:
The patent employs periodic sampling of lip movement data during speech detection rather than continuous monitoring. By analyzing lip movement at regular intervals and correlating these periodic samples with audio signal characteristics, the system achieves accurate endpoint detection while reducing the computational burden and processing time compared to continuous analysis.
Solution Approach 2:
The patent performs preliminary analysis of audio signals to identify potential speech segments before conducting detailed lip movement correlation. By pre-processing the audio signal to identify regions of interest and only performing intensive lip-movement-to-sound correlation in those regions, the system reduces overall processing time while maintaining endpoint detection accuracy.
Data Source
AI summary
Method for voice-based interactive communication using a digital assistant, wherein the method comprises,an attention detection step, in which the digital assistant detects a user attention and as a result is set into a listening mode;a speaker detection step, in which the digital assistant detects the user as a current speaker;a speech sound detection step, in which the digital assistant detects and records speech uttered by the current speaker, which speech sound detection step further comprises a lip movement detection step, in which the digital assistant detects a lip movement of the current speaker;a speech analysis step, in which the digital assistant parses said recorded speech and extracts speech-based verbal informational content from said recorded speech; anda subsequent response step, in which the digital assistant provides feed-back to the user based on said recorded speech.


