Digital Assistant Gaze and Lip Detection for Voice Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional digital assistants face challenges in accurately detecting user attention, distinguishing between users' phrases, and maintaining context in voice-based interactions, leading to complex and adaptive user experiences.

Innovation Solution

A method for voice-based interactive communication that includes attention detection, speaker detection, speech sound detection with lip movement analysis, and response provision, using sensors like sound, image, and gaze detection to set the digital assistant into a listening mode, track the current speaker, and parse speech for context-aware feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If conventional digital assistants use predetermined catch-phrases to detect user attention, then the detection mechanism is simple, but the accuracy of detecting user attention and phrase start point deteriorates

Engineering Contradiction:
Improvedetection mechanism complexityVSAvoiduser attention detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines multiple detection modalities (audio signal analysis, lip movement detection via image sensor, and gaze detection) into a unified attention detection system. This merging of detection methods allows the digital assistant to accurately determine when a user is paying attention and when speech begins, resolving the contradiction between simple detection mechanisms and accurate attention detection.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces lip movement detection as an intermediary signal between user intent and system response. By detecting lip movements through image sensors and correlating them with audio signals, the system gains an additional layer of information that improves the accuracy of detecting phrase start points and user attention without requiring complex predetermined catch-phrases.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple sensors and detection methods are used to improve user attention and speech detection accuracy, then detection accuracy improves, but device complexity increases

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the image sensor serve multiple functions: it detects lip movements for speech synchronization, detects gaze direction for attention determination, and can potentially identify speaker identity. This multi-functionality allows the system to achieve high speech detection accuracy using a single versatile component rather than multiple specialized sensors, thereby reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent divides the detection process into distinct functional modules: audio signal processing, image signal processing for lip movement detection, gaze detection for attention analysis, and context analysis for speaker identification. Each module handles a specific aspect of the detection task, making the overall complex system manageable and maintainable while achieving high accuracy through the coordinated operation of specialized sub-systems.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If the digital assistant continuously monitors speech and lip movements to accurately detect phrase endpoints, then endpoint detection accuracy improves, but processing time and energy consumption increase

Engineering Contradiction:
Improveendpoint detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent employs periodic sampling of lip movement data during speech detection rather than continuous monitoring. By analyzing lip movement at regular intervals and correlating these periodic samples with audio signal characteristics, the system achieves accurate endpoint detection while reducing the computational burden and processing time compared to continuous analysis.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The patent performs preliminary analysis of audio signals to identify potential speech segments before conducting detailed lip movement correlation. By pre-processing the audio signal to identify regions of interest and only performing intensive lip-movement-to-sound correlation in those regions, the system reduces overall processing time while maintaining endpoint detection accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11699439B2Digital assistant and a corresponding method for voice-based interactive communication based on detected user gaze indicating attention
Publication Date: 2023.07.11 TOBII TECH AB
  • US11699439B2 patent drawing
  • US11699439B2 patent drawing
  • US11699439B2 patent drawing

AI summary

Method for voice-based interactive communication using a digital assistant, wherein the method comprises,an attention detection step, in which the digital assistant detects a user attention and as a result is set into a listening mode;a speaker detection step, in which the digital assistant detects the user as a current speaker;a speech sound detection step, in which the digital assistant detects and records speech uttered by the current speaker, which speech sound detection step further comprises a lip movement detection step, in which the digital assistant detects a lip movement of the current speaker;a speech analysis step, in which the digital assistant parses said recorded speech and extracts speech-based verbal informational content from said recorded speech; anda subsequent response step, in which the digital assistant provides feed-back to the user based on said recorded speech.