Automated Assistant Hand Tracking for Reliable Sign Language Invocation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated assistants often rely on audio inputs and camera-based visual inputs, which can be limiting for users with hearing impairments and limited field of view, leading to unreliable invocation and feedback, privacy concerns, and inefficiencies in sign language communication.

Innovation Solution

An automated assistant that detects user presence and intent through motion and hand tracking, provides real-time hand rendering and feedback, and allows for local processing of sign language commands, offering selectable suggestions and streamlined interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a dedicated video camera is used to detect sign language commands, then the automated assistant can receive visual input from users with hearing impairments, but the field of view limitations and unreliable detection make the system unsuitable for sign language interpretation

Engineering Contradiction:
Improveability to receive sign language inputVSAvoidreliability of sign language detection
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the sign language detection process into multiple independent components: hand detection module, gesture recognition module, and command interpretation module. This segmentation allows each component to be optimized independently, improving overall detection reliability while maintaining adaptability to various sign language inputs

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that translates camera coordinates to display coordinates and maps hand positions to gesture meanings. This intermediary layer acts as a buffer that enhances detection reliability by providing multiple processing steps and validation mechanisms

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If a camera is used for non-verbal communications, then the automated assistant can detect user presence and gestures, but privacy concerns arise when the camera is always on

Engineering Contradiction:
Improveease of non-verbal interactionVSAvoidprivacy concerns
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The patent implements periodic camera activation based on user presence detection. The camera is activated only during periods when a user is detected near the device, and deactivated during periods of no user presence. This periodic operation maintains ease of non-verbal interaction while addressing privacy concerns by minimizing continuous camera operation

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system uses the camera's own detection capabilities to determine when it should be active. The camera detects user presence and automatically triggers its own activation state, eliminating the need for manual control and ensuring privacy-respecting operation

Inventive Principle:
Principle #25Self-service

3Loss of information

If audio feedback is used to indicate successful invocation, then the automated assistant can provide confirmation to users, but hearing-impaired users cannot perceive the feedback

Engineering Contradiction:
Improvefeedback deliveryVSAvoidfeedback accessibility to hearing-impaired users
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal feedback system that provides invocation confirmation through multiple channels simultaneously: visual feedback on the display, haptic feedback through the device, and audio feedback. This multi-functional feedback mechanism ensures that hearing-impaired users can perceive confirmation through visual and tactile means while maintaining compatibility with hearing users who can use audio feedback

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces a multi-modal feedback loop that provides real-time confirmation of user actions through multiple senses. The system monitors user gestures, provides immediate visual and haptic feedback, and maintains a feedback history that can be reviewed, ensuring that all users receive adequate confirmation of their interactions

Inventive Principle:
Principle #23Feedback

4Area of stationary object

If the camera field of view is limited, then the device can maintain a compact form factor, but users cannot provide non-verbal input when standing outside the field of view

Engineering Contradiction:
Improvecamera field of viewVSAvoidaccessibility of input
Core Design Contradiction:
Area of stationary objectVSEase of operation

Solution Approach 1:

The patent transforms the limited two-dimensional camera field of view into a three-dimensional interaction space by incorporating depth sensing and multiple camera angles. The system can detect hand gestures from a wider spatial range and map them to the display coordinate system, allowing users to interact from positions outside the traditional camera field of view while maintaining compact device dimensions

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250321643A1Automated assistant adapted to facilitate sign language interactions and discoverability of related functionality
Publication Date: 2025.10.16 GOOGLE LLC
  • US20250321643A1 patent drawing
  • US20250321643A1 patent drawing
  • US20250321643A1 patent drawing

AI summary

Implementations described herein relate to an automated assistant that is responsive to sign language commands and can provide feedback to assist a user with efficiently controlling the automated assistant using sign language. When the user is initially detected, and/or the automated assistant otherwise determines that the user intends to invoke the automated assistant, the automated assistant can render graphical output and/or a depiction of one or both hands of the user (or a representation thereof). In some implementations, this depiction can be a static representation of hands, or a dynamic representation (e.g., an avatar) that mimics the movement of one or both hands of the user. When the user provides a sign language command, an American Sign Language (ASL) Gloss interpretation (or corresponding natural language interpretation thereof) can be rendered at the display interface, along with any autocomplete suggestions and/or suggestions for other commands.