3D Optical Gesture Decoding for Accessible Virtual Assistant Communication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional virtual assistants are inadequate for users with special needs, such as speech- or hearing-impaired individuals, as they rely on verbal communication, lacking the ability to interpret non-verbal cues like sign language or lip movements.

Innovation Solution

Augmenting virtual assistant devices with three-dimensional optical sensors and AI-based Machine Learning models trained on user-specific data to decipher non-verbal communication, such as sign language and lip movements, and generating three-dimensional avatars for real-time or preconfigured responses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional virtual assistants use voice-based communication, then they can process information efficiently, but they cannot serve users with speech or hearing impairments

Engineering Contradiction:
Improvecommunication capability for special needs usersVSAvoidcommunication reliability for speech-impaired users
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The virtual assistant system is enhanced to perform multiple communication functions: it can process both verbal commands (through existing voice recognition) and non-verbal gestures (through new computer vision capabilities). This multi-functionality allows the same system to serve both able-bodied users and users with speech or hearing impairments, resolving the contradiction between efficiency and adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces computer vision technology as an intermediary that translates visual gestures into computational commands. This intermediary layer enables users who cannot speak or hear to communicate with the virtual assistant, bridging the gap between their non-verbal expressions and the system's processing capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If virtual assistants are augmented with three-dimensional optical sensors and AI models, then non-verbal communication can be deciphered accurately, but device complexity increases

Engineering Contradiction:
Improvegesture recognition accuracyVSAvoidsensor and processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The gesture recognition system is divided into specialized components: three-dimensional optical sensors for capturing spatial data, AI-based engines for pattern recognition, and Machine Learning models for interpretation. This segmentation allows each component to be optimized independently, achieving high measurement precision while managing overall system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system replaces traditional mechanical or manual input methods with optical sensing and AI-based gesture recognition. This substitution eliminates the need for physical contact or complex mechanical interfaces, achieving accurate non-verbal communication through field-based (optical) sensing and intelligent processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Speed

If the system captures and processes visual cues in real-time, then communication responsiveness improves, but energy consumption increases

Engineering Contradiction:
Improvecommunication response speedVSAvoidenergy consumption for sensor and processing operations
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system employs periodic action by capturing visual cues at optimized intervals rather than continuously. The three-dimensional optical sensors and AI-based engine process gestures when motion is detected or at scheduled frames, reducing computational load and energy consumption while maintaining real-time responsiveness for active communication.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system applies partial action by focusing processing resources only on relevant visual data. Rather than analyzing all captured frames equally, the AI-based engine identifies and processes only those visual cues that contain meaningful gesture information, reducing overall computational energy requirements while maintaining accurate gesture recognition.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12393281B2Multi-dimensional extrasensory deciphering of human gestures for digital authentication and related events
Publication Date: 2025.08.19 BANK OF AMERICA CORP
  • US12393281B2 patent drawing
  • US12393281B2 patent drawing
  • US12393281B2 patent drawing

AI summary

Virtual assistant devices and related methods are described that allow users having special needs to communicate with the virtual assistant device. The virtual assistant devices are augmented with additional intelligent sensors, including, but not limited to, three-dimensional optical sensors that are used to capture optical signals of user's non-verbal communication (e.g., sign language, lip movements, gestures or the like). Additionally, the virtual assistant device includes an Artificial Intelligence (AI)-based engine that includes one or more Machine Learning (ML) models trained on user-specific data and used to determine the non-verbal communication of the user (i.e., sign language, lip movements or other gestures) based on inputs derived from the sensors. Moreover, the virtual assistant devices may be configured to generate and display visual information, such as three-dimensional floating avatars of the user and/or the virtual assistant.