AR Keyboard Transcription via Hand Pose Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Virtual and augmented reality keyboards lack kinesthetic feedback, leading to inaccurate typing due to the absence of physical keys, and existing corrective algorithms are inapplicable to these environments.

Innovation Solution

A computer-implemented method that identifies and analyzes hand poses to generate a sequence of keystrokes using trained transcription models, allowing for accurate transcription of input without relying on kinesthetic feedback by selecting models based on user behavior and attention levels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If virtual or augmented reality keyboards are used, then users can interact in simulated environments, but typing accuracy deteriorates due to lack of kinesthetic feedback

Engineering Contradiction:
Improveability to interact in simulated environmentsVSAvoidtyping accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent replaces the mechanical keyboard system with an optical tracking system that captures hand gestures and poses. Instead of relying on physical key presses, the system uses cameras or sensors to detect and analyze hand movements, transforming the input mechanism from mechanical to optical while maintaining typing functionality in virtual environments

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary transcription model that acts as a bridge between hand gestures and text output. This model analyzes gesture sequences and translates them into intended text, serving as a mediator that compensates for the lack of direct key feedback by interpreting gesture patterns and predicting user intent

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If corrective algorithms are applied to graphical keyboards, then typing accuracy improves, but the algorithms become inapplicable to virtual keyboards that do not generate touch events

Engineering Contradiction:
Improvetyping accuracyVSAvoidapplicability across different keyboard types
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal transcription system that works across multiple keyboard types (graphical, virtual, augmented reality) by using gesture-based input as a common foundation. The system can process both traditional touch events and gesture sequences through a unified transcription model, making it adaptable to different keyboard implementations while maintaining improved accuracy

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If hand pose analysis is implemented, then typing accuracy improves without kinesthetic feedback, but system complexity increases

Engineering Contradiction:
Improvetyping accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-training transcription models on large datasets of gesture sequences and corresponding text. This advance preparation allows the system to make accurate predictions during actual use without requiring complex real-time processing, reducing operational complexity while maintaining high typing accuracy

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10719173B2Transcribing augmented reality keyboard input based on hand poses for improved typing accuracy
Publication Date: 2020.07.21 META PLATFORMS TECHNOLOGIES LLC
  • US10719173B2 patent drawing
  • US10719173B2 patent drawing
  • US10719173B2 patent drawing

AI summary

A transcription engine transcribes input received from an augmented reality keyboard based on a sequence of hand poses performed by a user when typing. A hand pose generator analyzes video of the user typing to generate the sequence of hand poses. The transcription engine implements a set of transcription models to generate a series of keystrokes based on the sequence of hand poses. Each keystroke in the series may correspond to one or more hand poses in the sequence of hand poses. The transcription engine monitors the behavior of the user and selects between transcription models depending on the attention level of the user. The transcription engine may select a first transcription model when the user types in a focused manner and then select a second transcription model when the user types in a less focused, conversational manner.