Visual Gesture Recognition for Hands-Free Mobile Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

User interfaces on mobile devices require manual interaction, such as tapping or swiping, which can be inconvenient when using the device with both hands, like when taking a picture, as it necessitates removing a hand to activate commands.

Innovation Solution

Implementing a system that allows users to select and activate features using visual gestures by physically moving and reorienting the device, leveraging sensors and augmented reality to overlay virtual objects on real-time images, enabling actions without direct screen interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual interaction (tapping, swiping) is required to activate features, then precise control and clear user intent are achieved, but ease of operation deteriorates when using the device with both hands

Engineering Contradiction:
Improveease of operationVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces mechanical touchscreen interaction (tapping, swiping) with visual gesture recognition through the camera. The system captures images of the user's hand gestures and processes them to determine user intent, substituting the mechanical contact interface with an optical recognition system that enables hands-free operation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary visual gesture recognition system between the user and the device features. Instead of direct touchscreen contact, the user performs gestures in front of the camera, and the system mediates by interpreting these gestures to activate corresponding features, enabling operation while holding the device with both hands.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If visual gesture recognition is implemented, then ease of operation improves for hands-free interaction, but device complexity increases due to additional sensors and processing

Engineering Contradiction:
Improveease of operationVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent makes the existing camera serve multiple functions: its primary function for capturing photos and videos, and a secondary function for visual gesture recognition. By reusing the camera hardware for both purposes, the system avoids adding dedicated gesture sensors, thereby limiting the increase in device complexity while still enabling hands-free operation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If visual gestures are used for feature activation, then productivity improves by eliminating manual input steps, but measurement precision deteriorates in determining user intent

Engineering Contradiction:
ImproveproductivityVSAvoidmeasurement precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent employs preliminary action by having the user perform a distinct gesture in front of the camera before feature activation. This preliminary visual action serves as a clear signal of user intent, allowing the system to accurately determine which feature to activate without ambiguity, thus maintaining measurement precision while improving productivity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10585473B2Visual gestures
Publication Date: 2020.03.10 META PLATFORMS TECHNOLOGIES LLC
  • US10585473B2 patent drawing
  • US10585473B2 patent drawing
  • US10585473B2 patent drawing

AI summary

A device has a display, a first camera, a second camera and a virtual gesture application. The first camera generates an image that depicts a physical object. The second camera tracks a position of a stare of a user. The virtual gesture application identifies the physical object using the image, generates a virtual object corresponding to the identified physical object, renders the virtual object in the display based a position of the display relative to the physical object, identifies an area in the display corresponding to the position of the stare of the user, determines that an interactive feature of the virtual object is located inside the area, and performs at least one action on the interactive feature in response to determining that the interactive feature is located inside the area.