Intent Detection via Gesture and Environment Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computing devices lack an effective method to detect user intent without traditional input devices like mice or keyboards, especially in scenarios where screens are absent, and existing technologies are limited in interpreting real-world objects and gestures.

Innovation Solution

A system that captures images using a single non-depth sensing camera, employs machine learned models to detect hand gestures and identify user intent based on both the gesture and environment, allowing for the execution of tasks such as searching or interacting with digital assistants.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional input devices like mice or keyboards are used, then user intent detection is reliable, but device complexity and physical hardware requirements increase

Engineering Contradiction:
Improveuser intent detection accuracyVSAvoidphysical input hardware
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces mechanical input devices (mouse, keyboard) with an optical system consisting of a camera and machine learning model. The camera captures images of hand gestures, and the ML model processes these images to detect gestures and determine user intent, eliminating the need for physical input hardware while maintaining reliable intent detection.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If a single non-depth sensing camera is used, then device complexity is reduced, but gesture detection precision deteriorates

Engineering Contradiction:
Improvecamera systemVSAvoidgesture detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent changes the processing parameters by applying machine learning algorithms that can interpret 2D camera data to detect 3D hand gestures. The ML model learns to infer depth and gesture meaning from single-camera 2D images through training on labeled gesture data, achieving accurate gesture detection without depth sensing hardware.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If machine learned models are used to detect hand gestures, then ease of operation is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvenatural gesture interactionVSAvoidgesture processing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-training the machine learning model on large datasets of hand gestures before deployment. This offline training phase prepares the model to quickly and accurately classify gestures in real-time during operation, reducing processing time during actual use while maintaining high accuracy for natural gesture interaction.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11960793B2Intent detection with a computing device
Publication Date: 2024.04.16 GOOGLE LLC
  • US11960793B2 patent drawing
  • US11960793B2 patent drawing
  • US11960793B2 patent drawing

AI summary

A method can perform a process with a method including capturing an image, determining an environment that a user is operating a computing device, detecting a hand gesture based on an object in the image, determining, using a machine learned model, an intent of a user based on the hand gesture and the environment, and executing a task based at least on the determined intent.