Stereo Camera Finger Gesture Detection via Offset Image Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current input mechanisms for computing systems, such as keyboards and mice, are not natural and have been difficult to implement effectively for user control.

Innovation Solution

A method using stereo cameras to detect objects, specifically fingers, by receiving images from multiple perspectives and applying machine-learning classifiers to accurately identify and track finger gestures, enabling more natural user interaction through augmented reality interfaces.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional input mechanisms (keyboards and mice) are used for user control, then device compatibility and reliability are maintained, but user interaction naturalness and ease of operation deteriorate

Engineering Contradiction:
Improveuser interaction naturalnessVSAvoidinput mechanism complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces mechanical input devices (keyboards, mice) with an optical-based gesture recognition system using stereo cameras and machine learning classifiers. This substitution enables natural hand and finger gesture recognition, improving ease of operation while eliminating the need for complex mechanical input mechanisms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If stereo camera-based gesture recognition is implemented, then user interaction naturalness is improved, but system complexity and difficulty of detecting and measuring increase

Engineering Contradiction:
Improvegesture recognition capabilityVSAvoidobject detection accuracy
Core Design Contradiction:
Ease of operationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies machine learning classifiers in advance to train the system to recognize specific hand and finger gestures. This preliminary training enables the system to accurately detect and measure gesture patterns, reducing the difficulty of object detection while maintaining high recognition accuracy for natural user interactions.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces machine learning classifiers as an intermediary between the stereo camera input and the gesture recognition output. These classifiers process the complex stereo image data and translate it into recognizable gesture patterns, simplifying the detection and measurement process while improving accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If offset images are processed through machine-learning classifiers, then object detection accuracy is improved, but processing time and computational complexity increase

Engineering Contradiction:
Improvefinger detection accuracyVSAvoidgesture recognition processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies multiple machine learning classifiers to different aspects of the offset images (hand detection, finger detection, gesture classification). By distributing the detection task across multiple specialized classifiers rather than using a single comprehensive classifier, the system achieves high accuracy while optimizing processing efficiency through parallel evaluation.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10592778B2Stereoscopic object detection leveraging expected object distance
Publication Date: 2020.03.17 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10592778B2 patent drawing
  • US10592778B2 patent drawing
  • US10592778B2 patent drawing

AI summary

A method of object detection includes receiving a first image taken from a first perspective by a first camera and receiving a second image taken from a second perspective, different from the first perspective, by a second camera. Each pixel in the first image is offset relative to a corresponding pixel in the second image by a predetermined offset distance resulting in offset first and second images. A particular pixel of the offset first image depicts a same object locus as a corresponding pixel in the offset second image only if the object locus is at an expected object-detection distance from the first and second cameras. The method includes recognizing that a target object is imaged by the particular pixel of the offset first image and the corresponding pixel of the offset second image.