Stereo Camera 3D Hover Gesture Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems for sensing user interaction in three-dimensional hover zones adjacent to display surfaces are costly, power-intensive, and require specialized components, while also facing challenges with occlusion, accuracy, and computational overhead, especially when trying to detect gestures before actual touch on the screen.
Innovation Solution
The use of two generic, pre-calibrated off-the-shelf cameras with a processor and software system that captures two-dimensional images from different vantage points, processes them to identify landmark points, and reconstructs three-dimensional data only for relevant points, allowing for gesture recognition without the need for specialized hardware or high computational power.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If time-of-flight (TOF) systems are used to implement three-dimensional sensing in hover zones, then true three-dimensional data acquisition is achieved, but system cost and power consumption increase significantly
Solution Approach 1:
The patent replaces active TOF sensing mechanisms with passive optical capture using standard cameras. Instead of emitting optical energy and measuring reflection time, the system uses ambient or display-emitted light to capture images, eliminating the need for power-intensive TOF hardware while achieving sufficient depth information through stereo vision geometry
Solution Approach 2:
The invention substitutes expensive specialized TOF sensors with inexpensive off-the-shelf cameras that can be mass-produced at low cost. The system achieves three-dimensional sensing capability using commodity hardware rather than proprietary expensive components, making the technology economically viable for consumer devices
2Measurement precision
If structured-light systems are used to obtain three-dimensional data, then depth information is acquired, but system complexity and cost increase
Solution Approach 1:
The patent replaces complex structured-light projection systems with simple passive optical capture. Instead of projecting coded light patterns and analyzing distortions, the system uses standard camera imaging with geometric triangulation, eliminating the need for pattern projectors, specialized optics, and complex decoding algorithms
Solution Approach 2:
The invention uses standard cameras that can perform both two-dimensional image capture and three-dimensional depth sensing through stereo geometry. This multi-functional approach eliminates the need for separate specialized depth-sensing hardware, reducing system complexity while maintaining depth acquisition capability
3Ease of manufacture
If camera-based optical sensing is used to implement two-dimensional touch screen system, then cost is reduced, but three-dimensional hover detection capability is lost
Solution Approach 1:
The patent extends two-dimensional camera imaging to three-dimensional sensing by using stereo vision geometry. By capturing images from multiple camera viewpoints and triangulating landmark positions, the system recovers depth information (z-coordinate) in addition to the standard (x,y) coordinates, enabling hover zone detection without adding specialized three-dimensional hardware
Solution Approach 2:
The system pre-identifies landmark points on the display surface and in the hover zone before actual touch occurs. By continuously tracking these landmarks as users approach and interact with the display, the system can detect gestures in the hover zone and distinguish between hovering and touching, enabling versatile interaction modes
4Reliability
If retro-reflective strips are added to display bezel for optical sensing, then touch detection is enabled, but display thickness and cost increase
Solution Approach 1:
The patent replaces physical retro-reflective strips with virtual optical pathways created by camera geometry. Instead of requiring physical reflective elements on the display bezel, the system uses computational methods and camera positioning to achieve optical sensing, eliminating the need for additional physical layers that increase thickness
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables natural user interaction with minimal resource expenditure, achieving high accuracy and low power consumption, and is compatible with various display sizes from smartphones to large screens, allowing for efficient gesture recognition in both two-dimensional touch and three-dimensional hover zones.
Implementation Method 1
a camera and an optical emitter, perhaps an IR LED, are disposed at each upper corner region of the screen with (x,y) fields of view (FOV) that ideally encompass all of the screen
Implementation Method 2
The vertical sides and the horizontal bottom of the inner surfaces of the display bezel are lined with retro-reflective strips that reflect-back energy from the two optical emitters
Implementation Method 3
You can determine the (x,y) location of the touch on the display screen surface by combining the centroid of the blob using triangulation providing information is present from both cameras
Implementation Method 4
Such TOF systems emit active optical energy and determine distance (x,y,z) to a target by counting how long it takes for reflected-back emitted optical energy to be sensed
Data Source
AI summary
User interaction with a display is detected using at least two cameras whose intersecting FOVs define a three-dimensional hover zone within which user interactions can be imaged. Each camera substantially simultaneously acquires from its vantage point two-dimensional images of the user within the hover zone. Separately and collectively the image data is analyzed to identify therein a relatively few landmarks definable on the user. A substantially unambiguous correspondence is established between the same landmark on each acquired image, and as to those landmarks a three-dimensional reconstruction is made in a common coordinate system. This landmark identification and position information can be converted into a command causing the display to respond appropriately to a gesture made by the user. Advantageously size of the hover zone can far exceed size of the display, making the invention usable with smart phones as well as large size entertainment TVs.


