Non-coaxial Camera Array for 3D Gesture Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current imaging systems for gesture recognition, such as the KINECT and PLAYSTATION MOVE, face limitations in depth detection sensitivity, requiring large movements and fixed camera positions, which restrict their ability to accurately recognize gestures and are computationally intensive, making them inflexible and costly to set up.
Innovation Solution
A system that combines disparate cameras with non-coaxial axes to detect and infer 3D gestures without the need for precise calibration or extensive computation, allowing for flexible placement and setup of camera components, including smartphones and webcams, to create a cohesive gesture recognition environment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If structured light projection technology and time-of-flight sensors are used for depth detection, then 3D gesture recognition capability is achieved, but depth detection sensitivity is limited and large movements are required
Solution Approach 1:
The patent replaces structured light projection and time-of-flight sensing mechanisms with a passive stereo vision system using conventional cameras. Instead of active illumination and temporal measurement, the system uses geometric triangulation from multiple camera viewpoints to achieve depth detection, thereby improving sensitivity without requiring large gesture movements.
Solution Approach 2:
The patent employs conventional cameras that can serve multiple purposes: capturing 2D images for standard photography and simultaneously providing depth information through stereo triangulation. This multi-functionality eliminates the need for specialized depth-sensing hardware, improving both sensitivity and operational ease.
2Measurement precision
If fixed camera positions with known calibration parameters are used for triangulation, then 3D coordinate recognition is achieved, but system setup complexity and cost increase
Solution Approach 1:
The patent implements self-calibration capability where the system automatically determines camera parameters and relative positions through computational methods rather than requiring manual precision calibration. This self-service approach maintains 3D recognition accuracy while dramatically reducing setup complexity and cost.
Solution Approach 2:
The patent changes the calibration approach from fixed, pre-determined parameters to dynamically computable parameters. By using image processing and geometric analysis to derive camera parameters from actual captured images, the system adapts to different camera configurations without requiring precise manual calibration.
3Measurement precision
If stereo imaging systems with fixed optical axes are used, then depth construction is achieved, but system flexibility and adaptability decrease
Solution Approach 1:
The patent transitions from static, fixed camera arrangements to a dynamic system where cameras can be positioned flexibly and the system adapts through computational calibration. The optical axes no longer need to be predetermined, allowing cameras to be placed in various configurations while maintaining depth construction accuracy through software-based adjustment.
Solution Approach 2:
The patent allows camera parameters such as position, orientation, and focal length to be determined computationally rather than fixed in advance. This parameter flexibility enables arbitrary camera placements while maintaining accurate depth construction through post-capture calibration and triangulation algorithms.
4Measurement precision
If feature matching and triangulation calculations are used for gesture recognition, then 3D gesture detection is achieved, but computational intensity increases
Solution Approach 1:
The patent applies partial action by performing feature matching and triangulation only on relevant regions of interest rather than processing entire images. By focusing computational resources on areas containing gesture information and using simplified triangulation models, the system maintains detection accuracy while reducing overall computational energy consumption.
Data Source
AI summary
The subject system hardware and methodology combine disparate cameras into a cohesive gesture recognition environment. To render an intended computer, gaming, display, etc. control function, two or more cameras with non-coaxial axes are trained on a space to detect and lock onto an object image regardless of its depth coordinate. Each camera captures one 2D view of the gesture and the plurality of 2D gestures are combined to infer the 3D input.


