3D Skeleton Capture Using ROI-Guided Multi-Camera Hand Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing wearable electronic devices with limited battery capacity face challenges in accurately estimating 3D hand poses and gestures while minimizing power consumption, particularly when using multiple cameras for hand interaction in AR environments.
Innovation Solution
An electronic device and method that reduces power consumption by detecting a Region Of Interest (ROI) in images from one camera and using multiple cameras to obtain 3D skeleton data, leveraging deep learning models to identify keypoints and project coordinates for accurate 3D pose estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If detection operations are performed on images from all cameras to obtain 3D skeleton data, then measurement precision is improved, but use of energy increases
Solution Approach 1:
The patent divides the image processing task into two stages: first, perform detection operations only on images from a reference camera to identify ROI; second, use the ROI information to guide processing of images from other cameras. This segmentation reduces the number of full detection operations while maintaining 3D skeleton data accuracy through selective processing.
Solution Approach 2:
The patent performs preliminary detection operations on the reference camera image to obtain ROI information before processing other camera images. This preliminary action identifies regions of interest that guide subsequent processing, reducing unnecessary computation on non-relevant areas of other camera images and lowering overall power consumption.
2Productivity
If detection operations are performed on images from all cameras, then productivity is improved, but loss of time increases
Solution Approach 1:
The patent segments the computation process into preliminary ROI detection on reference camera images and subsequent targeted processing of other camera images. This segmentation reduces total computation time by avoiding redundant full-image detection operations while maintaining productivity through efficient use of ROI guidance.
Solution Approach 2:
The patent performs preliminary ROI detection on reference camera images before processing other camera images. This preliminary action establishes regions of interest that accelerate subsequent processing by limiting computation to relevant areas, thereby reducing total computation time while maintaining high productivity in 3D skeleton data acquisition.
3Use of energy by moving object
If ROI detection is performed only on reference camera images, then use of energy is reduced, but measurement precision may deteriorate
Solution Approach 1:
The patent uses ROI information from reference camera images as an intermediary to guide processing of other camera images. This intermediary approach maintains measurement precision by ensuring that ROI guidance from the reference camera is applied to correctly identify corresponding regions in other camera images, while still reducing power consumption through selective processing.
Solution Approach 2:
The patent replaces the mechanical approach of performing full detection operations on all camera images with an information-based approach using ROI guidance. By substituting comprehensive detection with ROI-guided processing, the system reduces power consumption while maintaining measurement precision through intelligent use of reference camera information.
Data Source
AI summary
An electronic device performs a method of obtaining Three-Dimensional (3D) skeleton data of an object obtained by using a first camera and a second camera. The method includes: obtaining a first image using the first camera and obtaining a second image using the second camera; obtaining, from the first image, a first Region Of Interest (ROI) comprising the object; obtaining, from the first ROI, first skeleton data comprising at least one keypoint of the object; obtaining a second ROI from the second image, based on the first skeleton data and information about a relative position between the first camera and the second camera; obtaining, from the second ROI, second skeleton data comprising at least one keypoint of the object; and obtaining 3D skeleton data of the object, based on the first skeleton data and the second skeleton data.


