Multi-View Hand Pose Estimation via Stereo Imaging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current hand pose estimation methods face challenges in accuracy and stability due to interference from complex backgrounds and the lack of scale information in monocular images, leading to inaccurate finger pose prediction and 3D position estimation.
Innovation Solution
A method and apparatus for image processing that captures hand images from multiple viewing angles to determine finger and palm pose information separately, integrating these to improve prediction accuracy and efficiency, using a pose prediction model with feature extraction and fusion modules to handle diverse viewing angles and provide more accurate 3D pose estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If monocular images are used for hand pose estimation, then the processing speed is fast, but the measurement precision and reliability are poor due to lack of scale information and background interference
Solution Approach 1:
The patent transitions from 2D monocular images to 3D multi-view imaging by introducing depth information through a stereo camera system. Multiple cameras capture hand images from different spatial angles, creating a three-dimensional representation that provides scale information and reduces background interference, thereby improving measurement precision without excessive complexity increase
Solution Approach 2:
The patent separates the hand pose estimation process into distinct modules: background subtraction to isolate the hand region, feature point detection on the segmented hand, and pose calculation. This segmentation allows each module to optimize for its specific function, improving overall accuracy while managing system complexity
2Reliability
If multiple viewing angles are used to capture hand images, then the measurement precision and reliability improve, but the loss of time and device complexity increase
Solution Approach 1:
The patent performs background subtraction and hand region segmentation before pose estimation, creating a cleaned input that reduces noise and interference. This preliminary processing of multiple view images establishes reliable hand contours and eliminates background elements in advance, improving prediction stability while optimizing the subsequent pose calculation process
Solution Approach 2:
The patent fuses information from multiple camera views by integrating features detected across different angles. The system combines correspondence points from multiple views to calculate three-dimensional pose, merging redundant information to improve reliability while processing efficiency through coordinated multi-camera synchronization
Data Source
AI summary
Embodiments of the disclosure discloses a method, apparatus, and electronic device for image processing. The method for image processing includes obtaining a plurality of hand images, the plurality of hand images being images of a target hand captured from a plurality of viewing angles; determining finger pose information based on the plurality of hand images; determining palm pose information based on the plurality of hand images; and determining pose information of the target hand based on the finger pose information and the palm pose information. The method for image processing can improve the accuracy of hand pose prediction.


