Automatic Hand Key Point Labeling via 3D Reconstruction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating labeled hand data for gesture recognition in VR and AR are labor-intensive and inefficient, requiring manual labeling of key points in images taken from different angles, which is time-consuming and prone to errors.
Innovation Solution
A system and method for automatically generating labeled hand data by acquiring images from multiple angles, detecting key points, reconstructing a 3D representation of the hand, projecting these points onto the images, and using a given finger bone length to create accurate labeled data, reducing the need for manual labor and improving data universality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labeling of hand key points is used to generate training data, then the accuracy of labels can be ensured, but the time consumption and labor cost increase significantly
Solution Approach 1:
The system enables automatic self-labeling of hand key points through multi-view image acquisition and 3D reconstruction. The computer vision algorithm automatically detects and labels key points without human intervention, making the system serve itself in the data annotation process while maintaining high accuracy through geometric constraints from multiple viewing angles.
Solution Approach 2:
The patent replaces the mechanical manual labeling process with an automated computer vision system. Instead of human operators manually marking key points, the system uses image processing algorithms, 3D reconstruction, and projection mathematics to automatically generate accurate labels, substituting human labor with computational processes.
2Reliability
If manual labeling is used to provide training samples, then the quality of training data can be maintained, but the productivity and efficiency of data generation decrease
Solution Approach 1:
The system performs preliminary 3D reconstruction of the hand model from multi-view images before generating the final labeled data. This preliminary action creates an accurate 3D representation that serves as the basis for automatic key point detection and labeling, enabling high-quality data generation without manual intervention and significantly improving productivity.
Solution Approach 2:
The patent introduces a 3D hand model as an intermediary between multi-view images and labeled training data. The 3D reconstruction serves as a mediator that bridges the gap between 2D image data and the required labeled output, enabling automatic generation of accurate training samples while maintaining high data quality and efficiency.
3Adaptability or versatility
If images from multiple angles are acquired and processed, then the universality and comprehensiveness of training data improve, but the device complexity and processing difficulty increase
Solution Approach 1:
The patent transitions from 2D image data to 3D reconstruction by utilizing multiple viewing angles. This dimensional change from 2D to 3D space enables the system to capture comprehensive hand geometry information, improving data universality and adaptability while the structured 3D reconstruction process manages the complexity through mathematical modeling.
Data Source
AI summary
The present disclosure relates to a method for automatically generating labeled data of a hand, comprising: acquiring at least three images to be processed of the hand under different angles of view; detecting key points on the at least three images to be processed respectively; screening the detected key points by using an association relation among the at least three images to be processed, the association relation being the same frame of image of the at least three images to be processed from the hand under different angles of view; reconstructing a three-dimensional space representation of the hand with regard to the key points screened on the same frame of image, in combination with a given finger bone length; projecting the key points on the three-dimensional representation of the hand onto the at least three images to be processed; and generating the labeled data of the hand on the images to be processed by using the projected key points on the at least three images to be processed.


