Synthetic Hand Models for Deep Learning Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Training deep learning networks to recognize 3D objects with multiple moving parts, such as the human hand, is challenging due to the complexity of poses and the labor-intensive process of generating and labeling training images.
Innovation Solution
A training platform generates synthetic models of a hand, including various components like fingers, wrist, and palm, based on data defining potential poses. These synthetic models are then processed to create additional variations, which are used to train a deep learning network for tasks like image segmentation, object recognition, and motion recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If actual hand images are captured and labeled for training, then training data quality is improved, but time consumption and labor intensity increase significantly
Solution Approach 1:
The patent creates synthetic copies of hand images through 3D modeling and rendering instead of capturing actual images. A 3D hand model is generated with anatomically correct structures, then rendered into multiple 2D images showing different poses, angles, and lighting conditions. This copying approach provides high-quality training data without the time-consuming process of actual image capture and manual labeling.
Solution Approach 2:
The patent performs preliminary actions by pre-defining the 3D hand model structure, joint configurations, and pose parameters before generating training images. The system establishes the anatomical framework, skin texture, and structural relationships in advance, then systematically generates diverse training images from this prepared model, significantly reducing the overall time required for training data preparation.
2Reliability
If more training images with various poses are generated, then deep learning network performance is improved, but the complexity of data preparation increases
Solution Approach 1:
The patent implements dynamics by creating a parametric 3D hand model where joints and segments can be dynamically positioned to create various poses. The model allows flexible manipulation of finger joints, palm orientation, and wrist positions through defined rotation axes and angular ranges, enabling systematic generation of diverse training images without manually creating each pose separately.
Solution Approach 2:
The patent segments the hand into anatomically correct components including fingers, palm, wrist, and individual joint segments. Each segment is independently modeled and can be transformed separately, allowing systematic generation of various poses by rotating individual segments around defined axes. This segmentation simplifies the complexity of managing full-hand poses by breaking them down into manageable joint rotations.
3Use of energy by moving object
If synthetic models are generated instead of actual images, then processing resources are conserved, but the realism of training data may be reduced
Solution Approach 1:
The patent applies parameter changes by systematically varying rendering parameters including lighting directions, camera angles, background colors, and skin texture properties when generating synthetic hand images. These parameter variations create diverse and realistic-looking training images from the same 3D model, improving visual fidelity while maintaining the efficiency advantages of synthetic generation.
Data Source
AI summary
In some implementations, a training platform may receive data for generating synthetic models of a body part, such as a hand. The data may include information relating to a plurality of potential poses of the hand. The training platform may generate a set of synthetic models of the hand based on the information, where each synthetic model, in the set of synthetic models, representing a respective pose of the plurality of potential poses. The training platform may derive an additional set of synthetic models based on the set of synthetic models by performing one or more processing operations with respect to at least one synthetic model in the set of synthetic models, and causing the set of synthetic models and the additional set of synthetic models to be provided to a deep learning network to train the deep learning network to perform image segmentation, object recognition, or motion recognition.


