Synthetic 3D Data Generation for Deep Learning Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for extracting 3D measurements from 2D images require large amounts of training data, which is difficult to acquire, and impose restrictions on user pose, background, and clothing, making them cumbersome and inefficient.
Innovation Solution
A system and method for generating large datasets for training deep learning networks by using a 3D base mesh model and augmentation data such as skin color, face contours, hair styles, lighting, and virtual clothing to create thousands or millions of training data points from a single 3D model, allowing for accurate measurements from 2D photos taken with a mobile device without specific poses or backgrounds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If large amounts of training data are acquired from willing volunteers and manually segmented and annotated, then training data quantity is improved, but acquisition difficulty and manual labor increase significantly
Solution Approach 1:
The patent uses a single 3D base mesh model as a template that is copied and transformed multiple times through virtual augmentation. Instead of collecting data from many volunteers, the system creates synthetic training data by applying different virtual clothing, lighting conditions, camera angles, and background environments to the same base model, generating thousands of training samples from one source.
Solution Approach 2:
The system changes multiple parameters of the base 3D model including lighting conditions, camera positions and angles, background environments, and virtual clothing configurations. By systematically varying these parameters, the system generates diverse training data that simulates real-world variability without requiring diverse physical subjects.
2Measurement precision
If restricted pose and background requirements are imposed on data collection, then measurement accuracy is improved, but user convenience and operational ease deteriorate
Solution Approach 1:
The patent performs preliminary actions by pre-augmenting the training data with various poses, backgrounds, lighting conditions, and camera angles before the actual measurement task. The deep learning network is trained in advance on this diverse synthetic data, enabling it to handle real-world variations without requiring users to control poses or backgrounds during actual use.
Solution Approach 2:
The system makes the measurement process self-adaptive by training the network to automatically handle various conditions. Once trained on augmented data covering multiple scenarios, the network can process images taken in diverse real-world conditions without requiring user intervention to control environmental factors, making the system robust and convenient.
3Reliability
If specialized 3D cameras and controlled environments are used, then measurement reliability is improved, but device complexity and cost increase
Solution Approach 1:
The patent replaces specialized mechanical 3D imaging hardware with a software-based solution. Instead of using complex 3D cameras and controlled physical environments, the system uses a standard mobile device camera combined with deep learning algorithms trained on synthetic data. The reliability is achieved through computational methods rather than mechanical precision.
Solution Approach 2:
The patent introduces an intermediary element - the deep learning network trained on augmented synthetic data - that bridges the gap between simple 2D camera inputs and accurate 3D measurements. This intermediary processes the images and extracts measurements without requiring the camera itself to be complex or the environment to be controlled.
Data Source
AI summary
Disclosed are systems and methods for generating large data sets for training deep learning networks (DLNs) for 3D measurements extraction from 2D images taken using a mobile device camera. The method includes the steps of receiving a 3D model of a 3D object; extracting spatial features from the 3D model; generating a first type of augmentation data for the 3D model, such as but not limited to skin color, face contour, hair style, virtual clothing, and/or lighting conditions; augmenting the 3D model with the first type of augmentation data to generate an augmented 3D model; generating at least one 2D image from the augmented 3D model by performing a projection of the augmented 3D model onto at least one plane; and generating a training data set to train the deep learning network (DLN) for spatial feature extraction by aggregating the spatial features and the at least one 2D image.


