3D Base Model Augmentation for Synthetic Ground Truth Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generating diverse and inclusive ground truths for training machine learning models, particularly convolutional neural networks (CNNs), is time-consuming and challenging due to the need for extensive image modification and rendering.
Innovation Solution
The use of 3D base models for ground truth inputs and outputs, with augmentations applied to ensure diversity and inclusivity, including blend shape augmentations, mattes, and HDRI lighting, to automate the generation of diverse and inclusive ground truths.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional methods are used to generate ground truths for training machine learning models, then the ground truths can be created with manual effort, but the process is time-consuming and inefficient
Solution Approach 1:
The patent uses 3D base models as templates to generate multiple synthetic images through automated rendering. Instead of manually creating each ground truth image, the system copies and transforms the 3D model under various conditions (lighting, pose, occlusion) to produce diverse training data automatically, dramatically improving productivity while reducing time investment
Solution Approach 2:
The patent performs preliminary actions by pre-defining 3D base models with known ground truth annotations before generating the actual training images. The 3D models are prepared in advance with accurate geometric and semantic information, which then serves as the foundation for automatically generating multiple 2D images with pre-computed ground truths, eliminating the need for time-consuming manual annotation
2Adaptability or versatility
If diverse and inclusive ground truths are generated through extensive image modification and rendering, then the training data quality improves, but the complexity of the generation process increases
Solution Approach 1:
The patent segments the ground truth generation process into distinct modular components: 3D model preparation, augmentation application, image rendering, and ground truth computation. Each module handles a specific aspect of diversity generation (e.g., pose variation, lighting conditions, occlusion objects), making the complex process manageable and可调控 while maintaining high adaptability
Solution Approach 2:
The patent creates a universal 3D base model that can serve multiple functions: it can be rendered from different angles, under various lighting conditions, with different occlusions, and transformed into multiple 2D images. This single 3D model acts as a multi-functional source that generates diverse ground truths across multiple categories and scenarios, achieving high versatility without proportionally increasing complexity
3Reliability
If large datasets are generated to ensure robust training, then the machine learning model performance improves, but the time and effort required for dataset creation increases
Solution Approach 1:
The patent uses automated copying and transformation of 3D base models to generate large numbers of synthetic images with pre-computed ground truths. The system can rapidly produce thousands of diverse training images by rendering the same 3D model under various conditions, creating large robust datasets without the time investment required for manual image collection and annotation
Solution Approach 2:
The patent performs preliminary computation of ground truths during the 3D model preparation and rendering phase. By pre-calculating accurate annotations (segmentation masks, bounding boxes, semantic labels) from the 3D models before generating the final 2D images, the system creates large datasets with reliable ground truths automatically, eliminating the need for time-consuming post-processing and manual verification
Data Source
AI summary
A messaging system processes three-dimensional (3D) models to generate ground truths for training machine learning models for applications of the messaging system. A method of generating ground truths for machine learning includes generating a plurality of first rendered images from a first 3D base model where each first rendered image includes the 3D base model modified by first augmentations of a plurality of augmentations. The method further includes determining for a second 3D base model incompatible augmentations of the first plurality of augmentations, where the incompatible augmentations indicate changes to fixed features of the second 3D base model, and generating a plurality of second rendered images from a second 3D base model, each second rendered image comprising the second 3D base model modified by second augmentations, the second augmentations corresponding to the first augmentations of a corresponding first rendered image, where the second augmentations comprises augmentations of the first augmentations that are not incompatible augmentations.


