3D Model Training Images with Automatic Occlusion-Aware Polygons
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Creating training data sets for machine learning-based image recognition processes is resource-intensive and laborious, particularly in capturing images of objects from various angles and backgrounds with occlusions, which are crucial for robust recognition.
Innovation Solution
Utilizing computer-generated three-dimensional models to create annotated images of objects with bounding polygons, which efficiently capture objects from different viewpoints and occlusions, reducing the need for manual labor and increasing the efficiency of training data set creation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual image capture and annotation is used to create training data sets, then image recognition accuracy can be improved through diverse views and occlusions, but the process becomes labor intensive and resource consuming
Solution Approach 1:
The patent uses computer-generated three-dimensional models to create synthetic training images that replicate real-world object views, occlusions, and backgrounds. These virtual copies of real objects replace the need for manual photography and annotation, maintaining training quality while dramatically reducing labor requirements and increasing productivity
Solution Approach 2:
The patent replaces manual mechanical processes (physical image capture, manual bounding box drawing, label assignment) with automated computer-generated content. The system automatically generates diverse training images with various viewpoints and occlusions through virtual modeling, eliminating human labor from the training data creation pipeline
2Reliability
If a large number of diverse training images are captured manually, then the robustness of image recognition is improved, but the cost and time required increase significantly
Solution Approach 1:
The patent pre-generates a comprehensive library of three-dimensional models representing objects, scenes, and their various configurations before training begins. This preliminary creation of virtual training data allows the system to rapidly generate diverse training images on demand without time-consuming manual capture processes
Solution Approach 2:
By creating synthetic copies of real-world scenarios through three-dimensional modeling, the system can generate unlimited training images with various viewpoints, occlusions, and backgrounds instantly. This virtual replication approach provides diverse training data without the time constraints of manual photography and annotation
3Measurement precision
If manual annotation of objects in training images is performed, then labeling accuracy is improved, but labor costs and processing time increase
Solution Approach 1:
The system performs self-annotation by automatically generating bounding boxes and labels from three-dimensional models. The virtual models inherently provide precise spatial information and object identification, eliminating the need for human annotators to manually draw bounding boxes or assign labels, thereby reducing labor costs while maintaining precision
Data Source
AI summary
Systems and methods for defining bounding polygons in a view of a three-dimensional scene. Rays are defined that each extend from a viewpoint of a virtual three-dimensional model to a vertex of an object of interest in the virtual three-dimensional model. A set of occluded rays is determined that include rays intercepting occluding objects in the virtual three-dimensional model prior to reaching a vertex of the object of interest when extending from the viewpoint. A set of visible rays is defined with respect to the object of interest that excludes the occluded set of rays. A bounding polygon for the object of interest that encompasses each vertex intercepted by the set of visible rays and excludes at least one vertex intercepted by a respective ray in the set of occluded rays is defined in an image of the virtual three-dimensional model that is created from the viewpoint.


