Synthetic Training Data Meshes for Unknown Object Grasping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer vision methods lack efficient generation of accurate training data for object detection in unstructured scenes, particularly for unknown objects, which hinders the ability to identify and grasp objects accurately in robotic applications.
Innovation Solution
A method for generating synthetic training data by creating object meshes from images, adding keypoints, and determining material properties, allowing for the training of object detectors to identify graspable faces and select grasping points without requiring complex geometric scans or manual input.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If complex geometric scans or manual input are used to generate training data, then measurement precision and manufacturing precision improve, but device complexity and productivity worsen
Solution Approach 1:
The patent uses photogrammetry to create accurate 3D copies (meshes) of objects from multiple 2D images. Instead of complex geometric scanning, the system captures photographs from multiple angles and reconstructs the object geometry through computational algorithms, producing training data that accurately represents the object's shape and structure without requiring sophisticated scanning hardware
Solution Approach 2:
The patent replaces mechanical/geometric scanning systems with a photogrammetry-based approach using standard cameras and computational processing. The system substitutes complex mechanical measurement devices with optical capture and algorithmic reconstruction, achieving comparable or superior precision while reducing hardware complexity
2Manufacturing precision
If complex geometric scans or manual input are used to generate training data, then measurement precision and manufacturing precision improve, but productivity worsens
Solution Approach 1:
The photogrammetry process rapidly captures multiple images of an object and automatically reconstructs 3D meshes through computational algorithms. This copying approach from 2D images to 3D representations is significantly faster than traditional geometric scanning while maintaining high precision, enabling rapid generation of training data for multiple objects
Solution Approach 2:
The system performs automatic mesh generation, keypoint detection, and training data compilation without requiring manual input or intervention. The photogrammetry pipeline self-processes the images to produce ready-to-use training data, eliminating time-consuming manual operations and accelerating the overall workflow
3Adaptability or versatility
If synthetic training data is generated using photogrammetry and object meshes, then adaptability improves for unknown objects, but measurement precision may worsen compared to geometric scans
Solution Approach 1:
The photogrammetry-based system creates a universal workflow that handles any photographable object, including unknown objects without pre-existing models. The same image-capture-and-reconstruct pipeline works for diverse object types, making the system adaptable and versatile while maintaining sufficient geometric accuracy for robotic manipulation tasks
Solution Approach 2:
The system adjusts image capture parameters (number of images, angles, lighting conditions) and mesh reconstruction parameters to optimize the balance between adaptability and precision. By tuning these parameters, the system achieves adequate geometric accuracy for grasping applications while maintaining broad applicability to unknown objects
Data Source
AI summary
A method for generating training data can include: determining a set of images; determining a set of masks based on the images; determining a first mesh based on the set of masks; optionally determining a refined mesh by recomputing the first mesh; optionally determining one or more faces of the refined mesh; optionally adding one or more keypoints to the refined mesh; optionally determining a material property set for the object; optionally generating a full object mesh; determining one or more scenes; optionally determining training data based on the one or more scenes; optionally training one or more object detectors using the training data; and detecting one or more objects using the trained object detector.


