Robot Vision Learning for 3D Object Localization Without Manual Calibration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current robotic systems require significant effort and computational resources to customize image analysis tools for each object, struggle with translating 2D image positions to 3D positions, and are challenged by outdoor light effects and image resolution when identifying and locating objects in industrial and mobile applications.
Innovation Solution
A method for generating a dataset mapping visual features of objects by capturing images from different angles and distances, analyzing these features, and associating them with positional information to create a unique identification and localization system that reduces the need for manual calibration and specialized processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If specialized image analysis tools are customized for each object, then object identification accuracy is improved, but processing time and computational effort increase significantly
Solution Approach 1:
The system performs preliminary actions by pre-processing images from multiple angles and distances before actual object identification occurs. A dataset mapping visual features to 3D positions is created in advance through robotic capture of images at various positions, enabling faster real-time identification without requiring specialized processing for each object during operation
Solution Approach 2:
The patent creates a universal mapping dataset that can be applied to identify multiple different objects without requiring customized tools for each object. The system captures images from various positions and creates a generalizable mapping between visual features and 3D positions that works across different object types, reducing the need for object-specific customization
2Manufacturing precision
If sensors and software are calibrated to translate object positions to robot-relative coordinates, then task execution precision is improved, but calibration effort and complexity increase
Solution Approach 1:
The robotic system performs self-calibration by automatically capturing images from multiple positions and generating the mapping dataset without requiring manual calibration procedures. The robot autonomously collects the necessary data by moving to predetermined positions and capturing images, then processes this data to create the position mapping, eliminating the need for technically qualified individuals to perform calibration
Solution Approach 2:
The calibration process is performed in advance by pre-capturing images at multiple positions and pre-computing the mapping relationships. This preliminary action creates a ready-to-use mapping dataset that simplifies subsequent task execution without requiring complex real-time calibration procedures
3Measurement precision
If depth map sensors are used to translate 2D image positions to 3D positions, then localization accuracy is improved, but susceptibility to outdoor light effects and image resolution limitations increases
Solution Approach 1:
The patent captures images from multiple dimensions by taking photographs from various positions, angles, and distances rather than relying on a single depth map. This multi-dimensional approach creates a more robust dataset that is less susceptible to lighting conditions and resolution limitations, as the mapping is derived from multiple 2D views rather than a single 3D depth measurement
Data Source
Figure 1
Figure 2
Figure 3
AI summary
According to an aspect of some embodiments of the present invention there is provided a method for generating a dataset mapping visual features of each of a plurality of objects, comprising: for each of a plurality of different objects: instructing a robotic system to move an arm holding a respective the object to a plurality of positions, and when the arm is in each of the plurality of positions: acquiring at least one image depicting the respective object in the position, receiving positional information of the arm in respective the position, analyzing the at least one image to identify at least one visual feature of the object in the respective position, and storing, in a mapping dataset, an association between the at least one visual feature and the positional information, and outputting the mapping dataset.