Robot Vision Modeling for Unfamiliar Object Detection and Pose
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Robots face challenges in detecting and estimating the pose of unfamiliar objects in their environment, as existing 3D models and vision sensors are insufficient for recognition.
Innovation Solution
A method involving a vision sensor capturing data from multiple angles, generating a machine learning model, and creating rendered images with varying content to train a CNN for object detection and pose estimation, tailored to the robot's environment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a robot uses existing 3D models for object detection, then detection accuracy is improved for known objects, but the robot cannot detect or estimate pose of unfamiliar objects
Solution Approach 1:
The system performs preliminary actions by capturing images of unfamiliar objects from multiple viewpoints before model generation, creating a comprehensive visual dataset that enables subsequent accurate detection and pose estimation of previously unrecognized objects
Solution Approach 2:
The system changes parameters by generating machine learning models with varying degrees of additional content and diversity in training images, optimizing the model's ability to detect unfamiliar objects while maintaining accuracy for known objects
2Measurement precision
If a robot captures vision sensor data from multiple viewpoints, then object recognition accuracy is improved, but time and computational resources are increased
Solution Approach 1:
The system applies partial action by capturing images from a sufficient number of viewpoints rather than all possible angles, achieving adequate recognition accuracy while minimizing time and computational resource expenditure
3Reliability
If a machine learning model is trained with diverse rendered images including additional content, then robustness and environment-specific performance are improved, but model training complexity and time are increased
Solution Approach 1:
The system applies local quality by adding environment-specific additional content to rendered images, tailoring the model training to the robot's specific operational environment and improving robustness for that particular context
Solution Approach 2:
The system performs preliminary action by pre-generating diverse training images with varying additional content before model training, enabling comprehensive model preparation without increasing real-time operational complexity
Data Source
AI summary
Methods and apparatus related to generating a model for an object encountered by a robot in its environment, where the object is one that the robot is unable to recognize utilizing existing models associated with the robot. The model is generated based on vision sensor data that captures the object from multiple vantages and that is captured by a vision sensor associated with the robot, such as a vision sensor coupled to the robot. The model may be provided for use by the robot in detecting the object and/or for use in estimating the pose of the object.


