Robot Vision Modeling for Unfamiliar Object Detection and Pose

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Robots face challenges in detecting and estimating the pose of unfamiliar objects in their environment, as existing 3D models and vision sensors are insufficient for recognition.

Innovation Solution

A method involving a vision sensor capturing data from multiple angles, generating a machine learning model, and creating rendered images with varying content to train a CNN for object detection and pose estimation, tailored to the robot's environment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a robot uses existing 3D models for object detection, then detection accuracy is improved for known objects, but the robot cannot detect or estimate pose of unfamiliar objects

Engineering Contradiction:
Improveobject detection accuracyVSAvoidability to detect unfamiliar objects
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary actions by capturing images of unfamiliar objects from multiple viewpoints before model generation, creating a comprehensive visual dataset that enables subsequent accurate detection and pose estimation of previously unrecognized objects

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes parameters by generating machine learning models with varying degrees of additional content and diversity in training images, optimizing the model's ability to detect unfamiliar objects while maintaining accuracy for known objects

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If a robot captures vision sensor data from multiple viewpoints, then object recognition accuracy is improved, but time and computational resources are increased

Engineering Contradiction:
Improveobject recognition accuracyVSAvoidtime for data capture and processing
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies partial action by capturing images from a sufficient number of viewpoints rather than all possible angles, achieving adequate recognition accuracy while minimizing time and computational resource expenditure

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If a machine learning model is trained with diverse rendered images including additional content, then robustness and environment-specific performance are improved, but model training complexity and time are increased

Engineering Contradiction:
Improvemodel robustnessVSAvoidmodel training complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system applies local quality by adding environment-specific additional content to rendered images, tailoring the model training to the robot's specific operational environment and improving robustness for that particular context

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system performs preliminary action by pre-generating diverse training images with varying additional content before model training, enabling comprehensive model preparation without increasing real-time operational complexity

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11195041B2Generating a model for an object encountered by a robot
Publication Date: 2021.12.07 GDM HOLDING LLC
  • US11195041B2 patent drawing
  • US11195041B2 patent drawing
  • US11195041B2 patent drawing

AI summary

Methods and apparatus related to generating a model for an object encountered by a robot in its environment, where the object is one that the robot is unable to recognize utilizing existing models associated with the robot. The model is generated based on vision sensor data that captures the object from multiple vantages and that is captured by a vision sensor associated with the robot, such as a vision sensor coupled to the robot. The model may be provided for use by the robot in detecting the object and/or for use in estimating the pose of the object.