Robot Object Modeling From Multi-View Vision for Unknown Pose Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Robots face challenges in detecting and estimating the pose of unfamiliar objects in their environment, as existing 3D models and vision sensors are insufficient for recognition.

Innovation Solution

A method is implemented where a robot captures vision sensor data from multiple angles to generate a machine learning model, including convolutional neural networks, which creates rendered images of the object at various poses and environments, enabling the robot to detect and estimate the object's pose effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a robot uses existing 3D models for object detection, then detection accuracy is improved for known objects, but the robot cannot recognize unfamiliar objects

Engineering Contradiction:
Improveobject detection accuracyVSAvoidobject recognition capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system creates a copy or replica of the unfamiliar object by generating a 3D model from multiple 2D images captured from different viewpoints. This digital twin can then be used for detection and pose estimation without requiring pre-existing models, enabling the robot to recognize previously unseen objects.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary actions by capturing images from multiple viewpoints and generating a 3D model before the actual detection task. This advance preparation creates a reusable object model that can be stored and applied for future detection operations, improving both accuracy and adaptability.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If a robot captures images from multiple viewpoints to create comprehensive object models, then object recognition capability is improved, but the time and computational resources required increase

Engineering Contradiction:
Improveobject recognition capabilityVSAvoidmodel generation time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system applies partial action by selecting a sufficient but not exhaustive set of viewpoints for image capture. Instead of capturing images from all possible angles, the system identifies and captures images from key viewpoints that provide enough information for accurate 3D model reconstruction, reducing time and resource requirements while maintaining recognition capability.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If a robot generates custom 3D models for each unfamiliar object, then object detection accuracy is improved, but the complexity of the system increases

Engineering Contradiction:
Improveobject detection accuracyVSAvoidmodel generation system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system applies universality by designing a multi-functional pipeline that can handle multiple object types and detection scenarios using the same core technology. The 3D model generation system serves multiple purposes: creating digital twins for detection, enabling pose estimation, and providing a reusable framework for various object recognition tasks, thereby managing complexity through consolidation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP4122657A1Generating a model for an object encountered by a robot
Publication Date: 2023.01.25 GOOGLE LLC
  • EP4122657A1 patent drawingFigure 1
  • EP4122657A1 patent drawingFigure 2A~2D
  • EP4122657A1 patent drawingFigure 3A

AI summary

Methods and apparatus related to generating a model for an object encountered by a robot in its environment, where the object is one that the robot is unable to recognize utilizing existing models associated with the robot. The model is generated based on vision sensor data that captures the object from multiple vantages and that is captured by a vision sensor associated with the robot, such as a vision sensor coupled to the robot. The model may be provided for use by the robot in detecting the object and/or for use in estimating the pose of the object.