Robot Pickup Pose Estimation From Single-Camera Descriptor Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Robots face challenges in picking up objects in various positions within their working space due to the need for flexible recognition and orientation of objects, which existing technologies struggle to address effectively without requiring multiple camera perspectives or extensive training data.

Innovation Solution

A method using a machine learning model trained with supervised learning to map camera images onto descriptor images, allowing robots to identify reference points and ascertain pickup poses in three-dimensional space from a single camera image, enabling secure object grasping regardless of position, with flexible adaptation to different objects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple camera perspectives are used to recognize objects in various positions, then the recognition accuracy and pickup pose determination are improved, but the device complexity and cost increase

Engineering Contradiction:
Improveobject recognition accuracyVSAvoidcamera system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a descriptor image that serves as a simplified copy or representation of the object's visual features. Instead of using multiple physical cameras, a single camera image is transformed into a descriptor image containing encoded object characteristics that enable pose recognition from any viewpoint, eliminating the need for complex multi-camera setups

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the camera image into a descriptor image by changing the representation parameters. The descriptor image encodes object features in a different parameter space that is invariant to viewpoint changes, allowing the system to recognize objects from various positions using a single camera rather than multiple cameras

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If extensive training data is used to train the machine learning model, then the model's ability to recognize objects in various positions is improved, but the training time and computational resources increase

Engineering Contradiction:
Improveobject recognition versatilityVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-computing and storing descriptor images for training objects during an offline phase. These descriptor images are created in advance and can be directly compared with images from the single camera, eliminating the need for extensive real-time training and reducing both training time and computational requirements during operation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses descriptor images as simplified copies of objects for training purposes. Instead of training the model with extensive raw image data, the system trains by comparing descriptor images that already contain extracted features, significantly reducing the amount of training data and computational resources needed while maintaining recognition versatility

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11964400B2Device and method for controlling a robot to pick up an object in various positions
Publication Date: 2024.04.23 ROBERT BOSCH GMBH
  • US11964400B2 patent drawing
  • US11964400B2 patent drawing
  • US11964400B2 patent drawing

AI summary

A method for controlling a robot to pick up an object in various positions. The method includes: defining a plurality of reference points on the object; mapping a first camera image of the object in a known position onto a first descriptor image; identifying the descriptors of the reference points from the first descriptor image; mapping a second camera image of the object in an unknown position onto a second descriptor image; searching the identified descriptors of the reference points in the second descriptor image; ascertaining the positions of the reference points in the three-dimensional space in the unknown position from the found positions; and ascertaining a pickup pose of the object for the unknown position from the ascertained positions of the reference points.