Robot Grasp Depth Estimation Using Self-Supervised End Effector Labels

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Accurate depth estimation for robotic manipulation remains a challenge, especially on reflective or transparent surfaces, where existing techniques like structured light sensors fail.

Innovation Solution

A self-supervised neural network is trained to predict the position of a robot's end effector if a grasp were attempted at each pixel in an input image, using data from physical interactions in a pick-and-place environment, without human annotation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If structured light sensors are used for depth estimation, then depth information can be obtained, but they fail on reflective or transparent surfaces

Engineering Contradiction:
Improvedepth estimation reliabilityVSAvoidsurface type adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system uses the robot's own grasp attempts and forward kinematics to generate training labels, making the depth estimation system self-supervised and adaptable to various surfaces without requiring specialized sensors for each surface type

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces structured light sensing with a neural network-based depth estimation system that learns from visual data and robot interaction, eliminating the physical limitations of optical sensors on reflective or transparent surfaces

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If physical interaction data is collected for training, then accurate depth estimation is achieved, but data collection becomes expensive

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidtraining data cost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system generates its own training labels through forward kinematics calculations from recorded robot grasp attempts, eliminating the need for expensive manual annotation while achieving accurate depth estimation

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses the robot's own grasp outcomes and end effector positions as feedback to automatically generate training labels, creating a self-supervised learning loop that reduces external data collection costs

Inventive Principle:
Principle #23Feedback

3Ease of manufacture

If traditional depth estimation methods are used, then they work on standard surfaces, but they achieve higher root mean squared error

Engineering Contradiction:
Improvemethod simplicityVSAvoiddepth estimation accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent changes from traditional sensor-based depth measurement to a learned representation where depth is predicted through a neural network that processes visual features, achieving higher accuracy by learning complex surface properties

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12236340B2Computer-automated robot grasp depth estimation
Publication Date: 2025.02.25 OSARO
  • US12236340B2 patent drawing
  • US12236340B2 patent drawing
  • US12236340B2 patent drawing

AI summary

A computer system trains a neural network to predict, for each pixel in an input image, the position that a robot's end effector would reach if a grasp (“poke”) were attempted at that position. Training data consists of images and end effector positions recorded while a robot attempts grasps in a pick-and-place environment. For an automated grasping policy, the approach is self-supervised, as end effector position labels may be recovered through forward kinematics, without human annotation. Although gathering such physical interaction data is expensive, it is necessary for training and routine operation of state of the art manipulation systems. Therefore, the system comes “for free” while collecting data for other tasks (e.g., grasping, pushing, placing). The system achieves significantly lower root mean squared error than traditional structured light sensors and other self-supervised deep learning methods on difficult, industry-scale jumbled bin datasets.