Robot Grasp Prediction With Geometry-Aware 3D Encodings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current robotic systems face challenges in accurately predicting the success of grasp outcomes for end effectors, particularly due to limitations in representing and utilizing 3D geometry features in grasp pose evaluation.

Innovation Solution

The development of a deep machine learning method involving a geometry network and a grasp outcome prediction network, trained using user-guided grasp attempts in virtual reality, generates geometry-aware encodings that improve grasp outcome prediction accuracy by representing 3D features effectively, and automatically generates additional training instances to enhance neural network robustness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional 2D image-based methods are used for grasp detection, then the system complexity is low, but the grasp outcome prediction accuracy is insufficient due to lack of 3D geometry representation

Engineering Contradiction:
Improvegrasp outcome prediction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transitions from 2D image-based grasp detection to 3D geometry-aware representation by projecting 3D object models onto 2D images and back, enabling accurate depth and shape understanding while maintaining compatibility with standard vision systems. This dimensional enhancement resolves the contradiction by providing 3D information without completely abandoning 2D processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces an intermediate 3D object model as a mediator between the 2D camera input and the grasp prediction system. This intermediate representation contains geometric information that bridges the gap between simple 2D imaging and complex 3D analysis, improving accuracy without requiring direct complex 3D sensing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If extensive user-guided grasp attempts are collected for training, then the neural network robustness improves, but the training time and resource requirements increase significantly

Engineering Contradiction:
Improveneural network robustnessVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses virtual reality environments to create synthetic copies of physical grasp scenarios. Instead of collecting extensive real-world user-guided grasp attempts, the system generates training data by rendering virtual objects and simulating grasp interactions, achieving robust training with significantly reduced time and resource requirements.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary synthesis of training data in virtual reality before actual deployment. By pre-generating diverse grasp scenarios, object variations, and outcomes in the virtual environment, the system prepares comprehensive training sets in advance, reducing the need for extensive real-world data collection and accelerating the training process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3693138B1Robotic grasping prediction using neural networks and geometry aware object representation
Publication Date: 2022.08.03 GOOGLE LLC
  • EP3693138B1 patent drawingFigure 1
  • EP3693138B1 patent drawingFigure 2
  • EP3693138B1 patent drawingFigure 3

AI summary

Deep machine learning methods and apparatus, some of which are related to determining a grasp outcome prediction for a candidate grasp pose of an end effector of a robot. Some implementations are directed to training and utilization of both a geometry network and a grasp outcome prediction network. The trained geometry network can be utilized to generate, based on two-dimensional or two-and-a-half-dimensional image(s), geometry output(s) that are: geometry-aware, and that represent (e.g., high-dimensionally) three-dimensional features captured by the image(s). In some implementations, the geometry output(s) include at least an encoding that is generated based on a trained encoding neural network trained to generate encodings that represent three-dimensional features (e.g., shape). The trained grasp outcome prediction network can be utilized to generate, based on applying the geometry output(s) and additional data as input(s) to the network, a grasp outcome prediction for a candidate grasp pose.