Pixelwise Grasp Prediction Without Complete 3D Object Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robotic systems face challenges in generating stable grasp poses for object manipulation due to limitations in available 3D models, especially in realistic scenarios where complete object geometry is not available, leading to inefficiencies in grasp determination.

Innovation Solution

A data processing apparatus and method utilizing a neural network-based grasping model that generates pixelwise predictions from depth image data, aggregating these predictions to select optimal grasp poses for robotic manipulation, incorporating reinforcement learning for improved performance metrics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If geometry-inspired heuristics and geometric analysis are used to select grasp points, then grasp stability and reachability can be improved, but complete 3D models of the object are required which limits applicability in realistic scenarios

Engineering Contradiction:
Improvegrasp stabilityVSAvoidapplicability without complete 3D models
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent uses 2D image data as a copy or representation of the 3D object, allowing the system to perform grasp analysis without requiring complete 3D models. The neural network processes 2D images to predict grasp outcomes, effectively substituting the need for full 3D geometric representations while maintaining grasp reliability

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces traditional geometric analysis methods (mechanical approach) with a neural network-based system (data-driven approach). Instead of performing geometric calculations on 3D models, the system uses trained neural networks to predict grasp outcomes directly from 2D images, substituting mechanical computation with intelligent algorithms

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If multiple pixelwise predictions are aggregated into a single aggregated pixelwise prediction, then robustness and efficiency of grasp pose generation are improved, but the complexity of processing multiple predictions increases

Engineering Contradiction:
Improvegrasp pose generation efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent combines multiple pixelwise predictions from different neural network outcomes into a single aggregated pixelwise prediction. This merging process integrates information from multiple prediction sources to produce a unified grasp probability map, improving robustness while maintaining processing efficiency through consolidated output

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The aggregated pixelwise prediction serves multiple functions: it represents the final grasp probability map, enables selection of optimal grasp points, and provides a unified interface for downstream grasp pose generation. This multi-functionality reduces the need for separate processing steps for each prediction outcome

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240033907A1Pixelwise predictions for grasp generation
Publication Date: 2024.02.01 OCADO INNOVATION LTD
  • US20240033907A1 patent drawing
  • US20240033907A1 patent drawing
  • US20240033907A1 patent drawing

AI summary

Provided are an apparatus and method for grasp generation, involving obtaining image data, comprising depth data, representative of an image of an object captured by a camera, and providing the image data to a grasping model comprising a neural network trained to predict a plurality of outcomes associated with a grasping operation independently based on images of graspable objects. A plurality of pixelwise predictions corresponding to the plurality of outcomes are obtained, each pixelwise prediction being a representation of pixelwise probability values corresponding to a given outcome of the plurality of outcomes associated with the grasping operation. The plurality of pixelwise predictions are aggregated to obtain an aggregated pixelwise prediction which is output for selection therefrom of one or more pixels on which to base generation of one or more grasp poses to grasp the object.