Wrist-Mounted Vision for Robotic Grasp Pose Versatility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robotic control systems are limited in their ability to execute a wide variety of grasp poses due to restricted camera views, typically only allowing top-down grasps, which restricts the number of possible grasp orientations and positions a robotic gripper can achieve.

Innovation Solution

A system employing a wrist-mounted camera oriented towards the gripper, utilizing deep reinforcement learning and a double deep Q-network to learn a mapping from images to estimated Q-values, allowing the robotic gripper to grasp objects from various angles and orientations by training in simulation and transferring to real-world scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a top-down camera view is used, then the system is simple to implement, but the number of possible grasp orientations and positions is limited

Engineering Contradiction:
Improvenumber of grasp posesVSAvoidcamera positioning complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Instead of using a traditional top-down overhead camera view, the patent inverts the camera positioning to be wrist-mounted on the robotic arm, orienting the camera towards the gripper. This inversion allows the camera to capture images from the gripper's perspective, enabling the system to determine grasp poses for objects viewed from side, angled, and other non-top-down orientations, thereby significantly increasing the number of achievable grasp poses.

Inventive Principle:
Principle #13The other way round (Inversion)

2Adaptability or versatility

If a wrist-mounted camera is used, then more grasp poses are achievable, but the system complexity increases

Engineering Contradiction:
Improvegrasp orientation capabilityVSAvoidcamera system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The wrist-mounted camera serves multiple functions: it captures images of objects from the gripper's perspective for pose estimation, provides visual feedback for grasp execution, and enables the system to handle various object types and orientations. This multi-functionality justifies the increased device complexity by delivering comprehensive grasp capabilities across diverse scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If deep reinforcement learning is used for grasp planning, then grasp versatility improves, but computational resources required increase

Engineering Contradiction:
Improvegrasp planning capabilityVSAvoidcomputational energy consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary action by training the deep reinforcement learning model offline using simulation data before actual grasp execution. During real-world operation, the pre-trained model quickly processes images and determines grasp poses without requiring extensive real-time computation. This separates the computationally intensive training phase from the efficient inference phase, reducing real-time energy consumption while maintaining high grasp versatility.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11833681B2Robotic control system
Publication Date: 2023.12.05 NVIDIA CORP
  • US11833681B2 patent drawing
  • US11833681B2 patent drawing
  • US11833681B2 patent drawing

AI summary

In at least one embodiment, under the control of a robotic control system, a gripper on a robot is positioned to grasp a 3-dimensional object. In at least one embodiment, the relative position of the object and the gripper is determined, at least in part, by using a camera mounted on the gripper.