Vision-Based Robotic Teleoperation for Intuitive Dexterous Grasping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robotic control systems are non-intuitive and require significant training to perform complex tasks, especially when interacting with objects, due to the lack of intuitive human-machine interfaces.

Innovation Solution

A vision-based teleoperation system that estimates the pose of a human hand using depth cameras and translates it into corresponding robotic hand movements, allowing direct imitation and control of robotic systems without the need for gloves or tactile feedback, utilizing DART for hand tracking, kinematic retargeting, and Riemannian motion policies for real-time control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If direct human control interfaces (joystick or programmatic interface) are used to control robotic systems, then the operator can control the robot to perform tasks, but the interface becomes non-intuitive and difficult to use, requiring significant training and practice

Engineering Contradiction:
Improveease of useVSAvoidtraining time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system captures and copies the operator's hand pose and motion in real-time using depth cameras and pose estimation algorithms, then directly replicates these movements on the robotic system. This eliminates the need for traditional control interfaces by creating a direct visual-motor copy of human actions, making the system immediately intuitive to use without training

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces mechanical control interfaces (joysticks, buttons, programmatic controls) with a vision-based system that uses depth cameras, point cloud processing, and neural network-based pose estimation. This substitution transforms physical control mechanisms into optical field-based control, enabling natural hand gesture recognition and automatic translation to robotic movements

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If traditional control interfaces are used for complex tasks such as interacting with other objects, then the operator can attempt to perform the tasks, but the complexity of the interface makes it difficult to use effectively

Engineering Contradiction:
Improvetask capabilityVSAvoidease of use
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system enables the operator to perform complex manipulation tasks by simply demonstrating the desired hand motions naturally. The vision system automatically tracks hand pose, estimates finger configurations, and translates these into corresponding robotic hand and arm movements, allowing the operator to self-guide the robot through complex tasks without learning complex control procedures

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The vision-based teleoperation system provides a universal control interface that can handle diverse tasks including grasping, manipulation, and interaction with various objects. By capturing full hand pose (position, orientation, and finger configurations) and translating to robotic hand-arm coordination, the system achieves multi-functional capability across different task types through a single intuitive interface

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12399567B2Vision-based teleoperation of dexterous robotic system
Publication Date: 2025.08.26 NVIDIA CORP
  • US12399567B2 patent drawing
  • US12399567B2 patent drawing
  • US12399567B2 patent drawing

AI summary

A human pilot controls a robotic arm and gripper by simulating a set of desired motions with the human hand. In at least one embodiment, one or more images of the pilot's hand are captured and analyzed to determine a set of hand poses. In at least one embodiment, the set of hand poses is translated to a corresponding set of robotic-gripper poses. In at least one embodiment, a set of motions is determined that perform the set of robotic-gripper poses, and the robot is directed to perform the set of motions.