Eye-on-Hand RL Grasping With Active Pose Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robotic systems face challenges in dynamically grasping moving objects in unstructured environments due to decoupled perception and manipulation systems, occlusions, and the need for adaptive grasp planning in dynamic scenarios, particularly when using external cameras.

Innovation Solution

A system with a wrist-mounted camera and coupled perception and manipulation subsystems employs a curriculum-trained model-free reinforcement learning policy for active pose tracking and grasp synthesis, enabling full 6-DoF dynamic grasping of novel objects without prior knowledge of their motion profiles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a fixed workspace camera is used for vision-based manipulation, then the perception system can maintain a stable viewpoint, but the system requires large clearances above the workspace and cannot handle occlusions or confined spaces

Engineering Contradiction:
Improvestable viewpointVSAvoidworkspace flexibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent merges the perception system (camera) with the manipulation system (robotic arm) by mounting the camera on the wrist of the manipulator. This coupling allows the camera to move with the arm, eliminating the need for large clearances and enabling operation in confined spaces while maintaining a stable viewpoint relative to the object being manipulated.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from a static fixed camera to a dynamic wrist-mounted camera that moves with the robotic arm. This dynamic positioning allows the camera to adapt to different workspace configurations and maintain optimal viewing angles throughout the manipulation task, enabling operation in previously inaccessible confined spaces.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If the perception system is decoupled from the manipulation system, then the camera can be positioned for optimal viewing, but the system experiences occlusions and loss of tracking when the target object is moving

Engineering Contradiction:
Improveviewing qualityVSAvoidtracking continuity
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

By merging the perception and manipulation systems through wrist-mounted camera attachment, the patent ensures that the camera maintains a consistent relative position to the object being manipulated. This coupling eliminates occlusions and tracking loss that occur with decoupled systems, as the camera moves synchronously with the manipulator.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements active pose estimation using the wrist-mounted camera to continuously track the object's position and orientation. This feedback mechanism allows the system to adapt to object motion in real-time, maintaining reliable tracking even during dynamic manipulation tasks.

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If model-based approaches are used for vision-based grasping, then the system can leverage known object information, but the system cannot generalize to novel objects without prior models

Engineering Contradiction:
Improvegrasp accuracyVSAvoidobject generalization
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent employs model-free reinforcement learning that allows the system to learn grasping strategies directly from interaction data without requiring pre-built object models. The system serves itself by learning from experience, enabling generalization to novel objects while maintaining high grasp accuracy through learned policies.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transitions from model-based approaches that require specific object parameters to model-free reinforcement learning that learns from raw sensory inputs. This parameter-free approach enables the system to handle diverse and novel objects by learning invariant features directly from data.

Inventive Principle:
Principle #35Parameter changes

4Ease of manufacture

If reinforcement learning methods use discrete action sets, then the training data requirements are reduced, but the task scope is limited to constrained state-action spaces

Engineering Contradiction:
Improvetraining efficiencyVSAvoidtask scope
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent implements continuous action spaces for the reinforcement learning policy, allowing the robotic arm to execute smooth, dynamic movements rather than discrete jumps. This enables the system to handle complex manipulation tasks with full 6-DoF freedom while maintaining training efficiency through modern continuous control algorithms.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent develops a universal reinforcement learning policy that works across diverse manipulation tasks and object types. The continuous action space and model-free approach create a versatile system that can adapt to various task scopes without requiring task-specific discretization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12583111B2Eye-on-hand reinforcement learner for dynamic grasping with active pose estimation
Publication Date: 2026.03.24 MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC
  • US12583111B2 patent drawing
  • US12583111B2 patent drawing
  • US12583111B2 patent drawing

AI summary

A controller is provided for performing dynamic grasping of a target object using visual sensory inputs. The controller includes a robotic interface connected to a robotic arm including links connected by joints having actuators and encoders, and a gripper of the end-effector of the robotic arm configured to grasp the target object in response to robot control signals, and a vision sensor configured to continuously provide visual observations for tracking poses of the target object in a workspace and compute grasp poses, wherein the vision sensor is mounted on a distal end of the robotic arm adjacent to the gripper. The controller trains the Eye-on-Hand reinforcement learner policy, tracks the poses of the target object, and generates robot control signals to follow the target object while keeping it in the field of view of the vision sensor and grasp the target object in the workspace.