Recurrent Visual Servoing for Robot End Effectors Across Viewpoints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning-based approaches for robotic grasping and manipulation are limited by their reliance on training data from specific viewpoints, leading to inaccuracies when faced with variations in viewpoint and requiring extensive real-world testing, which is time-consuming, resource-intensive, and causes wear and tear on robots.

Innovation Solution

A recurrent neural network model is trained using simulated data to generate action predictions for robotic end effectors, incorporating long short-term memory units and gated recurrent units to adapt to various viewpoints, and further adapted with real-world data for improved robustness and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If machine learning models are trained using data from real-world physical robots, then the models can learn from actual robotic grasping scenarios, but this approach is time-consuming, resource-intensive, and causes wear and tear on the robots

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates virtual copies of the physical robot and its environment through simulation. A digital twin of the robot system is built where training data is generated in the virtual environment rather than requiring extensive physical robot operation. This copying approach allows unlimited training iterations without wear and tear on physical hardware while maintaining realistic robotic dynamics and sensor models.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary actions by pre-simulating numerous robotic grasping scenarios and storing the results as training data before actual robot operation. The simulation environment pre-generates diverse training examples including various object types, lighting conditions, and robot configurations, which are then used to train the machine learning model before deployment on physical robots.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If machine learning models are trained with input images captured from the same or similar viewpoint, then the models are adaptable for use in robots capturing images from the same viewpoint, but they become inaccurate and fail when robots capture images from different viewpoints

Engineering Contradiction:
Improveviewpoint adaptabilityVSAvoidgrasping prediction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent trains the machine learning model with images captured from multiple different viewpoints simultaneously, making the model universal across various camera positions and orientations. The simulation environment generates training data from numerous viewpoint configurations, enabling the model to handle images from any viewpoint rather than being specialized for a single perspective.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent adds the viewpoint dimension to the training data by systematically varying camera positions, angles, and orientations in the simulation environment. Instead of training on 2D image data from a fixed viewpoint, the model receives training examples that span a multi-dimensional viewpoint space, enabling it to generalize to unseen viewpoints during deployment.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11701773B2Viewpoint invariant visual servoing of robot end effector using recurrent neural network
Publication Date: 2023.07.18 GOOGLE LLC
  • US11701773B2 patent drawing
  • US11701773B2 patent drawing
  • US11701773B2 patent drawing

AI summary

Training and/or using a recurrent neural network model for visual servoing of an end effector of a robot. In visual servoing, the model can be utilized to generate, at each of a plurality of time steps, an action prediction that represents a prediction of how the end effector should be moved to cause the end effector to move toward a target object. The model can be viewpoint invariant in that it can be utilized across a variety of robots having vision components at a variety of viewpoints and/or can be utilized for a single robot even when a viewpoint, of a vision component of the robot, is drastically altered. Moreover, the model can be trained based on a large quantity of simulated data that is based on simulator(s) performing simulated episode(s) in view of the model. One or more portions of the model can be further trained based on a relatively smaller quantity of real training data.