Robot End-Effector Visual Servoing Across Changing Camera Viewpoints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models for robotic manipulation tasks are inaccurate and resource-intensive due to reliance on real-world data and limited viewpoint adaptability, leading to inefficiencies and wear on physical robots.

Innovation Solution

A recurrent neural network model trained using simulated data and reinforced learning is used for visual servoing, incorporating a recurrent layer to maintain action memory and adapt to varying viewpoints, with further adaptation using real-world data to enhance robustness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning models are trained using data from real-world physical robots, then the models can learn accurate manipulation predictions, but the training process is time-consuming, resource-intensive, and causes wear and tear on robots

Engineering Contradiction:
Improveprediction accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates virtual copies of the physical robot and its environment through simulated robots and simulated environments. These virtual replicas allow training data to be generated without using the actual physical robot, eliminating wear and tear while maintaining training effectiveness through realistic simulation physics and sensor models.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary actions by pre-training machine learning models using simulated data before deploying them to physical robots. This preliminary training in virtual environments prepares the models in advance, reducing the need for extensive real-world testing and accelerating the overall deployment timeline.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If machine learning models are trained using data from real-world physical robots, then the models can learn accurate manipulation predictions, but extensive real-world testing is required which consumes large amounts of resources and causes wear and tear

Engineering Contradiction:
Improveprediction accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent replaces energy-intensive physical robot operations with computationally efficient virtual simulations. The simulated robots consume significantly less power while generating equivalent training data, and the patent further optimizes by using pre-generated simulated datasets to reduce the frequency of energy-consuming real-world robot activations.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary training actions in low-power simulated environments before physical deployment. This preliminary action reduces the total energy consumption by minimizing the duration and intensity of real-world robot operation needed for training and testing.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If machine learning models are trained with images captured from the same or similar viewpoint, then the models are adaptable for use in robots capturing images from the same viewpoint, but they become inaccurate and fail when robots capture images from different viewpoints

Engineering Contradiction:
Improveviewpoint adaptabilityVSAvoidprediction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent trains machine learning models using diverse simulated viewpoints that encompass a wide range of camera angles and positions. This universal training approach enables the models to function accurately across multiple viewpoints, making them versatile for different robot configurations and camera placements while maintaining high prediction accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces dynamic viewpoint variations in the simulated training environment, where camera positions and angles change continuously throughout training. This dynamic exposure to varying perspectives teaches the models to adapt to different viewpoints, improving their robustness and accuracy when deployed with robots capturing images from arbitrary angles.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12611768B2Viewpoint invariant visual servoing of robot end effector using recurrent neural network
Publication Date: 2026.04.28 GOOGLE LLC
  • US12611768B2 patent drawing
  • US12611768B2 patent drawing
  • US12611768B2 patent drawing

AI summary

Training and/or using a recurrent neural network model for visual servoing of an end effector of a robot. In visual servoing, the model can be utilized to generate, at each of a plurality of time steps, an action prediction that represents a prediction of how the end effector should be moved to cause the end effector to move toward a target object. The model can be viewpoint invariant in that it can be utilized across a variety of robots having vision components at a variety of viewpoints and/or can be utilized for a single robot even when a viewpoint, of a vision component of the robot, is drastically altered. Moreover, the model can be trained based on a large quantity of simulated data that is based on simulator(s) performing simulated episode(s) in view of the model. One or more portions of the model can be further trained based on a relatively smaller quantity of real training data.