Robotic Policy Control Using Viewpoint-Invariant Image Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning systems face challenges in optimizing policy controllers for robotic agents to perform tasks effectively, especially when dealing with variations in viewpoint, occlusions, and other transformations, requiring extensive labeled data and explicit joint-level correspondence.

Innovation Solution

A system utilizing a time-contrastive neural network trained to generate invariant numeric embeddings, allowing the optimization of policy controllers using only raw video demonstrations without explicit joint-level correspondence, enabling robotic agents to perform tasks from different viewpoints and improve imitation capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If reinforcement learning systems use traditional policy controllers with explicit joint-level correspondence and labeled data, then they can achieve accurate task execution, but the system complexity and data requirements increase significantly

Engineering Contradiction:
Improvetask execution accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical control approaches (explicit joint-level correspondence, labeled data requirements) with a neural network-based system that learns from raw video demonstrations. The time-contrastive neural network substitutes complex labeled data processing with unsupervised learning from raw video, reducing system complexity while maintaining task execution accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system uses video demonstrations as copies of expert behavior to train the policy controller. Instead of requiring explicit labeled data about joint movements, the system copies motion patterns directly from video demonstrations, simplifying the data acquisition process while preserving task execution accuracy.

Inventive Principle:
Principle #26Copying

2Measurement precision

If the policy controller is trained with viewpoint-specific data, then it achieves high performance for that specific viewpoint, but it fails to generalize to different viewpoints and transformations

Engineering Contradiction:
Improveperformance accuracyVSAvoidviewpoint generalization
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The time-contrastive neural network is designed to be viewpoint-invariant, allowing it to process images from any viewpoint and generate consistent embeddings. This universal approach enables the policy controller to generalize across different viewpoints, camera angles, and transformations while maintaining performance accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the parameter space by using learned embeddings that are invariant to viewpoint transformations. Instead of training separate controllers for each viewpoint, the system transforms the input space into a viewpoint-invariant embedding space where the same task state maps to similar representations regardless of viewpoint.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If the system uses raw video demonstrations without labeled data, then the data preparation process is simplified, but the optimization of policy controllers becomes more challenging

Engineering Contradiction:
Improvedata preparation easeVSAvoidoptimization complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The time-contrastive neural network serves as an intermediary that transforms raw video demonstrations into structured embeddings suitable for policy controller optimization. This intermediary component automatically extracts relevant features and temporal relationships from raw video, simplifying data preparation while providing structured input for optimization.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary processing by training the time-contrastive neural network on raw video demonstrations before optimizing the policy controller. This preliminary action creates viewpoint-invariant embeddings that facilitate subsequent policy optimization, reducing the overall optimization complexity.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If the neural network is trained to be invariant to transformations, then it improves robustness to viewpoint changes and occlusions, but the training process requires more complex loss functions and optimization

Engineering Contradiction:
Improverobustness to transformationsVSAvoidtraining complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The time-contrastive loss function provides feedback by comparing embeddings of the same task state from different viewpoints and temporal positions. This feedback mechanism guides the neural network to learn viewpoint-invariant representations, improving robustness while managing training complexity through a well-defined loss function.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system changes the training objective by using a contrastive loss function that operates in the embedding space rather than directly on raw pixels. This parameter transformation simplifies the optimization process for learning invariance compared to traditional approaches that operate directly on image data.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11559887B2Optimizing policy controllers for robotic agents using image embeddings
Publication Date: 2023.01.24 GOOGLE LLC
  • US11559887B2 patent drawing
  • US11559887B2 patent drawing
  • US11559887B2 patent drawing

AI summary

There are provided systems, methods, and apparatus, for optimizing a policy controller to control a robotic agent that interacts with an environment to perform a robotic task. One of the methods includes optimizing the policy controller using a neural network that generates numeric embeddings of images of the environment and a demonstration sequence of demonstration images of another agent performing a version of the robotic task.