Robot Motion Prediction Using Action-Conditioned Pixel Flow

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Robots face challenges in predicting how their actions affect objects in their environment without relying on manually labeled object information, which becomes impractical for scaling real-world interaction learning across various scenes and objects.

Innovation Solution

An action-conditioned motion prediction model that explicitly models pixel motion by predicting a distribution over pixel motion from previous frames, allowing robots to 'visually imagine' different futures based on candidate movements without reconstructing the entire image, using datasets like 50,000 robot interactions for training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If robots rely on manually labeled object information to predict object motion, then prediction accuracy for known objects is improved, but scalability across various scenes and objects deteriorates

Engineering Contradiction:
Improveprediction accuracyVSAvoidscalability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent extracts and focuses specifically on pixel motion information from images, separating it from complete object recognition and labeling. By predicting pixel motion distributions directly from image data and robot actions without requiring manually labeled object information, the system achieves both accurate motion prediction and scalability to unseen objects and environments.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the traditional mechanical approach of manual object labeling and classification with a data-driven pixel motion prediction model. Instead of relying on pre-labeled object databases, the system uses neural networks to learn pixel motion patterns directly from image sequences, enabling automatic generalization to new objects and scenes without manual intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If the motion prediction model focuses on reconstructing the entire image, then comprehensive environmental prediction is improved, but computational complexity and processing time deteriorate

Engineering Contradiction:
Improveenvironmental prediction completenessVSAvoidcomputational complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts and predicts only the motion component of environmental changes rather than reconstructing entire future images. By focusing specifically on pixel motion distributions and applying them to the current image to generate predicted images, the system reduces computational complexity while preserving essential motion information for decision-making.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the image prediction task into separate components: current image input, robot action input, pixel motion distribution prediction, and predicted image generation. This segmentation allows the model to specialize in predicting motion transformations rather than generating complete images from scratch, reducing computational burden while maintaining predictive accuracy.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If the model predicts pixel motion distributions from previous frames, then generalization to unseen objects is improved, but prediction precision for specific object details may deteriorate

Engineering Contradiction:
Improvegeneralization capabilityVSAvoidobject detail precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent replaces object-centric recognition systems with pixel-centric motion prediction. Instead of identifying and tracking specific objects based on their features and labels, the system predicts pixel motion distributions directly from image data, enabling natural generalization to unseen objects while maintaining precision through data-driven learning of motion patterns.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11173599B2Machine learning methods and apparatus related to predicting motion(s) of object(s) in a robot's environment based on image(s) capturing the object(s) and based on parameter(s) for future robot movement in the environment
Publication Date: 2021.11.16 GOOGLE LLC
  • US11173599B2 patent drawing
  • US11173599B2 patent drawing
  • US11173599B2 patent drawing

AI summary

Some implementations of this specification are directed generally to deep machine learning methods and apparatus related to predicting motion(s) (if any) that will occur to object(s) in an environment of a robot in response to particular movement of the robot in the environment. Some implementations are directed to training a deep neural network model to predict at least one transformation (if any), of an image of a robot's environment, that will occur as a result of implementing at least a portion of a particular movement of the robot in the environment. The trained deep neural network model may predict the transformation based on input that includes the image and a group of robot movement parameters that define the portion of the particular movement.