Transporter Network Robot Actions From Visual Spatial Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models for robotic manipulation require large amounts of data and struggle with unseen classes of objects, occluded objects, and highly deformable objects, imposing data collection burdens and computational overheads, and are not robust due to rigid representational constraints.

Innovation Solution

A computer-implemented method using a transporter network that learns to infer robot actions from visual inputs by leveraging deep feature template matching, preserving spatial structure without object-centric assumptions, enabling efficient learning and generalization to new objects and configurations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If object-centric assumptions (object key points, embeddings, or dense descriptors) are integrated into machine learning approaches for robotic manipulation, then sample efficiency is improved, but data collection burden increases and computational overhead increases

Engineering Contradiction:
Improvesample efficiencyVSAvoiddata collection burden
Core Design Contradiction:
Loss of timeVSQuantity of substance

Solution Approach 1:

The patent extracts and removes object-centric assumptions (object key points, embeddings, dense descriptors) from the machine learning pipeline. By taking out these intermediate representational layers, the system processes raw pixel inputs directly through the transporter network, eliminating the need for separate object detection and description steps, thereby reducing data collection burden while maintaining sample efficiency.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The transporter network is designed as a universal architecture that handles multiple object types and manipulation tasks without requiring object-specific features. The network learns generalizable spatial transformations that apply across diverse objects, eliminating the need for object-centric representations while maintaining broad applicability and sample efficiency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of time

If object-centric assumptions (object key points, embeddings, or dense descriptors) are integrated into machine learning approaches for robotic manipulation, then sample efficiency is improved, but computational burden increases

Engineering Contradiction:
Improvesample efficiencyVSAvoidcomputational burden
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent removes object-centric processing modules (object detectors, embedding generators, descriptor extractors) from the computational pipeline. By extracting these separate computational components, the system reduces overall computational burden while maintaining sample efficiency through the streamlined transporter network that processes pixels directly to actions.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If existing machine learning models map directly from pixels to robotic actions, then complex manipulation skills can be learned, but copious amounts of training data are required

Engineering Contradiction:
Improvecomplex manipulation skillsVSAvoidtraining data
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent introduces the transporter network as an intermediary computational structure that bridges pixel inputs and robotic actions. This intermediary learns spatial transformation patterns that act as a efficient mediator, enabling complex manipulation skill acquisition with reduced training data by capturing essential spatial relationships without requiring exhaustive object-centric annotations.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Loss of time

If object-centric representations are used in machine learning approaches for robotic manipulation, then sample efficiency is improved, but robustness is debilitated by rigidness of representational constraints

Engineering Contradiction:
Improvesample efficiencyVSAvoidrobustness
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent replaces static object-centric representations with dynamic spatial transformations learned by the transporter network. The network adapts its internal representations based on the specific spatial relationships in each input, providing dynamic flexibility that maintains sample efficiency while improving robustness to unseen objects, occlusions, and deformations through learned rather than constrained representations.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20260001219A1Transporter Network for Determining Robot Actions
Publication Date: 2026.01.01 GOOGLE LLC
  • US20260001219A1 patent drawing
  • US20260001219A1 patent drawing
  • US20260001219A1 patent drawing

AI summary

A transporter network for determining robot actions based on sensor feedback can afford robots with efficient autonomous movement. The transporter network may exploit spatial symmetries and may not need assumptions of objectness to provide accurate instructions on object manipulation. The machine-learned model of the transporter network may also allow for learning various tasks with less training examples than other machine-learned models. The machine-learned model of the transporter network may intake observation data as input and may output actions in response to the processed observation data.