Virtual View Transformation for Low-Cost 3D Robot Reasoning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current view-based methods for robotic object manipulation in 3D environments are limited by high computational requirements due to the need for high-resolution voxel representations, which can be resource-intensive and inefficient.

Innovation Solution

The Robot View Transformer (RVT) generates a point-cloud representation using a multi-view transformer, allowing for scalable and accurate 3D object manipulation by transforming input images into a virtual environment with higher resolution and lower processing costs, improving computing speed and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If explicit 3D voxel representations are used for 3D reasoning, then measurement precision is improved, but device complexity and computational resources increase significantly

Engineering Contradiction:
Improve3D reasoning accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a virtual copy of the physical environment with 3D geometric representations of objects, surfaces, and spatial relationships. This virtual model allows the robot to perform 3D reasoning and simulation without requiring high-resolution voxel representations of the entire scene, thereby reducing computational resources while maintaining measurement precision for manipulation tasks.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent segments the 3D environment into discrete objects with geometric primitives (planes, cylinders, spheres) rather than using a dense voxel grid. This segmentation approach represents only the essential geometric features needed for manipulation, significantly reducing the number of elements from cubic scaling to a manageable set of parameterized shapes.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If high-resolution voxelization is used for neural network input, then manufacturing precision is improved, but loss of energy and computational cost increase

Engineering Contradiction:
Improveobject manipulation precisionVSAvoidcomputational energy
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent changes the parameterization approach from dense voxel grids to parameterized geometric primitives with a small number of degrees of freedom. Each object is represented by a limited set of parameters (position, orientation, dimensions of geometric primitives) rather than millions of voxel values, maintaining manipulation precision while dramatically reducing energy consumption for neural network processing.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If view-based methods are used for object manipulation, then device complexity is reduced, but measurement precision for full 3D reasoning deteriorates

Engineering Contradiction:
Improvesystem simplicityVSAvoid3D reasoning capability
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces a virtual environment as an intermediary between the simple view-based input and the required 3D reasoning capability. The virtual model serves as a mediator that transforms 2D image inputs into a structured 3D geometric representation, enabling full 3D reasoning without directly increasing the complexity of the perception system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240273810A1View transformation for machine-learned three-dimensional reasoning
Publication Date: 2024.08.15 NVIDIA CORP
  • US20240273810A1 patent drawing
  • US20240273810A1 patent drawing
  • US20240273810A1 patent drawing

AI summary

In various examples, a machine may generate, using sensor data capturing one or more views of an environment, a virtual environment including a 3D representation of the environment. The machine may render, using one or more virtual sensors in the virtual environment, one or more images of the 3D representation of the environment. The machine may apply the one or more images to one or more machine learning models (MLMs) trained to generate one or more predictions corresponding to the environment. The machine may perform one or more control operations based at least on the one or more predictions generated using the one or more MLMs.