Neural Network View Generation for Articulable Objects

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating image data that represent objects in different poses and viewpoints are limited, as they often require extensive labeled training data and can only modify poses or viewpoints but not both effectively, and are impractical for capturing additional image or video data.

Innovation Solution

A neural network is trained using unlabeled video data to generate image data for various articulations and viewpoints, capable of identifying and rendering articulable objects' parts and joints without annotation, by encoding image features into a latent space and using an implicit function to predict color and transparency for 3D points, allowing for volumetric rendering of objects in new poses and views.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional methods are used to generate image data for objects in different poses and viewpoints, then the quality of generated images can be maintained, but extensive labeled training data is required which increases data collection time and complexity

Engineering Contradiction:
Improveamount of labeled training dataVSAvoidtime to obtain labeled training data
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system uses unlabeled video data where the neural network automatically learns to identify and track articulable objects and their parts without requiring human annotation. The network self-trains by detecting joints, parts, and poses directly from raw video frames, eliminating the need for manual labeling while maintaining training effectiveness

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an implicit function as an intermediary component that bridges the gap between raw video data and the desired 3D object representations. This implicit function learns to map 2D image features to 3D spatial relationships and articulation information, enabling the system to process unlabeled data effectively

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If neural networks are trained to generate images for both pose modification and viewpoint change, then the versatility of the system is improved, but the complexity of training and model architecture increases

Engineering Contradiction:
Improveability to modify poses and viewpointsVSAvoidtraining complexity and model architecture
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The neural network is designed as a universal model that simultaneously handles multiple tasks: identifying articulable objects, detecting joints and parts, estimating 3D pose, and generating images for both pose modification and viewpoint change. This multi-functional approach consolidates what would traditionally require separate specialized models into a single integrated system

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system transitions from 2D image processing to 3D spatial representation by learning implicit 3D models of articulable objects. The network encodes 2D video frames into 3D articulated representations, enabling it to generate images for arbitrary poses and viewpoints by manipulating the 3D structure rather than processing multiple 2D images separately

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20230137403A1View generation using one or more neural networks
Publication Date: 2023.05.04 NVIDIA CORP
  • US20230137403A1 patent drawing
  • US20230137403A1 patent drawing
  • US20230137403A1 patent drawing

AI summary

Apparatuses, systems, and techniques are presented to generate one or more images. In at least one embodiment, one or more neural networks are used to generate one or more images of one or more objects in two or more different poses from two or more different points of view.