Category-Level 6-DoF Pose Tracking Using RGB Sequences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for object detection and pose estimation in real-world applications, such as autonomous driving, often rely on detailed 3D sensor data or specific mathematical models, which may not be available, making accurate detection and location difficult, especially when only category-level models are used.

Innovation Solution

A system and method for category-level 6-DoF pose estimation using a sequence of RGB camera images without depth information, employing a tracklet-conditioned network and filtering process to predict the pose of objects within a category, such as 'mugs' or 'shoes', by leveraging probabilistic filtering and neural networks to estimate keypoint distributions and refine object dimensions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If detailed 3D sensor data and mathematical models are used for object detection, then detection accuracy is improved, but device complexity and data requirements increase

Engineering Contradiction:
Improvedetection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces expensive, complex 3D sensor systems with inexpensive 2D RGB cameras. Instead of using detailed mathematical models of individual objects, the system uses category-level models that are computationally efficient and do not require precise 3D information. This substitution of simpler, cheaper sensing and modeling approaches maintains adequate detection accuracy while reducing system complexity and hardware costs.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Measurement precision

If instance-specific 3D models are used for pose estimation, then pose accuracy is improved, but adaptability to different objects decreases

Engineering Contradiction:
Improvepose accuracyVSAvoidgeneralization capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent employs category-level models that can be applied to any object within a category (e.g., any mug, any shoe) rather than requiring instance-specific models. The system processes sequences of RGB images and uses probabilistic filtering to estimate pose parameters for objects of the same category, making the system universally applicable across multiple instances while maintaining reasonable pose accuracy through temporal consistency and statistical modeling.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If detailed mathematical models are used for each object, then detection precision is improved, but loss of time for processing increases

Engineering Contradiction:
Improvedetection precisionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent pre-establishes category-level models and statistical parameters for object categories before actual detection occurs. During runtime, the system processes sequences of images using these pre-computed models and applies efficient probabilistic filtering algorithms. This preliminary preparation eliminates the need for complex real-time computations for each individual object, reducing processing time while maintaining detection precision through the use of pre-characterized category properties.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240005547A1Object pose tracking from video images
Publication Date: 2024.01.04 NVIDIA CORP
  • US20240005547A1 patent drawing
  • US20240005547A1 patent drawing
  • US20240005547A1 patent drawing

AI summary

Apparatuses, systems, and techniques to determined a pose of an object from a plurality of images. In at least one embodiment, the pose of an object is determined from at least two images of a video sequence using one or more neural networks, in which the neural network produces a distribution of pose information that is filtered to determine the current pose.