Automated Object Annotation via 3D Pose Propagation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated data capture systems for image frames rely heavily on manual annotation and are prone to errors when propagating object annotations across multiple images, leading to decreased accuracy due to the use of 2D representations alone, which can result in 'drift' and increased errors over a sequence of frames.

Innovation Solution

A data capture system that utilizes 3D representations of scenes, including depth information and camera pose data, to accurately propagate object annotations across images, reducing manual effort and maintaining high accuracy by projecting 3D volumes onto 2D frames, and accounting for occlusions and perspective changes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If 2D image propagation is used to annotate objects across multiple images, then manual annotation effort is reduced, but annotation accuracy deteriorates due to drift and compounding errors

Engineering Contradiction:
Improvemanual annotation effortVSAvoidannotation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent transitions from 2D image propagation to 3D scene representation by introducing depth information and camera pose data. This dimensional upgrade allows the system to maintain accurate spatial relationships across multiple images, eliminating the drift problem inherent in 2D propagation while still reducing manual annotation effort through automated 3D-to-2D projection methods.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If manual annotation is used for each image frame, then annotation accuracy is maintained, but productivity and efficiency decrease due to time-consuming repetitive work

Engineering Contradiction:
Improveannotation accuracyVSAvoidannotation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary 3D scene reconstruction and object localization in a unified coordinate system before generating 2D annotations for each image frame. This preliminary action in 3D space ensures geometric consistency across all views, allowing automated annotation generation that maintains high accuracy while dramatically improving productivity compared to frame-by-frame manual annotation.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If 3D representation with depth information is used, then annotation accuracy across multiple images is improved, but system complexity increases due to additional sensors and processing requirements

Engineering Contradiction:
Improveannotation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a 3D scene representation as an intermediary data structure that mediates between multiple 2D image inputs and the final annotation outputs. This intermediary layer, built using depth information from sensors like LiDAR or stereo cameras, serves as a unified geometric model that simplifies the annotation process across multiple views while improving accuracy, despite the added sensor and processing requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11727593B1Automated data capture
Publication Date: 2023.08.15 GDM HOLDING LLC
  • US11727593B1 patent drawing
  • US11727593B1 patent drawing
  • US11727593B1 patent drawing

AI summary

Methods for annotating objects within image frames are disclosed. Information is obtained that represents a camera pose relative to a scene. The camera pose includes a position and a location of the camera relative to the scene. Data is obtained that represents multiple images, including a first image and a plurality of other images, being captured from different angles by the camera relative to the scene. A 3D pose of the object of interest is identified with respect to the camera pose in at least the first image. A 3D bounding region for the object of interest in the first image is defined, which indicates a volume that includes the object of interest. A location and orientation of the object of interest is determined in the other images based on the defined 3D bounding region of the object of interest and the camera pose in the other images.