3D Mesh Annotation System for Reducing Dataset Labeling Hours

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for creating labeled datasets for machine learning and computer vision systems require significant human-hours for image capture and manual labeling, which is inefficient and labor-intensive.

Innovation Solution

A system that allows users to mix static scene and live annotations by capturing a 3D mesh of a scene and marking annotations in either online (live view) or offline (static view) modes, enabling seamless switching between modes and collaborative annotation across multiple devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling of individual images is performed, then labeled datasets can be created, but the process requires significant human-hours and is labor-intensive

Engineering Contradiction:
Improvelabeling accuracyVSAvoidhuman-hours required
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent combines multiple images captured from different viewpoints into a single 3D mesh representation. By merging the spatial information from multiple 2D images into a unified 3D model, annotators can label objects in three-dimensional space once, and the annotations are automatically projected back to all corresponding 2D images, eliminating redundant labeling work while maintaining labeling accuracy across all views

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from 2D image annotation to 3D mesh annotation by creating a three-dimensional representation of the scene. Annotations made on the 3D mesh surface are then projected onto multiple 2D images, allowing a single 3D annotation to generate multiple 2D annotations automatically. This dimensionality change reduces the total number of annotation operations required while preserving measurement precision across all views

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If multiple images are captured from different poses and lighting conditions, then comprehensive training data can be collected, but the number of images in the database becomes large

Engineering Contradiction:
Improvedataset coverageVSAvoidnumber of images
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent creates a virtual 3D copy of the physical scene and objects, which can be annotated and then projected to generate multiple 2D image representations. This virtual copy serves as a master model from which infinitely many 2D projections can be generated without capturing additional physical images, thereby maintaining dataset coverage while reducing the actual number of image files stored

Inventive Principle:
Principle #26Copying

3Reliability

If human technicians are deployed to capture images in the field, then diverse training data can be obtained, but the process is intensive and requires significant resources

Engineering Contradiction:
Improvedata qualityVSAvoiddata collection efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary 3D mesh construction and annotation in a controlled environment before generating the final 2D image dataset. By completing the complex 3D modeling and annotation work upfront, the system can then automatically generate numerous 2D projections without requiring additional field deployment or manual intervention, thereby maintaining data quality while dramatically improving productivity

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12223595B2Method and system for mixing static scene and live annotations for efficient labeled image dataset collection
Publication Date: 2025.02.11 GENESEE VALLEY INNOVATIONS LLC
  • US12223595B2 patent drawing
  • US12223595B2 patent drawing
  • US12223595B2 patent drawing

AI summary

A system is provided which mixes static scene and live annotations for labeled dataset collection. A first recording device obtains a 3D mesh of a scene with physical objects. The first recording device marks, while in a first mode, first annotations for a physical object displayed in the 3D mesh. The system switches to a second mode. The system displays, on the first recording device while in the second mode, the 3D mesh including a first projection indicating a 2D bounding area corresponding to the marked first annotations. The first recording device marks, while in the second mode, second annotations for the physical object or another physical object displayed in the 3D mesh. The system switches to the first mode. The first recording device displays, while in the first mode, the 3D mesh including a second projection indicating a 2D bounding area corresponding to the marked second annotations.