Pose Estimator Training Data Capture System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for training machine learning models for computer vision tasks face challenges in generating sufficient real-world data, particularly for tasks like object pose detection, which requires nuanced and complex data that synthetic data often fails to capture.

Innovation Solution

A data capture stage comprising a frame with a rotation device, multiple cameras, sensors, and light sources, controlled by an augmentation data generator to capture images and mapping data from various angles and lighting conditions, generating labeled training data for machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If synthetic data is used for training machine learning models, then data generation speed and cost are improved, but data quality and realism are worsened

Engineering Contradiction:
Improvedata generation speedVSAvoiddata quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent uses physical objects as templates to create accurate copies through photography and 3D scanning. Real objects are captured from multiple angles and viewpoints, creating photorealistic training data that maintains the nuanced details and complexities of real-world scenes while enabling automated data generation at scale.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary actions by physically setting up objects in controlled scenes before data capture. Objects are arranged in predetermined configurations, and multiple viewpoints are pre-planned and captured systematically. This preliminary physical setup enables automated generation of diverse training data without requiring manual scene reconstruction for each training sample.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If manual data collection and labeling is performed, then data quality and accuracy are improved, but time consumption and cost are worsened

Engineering Contradiction:
Improvelabeling accuracyVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service data labeling by automatically capturing images from multiple predetermined viewpoints and using the physical object's known pose and position as ground truth labels. The multi-view geometry and calibrated camera positions allow automatic computation of object poses, eliminating the need for manual annotation while maintaining high labeling accuracy through the physical setup's inherent geometric constraints.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual labeling processes with automated computational methods. Instead of human annotators manually labeling images, the system uses calibrated camera positions, known object poses, and multi-view geometry to automatically compute ground truth labels. This substitution of mechanical/manual processes with computational automation dramatically reduces time consumption while maintaining precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If diverse real-world data is collected to capture nuances and complexities, then model performance is improved, but data collection complexity and resource requirements are worsened

Engineering Contradiction:
Improvedata diversityVSAvoidcollection system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the data collection process into distinct, manageable components: multiple cameras positioned at predetermined viewpoints, individual objects placed in controlled scenes, and systematic capture of images from each viewpoint. This segmentation allows diverse data collection to be broken down into repeatable, modular units that can be automated and scaled without proportionally increasing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs universal, calibrated cameras and standardized capture procedures that can be used across multiple scenes and objects. The same multi-view geometry principles and camera calibration methods apply universally to different objects and configurations, allowing diverse data collection through a unified, manageable framework rather than requiring specialized equipment for each scenario.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12293535B2Systems and methods for training pose estimators in computer vision
Publication Date: 2025.05.06 INTRINSIC INNOVATION LLC
  • US12293535B2 patent drawing
  • US12293535B2 patent drawing
  • US12293535B2 patent drawing

AI summary

A data capture stage includes a frame at least partially surrounding a target object, a rotation device within the frame and configured to selectively rotate the target object, a plurality of cameras coupled to the frame and configured to capture images of the target object from different angles, a sensor coupled to the frame and configured to sense mapping data corresponding to the target object, and an augmentation data generator configured to control a rotation of the rotation device, to control operations of the plurality of cameras and the sensor, and to generate training data based on the images and the mapping data.