Training Data Generation for Real-World Object Pose Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating training data for machine learning models to recognize object position and attitude suffer from accuracy drops in real-world environments due to discrepancies between simulated and actual conditions, particularly with specular reflection, leading to inefficient training processes.

Innovation Solution

A method involving prior learning using simulation data, followed by capturing images from different directions to adjust and correct object positions and attitudes, incorporating specular reflection considerations, to generate accurate training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If training data is generated by simulation, then data generation efficiency is improved, but recognition accuracy in actual environment deteriorates

Engineering Contradiction:
Improvedata generation efficiencyVSAvoidrecognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by first performing prior learning using simulation data to initialize the machine learning model, then using this pre-trained model to guide subsequent real environment data collection. This two-stage approach allows efficient simulation-based pre-training while ensuring the model adapts to real-world conditions, resolving the contradiction between data generation efficiency and recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If images are captured from single direction, then data collection time is reduced, but training data quality deteriorates due to insufficient viewpoint coverage

Engineering Contradiction:
Improvedata collection timeVSAvoidtraining data quality
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent applies dynamics by making the image capture process adaptive rather than static. The system dynamically determines the number and directions of image captures based on the object's characteristics and the current training progress. This allows the system to capture images from multiple directions when needed for high-quality training data, while reducing captures when sufficient data already exists, thus balancing time loss with training data quality.

Inventive Principle:
Principle #15Dynamics

3Speed

If machine learning model is trained with simulation data only, then training speed is improved, but adaptability to real environment deteriorates

Engineering Contradiction:
Improvetraining speedVSAvoidadaptability to real environment
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The patent uses real environment images captured by a camera as an intermediary between simulation data and the final deployment environment. The system captures actual images of the object in the real environment, uses the pre-trained model to recognize position and attitude, verifies correctness, and then uses this verified real data to fine-tune the model. This intermediary real-world data bridge maintains training speed while improving adaptability to the actual environment.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12567241B2Method for generating training data used to learn machine learning model, system, and non-transitory computer-readable storage medium storing computer program
Publication Date: 2026.03.03 SEIKO EPSON CORP
  • US12567241B2 patent drawing
  • US12567241B2 patent drawing
  • US12567241B2 patent drawing

AI summary

A method includes: (a) executing prior learning of the machine learning model, using simulation data of an object; (b) capturing a first image of the object from a first direction of image capture; (c) recognizing a first position and attitude of the object from the first image, using the machine learning model already learned through the prior learning; (d) performing a correctness determination about the first position and attitude; (e) capturing a second image of the object from a second direction of image capture that is different from the first direction of image capture when it is determined that the first position and attitude is correct, then converting the first position and attitude according to a change from the first direction of image capture to the second direction of image capture and thus calculating a second position and attitude, and assigning the second position and attitude to the second image and thus generating training data; and (f) changing an actual position and attitude of the object and repeating the (b) to (e).