Surrogate Image Generation for Landmark Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image processing techniques for annotating landmarks on real two-dimensional images face challenges in accurately estimating object poses and generating reliable training datasets, especially when dealing with diverse real-world variations and large datasets required for machine learning models.

Innovation Solution

The method involves generating a training set of images by estimating object poses in real images, creating surrogate images using a three-dimensional model, perturbing characteristics to mimic real-world variations, and iteratively refining landmark annotations and pose estimations, ultimately improving the accuracy of landmark detection and pose estimation models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real images are used directly for training landmark detection models, then the training process is simple, but the accuracy of pose estimation and landmark detection is insufficient due to lack of diverse real-world variations

Engineering Contradiction:
Improveaccuracy of pose estimationVSAvoidcomplexity of training data generation
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates synthetic surrogate images by rendering 3D models of objects, copying their geometric and textural properties. These synthetic images serve as training data substitutes, providing diverse poses and variations without requiring extensive manual annotation of real images, thereby improving pose estimation accuracy while managing data generation complexity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system varies multiple parameters in the 3D model rendering process including lighting conditions, camera angles, object poses, and environmental factors. By systematically changing these parameters, the method generates diverse training images that capture real-world variations, improving model generalization and estimation accuracy

Inventive Principle:
Principle #35Parameter changes

2Reliability

If large datasets are collected for machine learning training, then model robustness improves, but the time and resources required for data collection and annotation increase significantly

Engineering Contradiction:
Improverobustness of machine learning modelVSAvoidtime for data collection and annotation
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system automatically generates its own training data by rendering 3D models with known ground truth annotations. This self-service approach eliminates the need for manual image collection and annotation, as the synthetic data generation process inherently provides labeled training examples, significantly reducing time investment while building robust models

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The method performs preliminary rendering of 3D models to create synthetic training images before the actual model training begins. By pre-generating diverse training data with known annotations, the system prepares comprehensive datasets in advance, avoiding time-consuming data collection during the training phase and enabling faster model development

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If pose estimation is performed on all real images, then comprehensive training data is obtained, but computational resources and processing time are excessively consumed

Engineering Contradiction:
Improvevolume of training dataVSAvoidcomputational resources for pose estimation
Core Design Contradiction:
Quantity of substanceVSUse of energy by moving object

Solution Approach 1:

Instead of estimating poses from 2D images and hoping for accurate results, the patent inverts the approach by rendering 3D models with known poses to create synthetic images. This inversion provides ground truth pose information directly, eliminating the need for computationally intensive pose estimation on all real images while still generating comprehensive training data with accurate annotations

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS11971953B2Machine annotation of photographic images
Publication Date: 2024.04.30 INAIT SA
  • US11971953B2 patent drawing
  • US11971953B2 patent drawing
  • US11971953B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for image processing that involves annotating landmarks on real two-dimensional images. In one aspect, the methods include generating a training set of images of an object for landmark detection. This includes receiving a collection of real images of an object, estimating a pose of the object in each real image in a proper subset of the collection of real images, creating a collection of surrogate images of the object for the training set using the estimated poses and a three-dimensional model of the object.