Synthetic Surrogate Image Generation for Landmark Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image processing technologies face challenges in efficiently annotating landmarks on real two-dimensional images, particularly in generating accurate training datasets for machine learning models used in applications like pose estimation and image classification.

Innovation Solution

The method involves generating a training set of images by receiving real images of an object, estimating the pose of the object in a subset of these images, creating surrogate images using a three-dimensional model of the object, and perturbing characteristics of the object to mimic real-world variations. Landmarks are then labeled on these surrogate images based on the three-dimensional model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real images are manually annotated with landmarks, then training dataset accuracy is improved, but time consumption and labor cost increase significantly

Engineering Contradiction:
Improvelandmark annotation accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates synthetic surrogate images by rendering 3D models of objects from multiple camera positions and perspectives. These synthetic images with automatically generated landmark annotations serve as copies that replace the need for manual annotation of real images, significantly reducing time and labor while maintaining training quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system pre-computes surrogate images and their corresponding landmark annotations before actual training begins. By preparing this synthetic training data in advance through automated 3D rendering and pose estimation, the time-consuming annotation task is performed once on synthetic data rather than repeatedly on real images

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If more diverse training images are generated, then model generalization is improved, but data processing complexity increases

Engineering Contradiction:
Improvemodel generalizationVSAvoiddata processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system varies multiple parameters in the 3D rendering process including camera position, lighting conditions, object poses, and background environments to generate diverse synthetic training images. This automated parameter variation achieves model generalization without manual intervention, managing complexity through systematic control

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The 3D model rendering system serves multiple functions simultaneously: generating training images, creating augmented views, establishing pose-label correspondences, and simulating various real-world conditions. This multi-functional approach achieves diversity efficiently through a single unified processing framework

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If pose estimation accuracy is improved, then surrogate image quality is improved, but computational requirements increase

Engineering Contradiction:
Improvepose estimation accuracyVSAvoidcomputational energy
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system uses the trained machine learning model itself to generate the training data it needs. The model estimates poses on real images, which are then used to create synthetic surrogate images with automatic annotations, creating a self-reinforcing cycle that improves accuracy without proportional increases in computational cost

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Pose estimation is performed on a subset of real images to generate the initial set of surrogate images. This preliminary pose estimation on limited data reduces immediate computational requirements while still enabling the creation of sufficient training data to improve overall model performance

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12235928B2Machine annotation of photographic images
Publication Date: 2025.02.25 INAIT SA
  • US12235928B2 patent drawing
  • US12235928B2 patent drawing
  • US12235928B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for image processing that involves annotating landmarks on real two-dimensional images. In one aspect, the methods include generating a training set of images of an object for landmark detection. This includes receiving a collection of real images of an object, estimating a pose of the object in each real image in a proper subset of the collection of real images, creating a collection of surrogate images of the object for the training set using the estimated poses and a three-dimensional model of the object.