Synthetic Training Image Generation for Object Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual annotation of images for training object recognition models is time-consuming, error-prone, and requires a large set of suitable images, which may not always be available, leading to biased and less accurate models, especially when recognizing objects in varying environments.

Innovation Solution

Using computer-generated training images that depict a digital training model of real-world objects, allowing for faster and more efficient training of object recognition models, and providing augmentations such as visual content, instructions, and virtual controls after recognition, enabling accurate object identification and manipulation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual annotation of images is used to train object recognition models, then the models can be trained with real-world data, but the process is time-consuming and error-prone

Engineering Contradiction:
Improveaccuracy of object recognition modelVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses digital training models (virtual copies) of real-world objects to generate synthetic training images, replacing the need for manual annotation of physical objects. This copying approach maintains training effectiveness while eliminating time-consuming manual processes

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary actions by pre-creating digital training models and generating synthetic training images in advance. This prepares training data beforehand, eliminating the need for time-consuming manual annotation during the actual training process

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If a large set of annotated images is collected for training, then the model can be trained comprehensively, but the process becomes more complex and resource-intensive

Engineering Contradiction:
Improvemodel performance across varying environmentsVSAvoidcomplexity of training process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The digital training models serve multiple functions: they can be rendered in various environments, lighting conditions, and perspectives to generate diverse training images from a single model. This universal approach replaces the need for collecting numerous specific annotated images

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes parameters such as lighting conditions, background environments, object positions, and camera angles when rendering synthetic training images from digital models. This generates diverse training data through parameter variation rather than collecting diverse physical samples

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If manual annotation is performed by human annotators, then real-world objects can be accurately labeled, but human error and bias are introduced

Engineering Contradiction:
Improveaccuracy of object labelingVSAvoidconsistency and objectivity of annotations
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system uses automated rendering processes where digital training models generate their own training images with inherent labeling information. This self-service approach eliminates human annotators, removing the source of human error and bias while maintaining accuracy through precise digital modeling

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11132845B2Real-world object recognition for computing device
Publication Date: 2021.09.28 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11132845B2 patent drawing
  • US11132845B2 patent drawing
  • US11132845B2 patent drawing

AI summary

A method for object recognition includes, at a computing device, receiving an image of a real-world object. An identity of the real-world object is recognized using an object recognition model trained on a plurality of computer-generated training images. A digital augmentation model corresponding to the real-world object is retrieved, the digital augmentation model including a set of augmentation-specific instructions. A pose of the digital augmentation model is aligned with a pose of the real-world object. An augmentation is provided, the augmentation associated with the real-world object and specified by the augmentation-specific instructions.