3D Bounding Volume Projection for Automated Training Data Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating specialized training data sets for image recognition algorithms is time-consuming and costly due to the need for manual tagging of large numbers of images, limiting the practical applications of image recognition technologies.

Innovation Solution

A computing system that includes optical sensors, position sensors, and a processor to generate three-dimensional representations of environments, detect physical objects, and create two-dimensional bounding shapes with labels, allowing for automated annotation and generation of training data sets with user input for customization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual tagging of large numbers of images is used to generate training data sets, then the quality and accuracy of training data can be improved, but the time consumption and cost increase significantly

Engineering Contradiction:
Improvetraining data qualityVSAvoiddata generation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary automated annotation of images using machine learning models before manual review. This pre-annotation step prepares the data in advance, so that human annotators only need to review and correct predictions rather than create annotations from scratch, significantly reducing the time required while maintaining high quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an automated machine learning-based annotation system as an intermediary between raw images and final training data. This intermediary performs initial annotation work, which then serves as input for manual verification, creating a multi-stage process that balances automation efficiency with human quality control

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual tagging of large numbers of images is used to generate training data sets, then the quality and accuracy of training data can be improved, but the cost increases significantly

Engineering Contradiction:
Improvetraining data qualityVSAvoiddata generation cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system performs preliminary automated annotation of images using machine learning models before manual review. This pre-annotation step prepares the data in advance, so that human annotators only need to review and correct predictions rather than create annotations from scratch, significantly reducing the time required while maintaining high quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an automated machine learning-based annotation system as an intermediary between raw images and final training data. This intermediary performs initial annotation work, which then serves as input for manual verification, creating a multi-stage process that balances automation efficiency with human quality control

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated annotation systems are used to generate training data sets, then the time consumption and cost can be reduced, but the accuracy and precision of annotations may decrease

Engineering Contradiction:
Improvedata generation efficiencyVSAvoidannotation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system merges automated machine learning annotation with manual human review in a unified workflow. The automated system handles high-volume processing while human annotators handle quality control, combining the speed of automation with the precision of human judgment to achieve both high productivity and high accuracy

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system implements a feedback loop where manually corrected annotations are used to retrain and improve the automated annotation model. This continuous feedback mechanism allows the system to learn from human corrections and progressively improve annotation accuracy while maintaining high throughput

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3906527B1Image bounding shape using 3D environment representation
Publication Date: 2022.11.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3906527B1 patent drawingFigure 1
  • EP3906527B1 patent drawingFigure 2
  • EP3906527B1 patent drawingFigure 3

AI summary

A computing system is provided, including one or more optical sensors, a display, one or more user input devices, and a processor. The processor may receive optical data of a physical environment. Based on the optical data, the processor may generate a three-dimensional representation of the physical environment. For at least one target region of the physical environment, the processor may generate a three-dimensional bounding volume surrounding the target region based on a depth profile measured by the one or more optical sensors and/or estimated by the processor. The processor may generate a two-dimensional bounding shape at least in part by projecting the three-dimensional bounding volume onto an imaging surface of an optical sensor. The processor may output an image of the physical environment and the two-dimensional bounding shape for display. The processor may receive a user input and modify the two-dimensional bounding shape based on the user input.