3D Bounding Volume Projection for Automated Training Data Annotation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generating specialized training data sets for image recognition algorithms is time-consuming and costly due to the need for manual tagging of large numbers of images, limiting the practical applications of image recognition technologies.
Innovation Solution
A computing system that includes optical sensors, position sensors, and a processor to generate three-dimensional representations of environments, detect physical objects, and create two-dimensional bounding shapes with labels, allowing for automated annotation and generation of training data sets with user input for customization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual tagging of large numbers of images is used to generate training data sets, then the quality and accuracy of training data can be improved, but the time consumption and cost increase significantly
Solution Approach 1:
The system performs preliminary automated annotation of images using machine learning models before manual review. This pre-annotation step prepares the data in advance, so that human annotators only need to review and correct predictions rather than create annotations from scratch, significantly reducing the time required while maintaining high quality
Solution Approach 2:
The system introduces an automated machine learning-based annotation system as an intermediary between raw images and final training data. This intermediary performs initial annotation work, which then serves as input for manual verification, creating a multi-stage process that balances automation efficiency with human quality control
2Measurement precision
If manual tagging of large numbers of images is used to generate training data sets, then the quality and accuracy of training data can be improved, but the cost increases significantly
Solution Approach 1:
The system performs preliminary automated annotation of images using machine learning models before manual review. This pre-annotation step prepares the data in advance, so that human annotators only need to review and correct predictions rather than create annotations from scratch, significantly reducing the time required while maintaining high quality
Solution Approach 2:
The system introduces an automated machine learning-based annotation system as an intermediary between raw images and final training data. This intermediary performs initial annotation work, which then serves as input for manual verification, creating a multi-stage process that balances automation efficiency with human quality control
3Productivity
If automated annotation systems are used to generate training data sets, then the time consumption and cost can be reduced, but the accuracy and precision of annotations may decrease
Solution Approach 1:
The system merges automated machine learning annotation with manual human review in a unified workflow. The automated system handles high-volume processing while human annotators handle quality control, combining the speed of automation with the precision of human judgment to achieve both high productivity and high accuracy
Solution Approach 2:
The system implements a feedback loop where manually corrected annotations are used to retrain and improve the automated annotation model. This continuous feedback mechanism allows the system to learn from human corrections and progressively improve annotation accuracy while maintaining high throughput
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A computing system is provided, including one or more optical sensors, a display, one or more user input devices, and a processor. The processor may receive optical data of a physical environment. Based on the optical data, the processor may generate a three-dimensional representation of the physical environment. For at least one target region of the physical environment, the processor may generate a three-dimensional bounding volume surrounding the target region based on a depth profile measured by the one or more optical sensors and/or estimated by the processor. The processor may generate a two-dimensional bounding shape at least in part by projecting the three-dimensional bounding volume onto an imaging surface of an optical sensor. The processor may output an image of the physical environment and the two-dimensional bounding shape for display. The processor may receive a user input and modify the two-dimensional bounding shape based on the user input.