3D Cuboid Estimation from 2D Annotations and Sensor Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Preparing ground truth information for supervised machine learning techniques used in automated control systems for vehicles is a time-consuming manual process, which hampers efficient object detection and estimation in complex scenarios.
Innovation Solution
A method that involves obtaining a two-dimensional image annotation, determining a location proposal, classifying the object, estimating its size, and defining a three-dimensional cuboid based on the proposal and size, using automated processes such as trained machine learning models for annotation and classification, and integrating map information for accurate positioning and orientation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual processes are used to prepare ground truth information, then accuracy of object detection can be maintained, but the process becomes time-consuming and reduces productivity
Solution Approach 1:
The system uses automated machine learning models to generate ground truth information independently without requiring manual annotation. The trained models process sensor inputs and automatically produce object detections, classifications, and three-dimensional estimates, enabling the system to serve itself in creating training data.
Solution Approach 2:
The patent replaces manual mechanical annotation processes with automated computational systems. Machine learning models substitute human operators, automatically generating ground truth information from sensor data through algorithmic processing rather than manual labeling.
2Productivity
If automated processes are used to generate ground truth information, then productivity is improved, but the complexity of the system increases
Solution Approach 1:
The system employs multi-functional machine learning models that can perform multiple tasks including object detection, classification, and three-dimensional estimation simultaneously. This universal approach consolidates multiple specialized processes into integrated models, managing complexity through functional consolidation.
Solution Approach 2:
The patent transforms two-dimensional image annotations into three-dimensional object estimates by adding depth information. This dimensional transformation enables the system to work with richer data representations while using the same underlying annotation infrastructure, effectively managing complexity through dimensional enhancement rather than process multiplication.
3Productivity
If two-dimensional annotations are used instead of three-dimensional annotations, then the annotation process becomes simpler and faster, but the accuracy of three-dimensional object estimation may be reduced
Solution Approach 1:
The system uses two-dimensional annotations as an intermediary representation that is then transformed into three-dimensional estimates through machine learning models. The 2D annotations serve as a convenient intermediate format that balances annotation simplicity with the ability to generate accurate 3D information through computational transformation.
Solution Approach 2:
The patent changes the parameter representation from direct three-dimensional annotations to two-dimensional annotations with derived depth information. By altering how spatial parameters are represented and processed, the system achieves both annotation efficiency and three-dimensional estimation accuracy through parameter transformation rather than direct measurement.
Data Source
AI summary
A method includes obtaining a two-dimensional image, obtaining a two-dimensional image annotation that indicates presence of an object in the two-dimensional image, obtaining three-dimensional sensor information, generating a top-down representation of the three-dimensional sensor information, and obtaining a top-down annotation that indicates presence of the object in the top-down representation. The method also includes determining a bottom surface of a three-dimensional cuboid based on map information, determining a position, a length, a width, and a yaw rotation of the three-dimensional cuboid based on the top-down annotation, and determining a height of the three-dimensional cuboid based on a two-dimensional image annotation, and the position, the length, the width, and the yaw rotation of the three-dimensional cuboid.


