Mobile Object Key-Point Modeling Without 3D Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for creating machine learning models for object detection, such as those described in Garrick Brazil et al., require extensive time to collect three-dimensional data, leading to prolonged operation times.
Innovation Solution
A method and apparatus that utilize two-dimensional images to specify key points corresponding to a mobile object's projected shape on a road, creating training data and generating a machine learning model without requiring three-dimensional data, using a neural network structure with a Base Net, Spatial Net, and discriminator to output key points.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If three-dimensional data (LiDAR) is used to create training data for machine learning models, then detection accuracy is improved, but the time required to collect training data and create the model becomes enormous
Solution Approach 1:
The patent creates a two-dimensional image that copies the essential geometric information of the mobile object's projected shape on the road. Instead of using complex three-dimensional LiDAR data, the invention creates a simplified 2D representation containing the necessary spatial relationships (vertexes of the rectangular projection) to train the machine learning model effectively, thereby reducing data collection time while maintaining detection accuracy
Solution Approach 2:
The patent extracts only the necessary key information from the complex three-dimensional scene - specifically the vertexes of the mobile object's projected rectangular shape on the road. By taking out only these critical geometric features and representing them in a two-dimensional image with added coordinate information, the invention eliminates the need to collect and process enormous amounts of full 3D LiDAR data
2Reliability
If three-dimensional data is used for training, then the machine learning model achieves better performance, but the operation time to create the model becomes enormous
Solution Approach 1:
The patent changes the parameter representation from three-dimensional coordinates to two-dimensional image coordinates with added key point information. By transforming the data representation parameters from 3D spatial coordinates to 2D image plane coordinates with annotated vertex positions, the invention reduces the complexity of the training data while preserving the essential geometric relationships needed for accurate model performance
3Loss of time
If two-dimensional images are used instead of three-dimensional data, then data collection time is reduced, but sufficient information for accurate detection may be lost
Solution Approach 1:
The patent performs preliminary action by pre-annotating the two-dimensional image with key point information - specifically the coordinates of vertexes of the mobile object's projected rectangular shape. This preliminary addition of critical spatial information to the 2D image ensures that when the machine learning model is trained, it receives sufficient geometric data without requiring complex 3D inputs, thus preventing information loss while maintaining fast data collection
Data Source
AI summary
A model generation method includes specifying image coordinates that fall within a two-dimensional image obtained by capturing at least a mobile object and that correspond to at least one point among vertexes of a rectangular shape formed when an outer shape of the mobile object viewed from above is projected on a road, as a key point of the mobile object, and creating the two-dimensional image, to which information on the key point is added, as training data, and generating a machine learning model that outputs the key point from a two-dimensional image obtained by capturing at least a mobile object, by performing machine learning using the training data.


