Mobile Object Key-Point Modeling Without 3D Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for creating machine learning models for object detection, such as those described in Garrick Brazil et al., require extensive time to collect three-dimensional data, leading to prolonged operation times.

Innovation Solution

A method and apparatus that utilize two-dimensional images to specify key points corresponding to a mobile object's projected shape on a road, creating training data and generating a machine learning model without requiring three-dimensional data, using a neural network structure with a Base Net, Spatial Net, and discriminator to output key points.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If three-dimensional data (LiDAR) is used to create training data for machine learning models, then detection accuracy is improved, but the time required to collect training data and create the model becomes enormous

Engineering Contradiction:
Improvedetection accuracyVSAvoidtime to collect training data and create model
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates a two-dimensional image that copies the essential geometric information of the mobile object's projected shape on the road. Instead of using complex three-dimensional LiDAR data, the invention creates a simplified 2D representation containing the necessary spatial relationships (vertexes of the rectangular projection) to train the machine learning model effectively, thereby reducing data collection time while maintaining detection accuracy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent extracts only the necessary key information from the complex three-dimensional scene - specifically the vertexes of the mobile object's projected rectangular shape on the road. By taking out only these critical geometric features and representing them in a two-dimensional image with added coordinate information, the invention eliminates the need to collect and process enormous amounts of full 3D LiDAR data

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If three-dimensional data is used for training, then the machine learning model achieves better performance, but the operation time to create the model becomes enormous

Engineering Contradiction:
Improvemodel performanceVSAvoidoperation time to create model
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent changes the parameter representation from three-dimensional coordinates to two-dimensional image coordinates with added key point information. By transforming the data representation parameters from 3D spatial coordinates to 2D image plane coordinates with annotated vertex positions, the invention reduces the complexity of the training data while preserving the essential geometric relationships needed for accurate model performance

Inventive Principle:
Principle #35Parameter changes

3Loss of time

If two-dimensional images are used instead of three-dimensional data, then data collection time is reduced, but sufficient information for accurate detection may be lost

Engineering Contradiction:
Improvedata collection timeVSAvoidspatial information
Core Design Contradiction:
Loss of timeVSLoss of information

Solution Approach 1:

The patent performs preliminary action by pre-annotating the two-dimensional image with key point information - specifically the coordinates of vertexes of the mobile object's projected rectangular shape. This preliminary addition of critical spatial information to the 2D image ensures that when the machine learning model is trained, it receives sufficient geometric data without requiring complex 3D inputs, thus preventing information loss while maintaining fast data collection

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12462576B2Model generation method, model generation apparatus, non-transitory storage medium, mobile object posture estimation method, and mobile object posture estimation apparatus
Publication Date: 2025.11.04 TOYOTA JIDOSHA KK
  • US12462576B2 patent drawing
  • US12462576B2 patent drawing
  • US12462576B2 patent drawing

AI summary

A model generation method includes specifying image coordinates that fall within a two-dimensional image obtained by capturing at least a mobile object and that correspond to at least one point among vertexes of a rectangular shape formed when an outer shape of the mobile object viewed from above is projected on a road, as a key point of the mobile object, and creating the two-dimensional image, to which information on the key point is added, as training data, and generating a machine learning model that outputs the key point from a two-dimensional image obtained by capturing at least a mobile object, by performing machine learning using the training data.