Training Apparatus for Depth Image Object Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for recognizing objects in images with depth information fail to accurately identify objects due to incomplete three-dimensional representation, as parts not visible in the image are not represented in the generated three-dimensional data.
Innovation Solution
A training apparatus and method that generates training data representing different parts of an object from various viewpoints, allowing a machine learning model to be trained on these data, enabling accurate recognition of objects in images with depth information by using a combination of three-dimensional data and images with depth information as training and recognition targets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If three-dimensional data is generated on the basis of an image associated with depth information, then the recognition process can be executed using a trained model, but parts not represented in the image are not represented in the three-dimensional data resulting in incomplete three-dimensional shape representation
Solution Approach 1:
The training data is segmented into multiple parts, where each part represents a different visible portion of the object from a specific viewpoint. By dividing the complete three-dimensional object into multiple viewpoint-specific segments, the system captures comprehensive object information that compensates for the incompleteness of any single depth image, thereby resolving the information loss while maintaining automated recognition.
Solution Approach 2:
The patent transitions from two-dimensional depth images to three-dimensional training data by generating multiple three-dimensional representations from different virtual viewpoints. This dimensional expansion allows the system to capture object parts that are occluded in any single viewpoint, thereby recovering the complete three-dimensional shape information lost in the original depth image.
2Measurement precision
If a trained model is trained using three-dimensional data representing the complete three-dimensional shape of the object, then the model can potentially recognize objects accurately, but recognition failure occurs when the training data does not match the incomplete nature of input data from depth images
Solution Approach 1:
Each training data instance is optimized for its specific viewpoint, containing detailed information about the object parts visible from that particular angle. This local quality optimization ensures that the training data matches the characteristics of the input data from depth images, improving both recognition accuracy and reliability by eliminating the mismatch between training and inference conditions.
Solution Approach 2:
The system performs preliminary generation of multiple viewpoint-specific three-dimensional training data before the recognition process. By preparing comprehensive training data covering all possible viewpoints in advance, the system ensures that the trained model is ready to handle any input configuration, thereby improving recognition reliability without sacrificing accuracy.
3Loss of information
If training data representing different parts of the object from various positions is generated, then comprehensive object coverage is achieved, but the complexity of data generation and processing increases
Solution Approach 1:
Instead of physically capturing objects from multiple viewpoints, the system creates virtual copies of the three-dimensional object data and transforms them to represent different viewpoints. This copying approach achieves comprehensive object coverage without the complexity of physical multi-view acquisition systems, as the virtual copies can be generated computationally from a single complete three-dimensional representation.
Data Source
AI summary
Provided are a training apparatus, a recognition apparatus, a training method, a recognition method, and a program that can accurately recognize what an object represented in an image associated with depth information is. An object data acquiring section acquires three-dimensional data representing an object. A training data generating section generates a plurality of training data each representing a mutually different part of the object on the basis of the three-dimensional data. A training section trains a machine learning model using the generated training data as the training data for the object.


