Training Apparatus for Depth Image Object Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for recognizing objects in images with depth information fail to accurately identify objects due to incomplete three-dimensional representation, as parts not visible in the image are not represented in the generated three-dimensional data.

Innovation Solution

A training apparatus and method that generates training data representing different parts of an object from various viewpoints, allowing a machine learning model to be trained on these data, enabling accurate recognition of objects in images with depth information by using a combination of three-dimensional data and images with depth information as training and recognition targets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If three-dimensional data is generated on the basis of an image associated with depth information, then the recognition process can be executed using a trained model, but parts not represented in the image are not represented in the three-dimensional data resulting in incomplete three-dimensional shape representation

Engineering Contradiction:
Improveautomation of recognition processVSAvoidincomplete three-dimensional shape information
Core Design Contradiction:
Extent of automationVSLoss of information

Solution Approach 1:

The training data is segmented into multiple parts, where each part represents a different visible portion of the object from a specific viewpoint. By dividing the complete three-dimensional object into multiple viewpoint-specific segments, the system captures comprehensive object information that compensates for the incompleteness of any single depth image, thereby resolving the information loss while maintaining automated recognition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from two-dimensional depth images to three-dimensional training data by generating multiple three-dimensional representations from different virtual viewpoints. This dimensional expansion allows the system to capture object parts that are occluded in any single viewpoint, thereby recovering the complete three-dimensional shape information lost in the original depth image.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If a trained model is trained using three-dimensional data representing the complete three-dimensional shape of the object, then the model can potentially recognize objects accurately, but recognition failure occurs when the training data does not match the incomplete nature of input data from depth images

Engineering Contradiction:
Improverecognition accuracyVSAvoidrecognition reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

Each training data instance is optimized for its specific viewpoint, containing detailed information about the object parts visible from that particular angle. This local quality optimization ensures that the training data matches the characteristics of the input data from depth images, improving both recognition accuracy and reliability by eliminating the mismatch between training and inference conditions.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system performs preliminary generation of multiple viewpoint-specific three-dimensional training data before the recognition process. By preparing comprehensive training data covering all possible viewpoints in advance, the system ensures that the trained model is ready to handle any input configuration, thereby improving recognition reliability without sacrificing accuracy.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If training data representing different parts of the object from various positions is generated, then comprehensive object coverage is achieved, but the complexity of data generation and processing increases

Engineering Contradiction:
Improvecompleteness of object representationVSAvoidcomplexity of training data generation
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

Instead of physically capturing objects from multiple viewpoints, the system creates virtual copies of the three-dimensional object data and transforms them to represent different viewpoints. This copying approach achieves comprehensive object coverage without the complexity of physical multi-view acquisition systems, as the virtual copies can be generated computationally from a single complete three-dimensional representation.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11681910B2Training apparatus, recognition apparatus, training method, recognition method, and program
Publication Date: 2023.06.20 SONY INTERACTIVE ENTERTAINMENT LLC
  • US11681910B2 patent drawing
  • US11681910B2 patent drawing
  • US11681910B2 patent drawing

AI summary

Provided are a training apparatus, a recognition apparatus, a training method, a recognition method, and a program that can accurately recognize what an object represented in an image associated with depth information is. An object data acquiring section acquires three-dimensional data representing an object. A training data generating section generates a plurality of training data each representing a mutually different part of the object on the basis of the three-dimensional data. A training section trains a machine learning model using the generated training data as the training data for the object.