Time-of-flight Object Recognition via Synthetic Data Masking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing object recognition methods for time-of-flight camera data require extensive and diverse training data, which can be time-consuming and inefficient to collect, and are limited by biases in training datasets, leading to errors in object detection, such as misidentifying a hand on a chest as a buckled seatbelt.

Innovation Solution

The method generates time-of-flight training data by combining real background data with simulated object data using a mask applied to synthetic overlay image data, allowing for diverse and extensive training without the need for extensive real-world data collection, and uses a pretrained algorithm trained on this data to recognize objects in time-of-flight camera images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If extensive and diverse real training data is collected for object recognition, then detection accuracy is improved, but time consumption and resource requirements increase

Engineering Contradiction:
Improvedetection accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates synthetic training data by generating simulated time-of-flight images of objects (such as seatbelts, hands, gestures) and combining them with real background images. This copying approach allows the system to train on diverse object instances without physically collecting extensive real-world data, thereby reducing time consumption while maintaining detection accuracy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary data preparation by pre-processing synthetic object images and pre-computing their time-of-flight characteristics before actual training. This includes generating object masks, calculating depth information, and preparing training samples in advance, which streamlines the subsequent training process and reduces overall time consumption.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If extensive and diverse real training data is collected for object recognition, then detection accuracy is improved, but hardware and human resources increase

Engineering Contradiction:
Improvedetection accuracyVSAvoidresource requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent generates synthetic training data through computational methods, eliminating the need to physically collect and store extensive real-world training datasets. This reduces both storage requirements and the human resources needed for data collection, annotation, and management while providing diverse training samples for improved detection accuracy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent creates a universal training data generation system that can produce diverse object instances (different persons, postures, lighting conditions) from a single synthetic object model. This multi-functional approach allows one set of synthetic data generation tools to serve multiple training needs, reducing overall resource requirements compared to collecting specialized real-world data for each scenario.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If real training data is used for object recognition, then detection accuracy is improved, but biases in training datasets cause detection errors

Engineering Contradiction:
Improvedetection accuracyVSAvoiddetection reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies different characteristics to different parts of the training data: real background images provide authentic environmental context, while synthetic object images provide controlled, bias-free object representations. This local quality differentiation allows the system to benefit from the realism of actual backgrounds without inheriting their biases, improving both accuracy and reliability.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Instead of using real objects in real backgrounds (which introduces biases), the patent inverts the approach by placing synthetic objects into real backgrounds. This inversion allows control over object characteristics while maintaining environmental realism, thereby reducing bias-related detection errors while preserving detection accuracy.

Inventive Principle:
Principle #13The other way round (Inversion)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables efficient and robust object recognition with reduced human and hardware resources, improving detection accuracy by avoiding biases and allowing objects to be recognized on various backgrounds, such as a zebra on green or yellow grass.

Implementation Method 1

ToF cameras may measure a roundtrip delay of emitted light (which is reflected at a scene (e.g. object)) which may be indicative of a depth, i.e. the distance to the scene.

Methodology Applied
Scientific EffectTime of flight: Time of Flight

Data Source

PatentUS20240071122A1Object recognition method and time-of-flight object recognition circuitry
Publication Date: 2024.02.29 SONY SEMICON SOLUTIONS CORP
  • US20240071122A1 patent drawing
  • US20240071122A1 patent drawing
  • US20240071122A1 patent drawing

AI summary

The present disclosure generally pertains to an object recognition method for time-of-flight camera data, including: recognizing a real object based on a pretrained algorithm, wherein the pretrained algorithm is trained based on time-of-flight training data, wherein the time-of-flight training data are generated based on a combination of real time-of-flight data being indicative of a background, and simulated time-of-flight data generated by applying a mask on synthetic overlay image data representing a simulated object, thereby generating a masked simulated object, the mask being generated based on the synthetic overlay image data.