3D Map Labeling via 2D Pixel Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing devices face overload and long processing times when applying object recognition techniques to three-dimensional maps, and collecting learning data for three-dimensional maps is difficult and costly compared to two-dimensional images.

Innovation Solution

An image processing device that acquires two-dimensional input images frame by frame, performs object type recognition to label pixels, and creates a three-dimensional map by recognizing the three-dimensional position of objects, attaching labels to voxels based on pixel labels and determining labels by analyzing multiple frames.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If object recognition technique is applied to three-dimensional map, then object type recognition is achieved, but image processing device is overloaded and processing time increases

Engineering Contradiction:
Improveobject type recognition accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the object recognition process by first performing recognition on two-dimensional input images to obtain pixel labels, then mapping these labels to three-dimensional map voxels. This division separates the complex 3D recognition task into manageable 2D recognition followed by coordinate transformation, reducing computational overload while maintaining recognition accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary object type recognition on two-dimensional images before creating the three-dimensional map. By pre-processing and labeling pixels in 2D space first, the system prepares recognition results in advance, which are then efficiently transferred to 3D voxel space, avoiding the need to perform full 3D recognition and thus reducing processing time and device load.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If machine learning model is trained for three-dimensional map object recognition, then accurate recognition is achieved, but collecting learning data becomes difficult and costly

Engineering Contradiction:
Improveobject type recognition accuracyVSAvoidlearning data collection difficulty
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent uses two-dimensional images as copies or projections of the three-dimensional world. Instead of collecting difficult-to-obtain 3D labeled data, the system collects easily available 2D images, performs recognition on them, and then maps the results to 3D space. This copying approach allows using standard 2D image datasets for training while achieving 3D map recognition capabilities.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent inverts the conventional approach by not directly training a model on 3D data, but rather training on 2D data and then transforming the results to 3D. This inversion leverages the ease of 2D data collection while achieving the ultimate goal of 3D map object recognition, avoiding the bottleneck of 3D learning data collection.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS12211249B2Image processing device, image processing method, and program
Publication Date: 2025.01.28 SONY INTERACTIVE ENTERTAINMENT LLC
  • US12211249B2 patent drawing
  • US12211249B2 patent drawing
  • US12211249B2 patent drawing

AI summary

Provided are an image processing device, an image processing method, and a program for recognizing an object in a three-dimensional map, which do not require collecting learning data of the three-dimensional map and can perform high-speed processing with a small load. The image processing device includes an image acquiring section that sequentially acquires a two-dimensional input image for each frame; an object type recognition executing section that attaches, to each of pixels of the input image acquired for each frame, a label indicating a type of an object represented by the pixels; and a labeling section that executes three-dimensional position recognition of a subject represented in the input image to create a three-dimensional map, based on the input image sequentially input, and attaches, to each voxel included in the three-dimensional map, the label of the pixel corresponding to the voxel.