Sensor Fusion Pixel Correction Using Learned Pseudo Depth Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

When sensor fusion is used, identifying and appropriately processing correction target pixels, such as defective pixels, is challenging due to variations in detection principles across different sensors, which affects the quality of the resulting image data.

Innovation Solution

An information processing apparatus and method that utilize a learned model from machine learning to process images from multiple sensors, including depth and RGB sensors, by generating pseudo images and comparing them to identify correction target pixels, thereby improving pixel correction and image quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If sensor fusion is used to combine multiple sensors with different detection principles, then the quality and completeness of measurement data is improved, but the difficulty of detecting and measuring correction target pixels increases due to variations in detection principles

Engineering Contradiction:
Improvequality of measurement dataVSAvoiddifficulty of identifying correction target pixels
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent divides the complex task of correction target pixel detection into multiple processing stages: generating pseudo depth maps from RGB images through machine learning, comparing these pseudo depth maps with actual depth sensor data, and identifying discrepancies that indicate correction target pixels. This segmentation approach makes the detection process more manageable and accurate despite different sensor principles.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces pseudo depth maps as an intermediary element that bridges the gap between RGB sensor data and depth sensor data. By converting RGB images into pseudo depth information through machine learning models, the system creates a common reference framework that enables comparison and identification of correction target pixels across different sensor types.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If correction target pixels are identified and corrected in sensor fusion processing, then the accuracy of recognition processing is improved, but the device complexity increases due to additional processing steps

Engineering Contradiction:
Improveaccuracy of recognition processingVSAvoidcomplexity of processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system uses the RGB sensor data itself to generate pseudo depth maps that serve as the correction reference, eliminating the need for external reference devices or complex manual calibration processes. The machine learning model automatically performs the conversion and comparison tasks using the available sensor inputs.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces complex mechanical or manual correction methods with machine learning-based pseudo depth map generation. Instead of using elaborate physical calibration systems or manual pixel-by-pixel correction, the system employs neural networks to automatically generate correction references and identify target pixels.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240161254A1Information processing apparatus, information processing method, and program
Publication Date: 2024.05.16 SONY SEMICON SOLUTIONS CORP
  • US20240161254A1 patent drawing
  • US20240161254A1 patent drawing
  • US20240161254A1 patent drawing

AI summary

The present disclosure relates to an information processing apparatus, an information processing method, and a program capable of more appropriately processing a correction target pixel when sensor fusion is used. Provided is an information processing apparatus including a processing unit that performs processing using a learned model learned by machine learning on at least a part of a first image in which an object acquired by a first sensor is indicated by depth information, a second image in which an image of the object acquired by a second sensor is indicated by plane information, and a third image obtained from the first image and the second image to specify a correction target pixel included in the first image. The present disclosure can be applied to, for example, a device having a plurality of sensors.