Segmented Image Generation Using Multi-Signal Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer vision systems face challenges in accurately segmenting scenes that include transparent and reflecting objects, as these objects produce misleading information for traditional RGB and depth sensors, leading to difficulties in 3D reconstruction, especially in building interior and exterior scenes where biological entities like humans can further perturb the segmentation process.
Innovation Solution
A computer-implemented method that generates a segmented image of a scene using a combination of images from different physical signals, such as infrared, RGB, and depth images, employing Markov Random Field (MRF) energy minimization to improve segmentation accuracy. This method iteratively provides and processes multiple images from various viewpoints, excluding segments corresponding to biological entities to enhance 3D reconstruction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional RGB and depth sensors are used for segmentation, then the segmentation process is simple, but segmentation accuracy deteriorates due to misleading information from transparent and reflecting objects
Solution Approach 1:
The patent divides the segmentation task into multiple stages by processing different physical signals separately (infrared signal processing, visible light signal processing) and then combining results. This allows each signal type to be optimized for its specific characteristics, improving overall segmentation accuracy while managing complexity through modular processing.
Solution Approach 2:
The patent introduces infrared signals as an intermediary to capture thermal information that is independent of optical reflections and transparency. This intermediary signal provides complementary data that resolves ambiguities caused by transparent and reflecting objects in visible light imaging, thereby improving segmentation accuracy.
2Measurement precision
If multiple physical signals are processed to improve segmentation accuracy, then segmentation accuracy improves, but processing time increases
Solution Approach 1:
The patent performs preliminary processing of infrared and visible light signals separately before combining them. By pre-processing each signal type with appropriate algorithms optimized for that modality, the system reduces the computational burden of integrating multiple signals, thereby managing processing time while maintaining high segmentation accuracy.
Solution Approach 2:
The patent selectively processes signals based on scene characteristics. When transparent or reflecting objects are detected or suspected, the full multi-signal processing pipeline is activated. In other cases, simplified processing may suffice, reducing processing time while maintaining accuracy when needed.
3Measurement precision
If segments corresponding to biological entities are excluded to enhance 3D reconstruction accuracy, then 3D reconstruction accuracy improves, but the scope of segmentation is reduced
Solution Approach 1:
The patent extracts and separates biological entity segments from the overall segmentation for specialized processing. By identifying and isolating these segments, the system can apply specific handling rules that improve 3D reconstruction accuracy for static structures while preserving the ability to detect and track biological entities separately, thus maintaining versatility.
Solution Approach 2:
The patent applies different processing qualities to different parts of the scene. Biological entities receive specialized local processing that excludes them from certain 3D reconstruction operations, while the rest of the scene undergoes standard multi-signal integration. This local differentiation improves overall 3D reconstruction accuracy without completely sacrificing segmentation scope.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The method achieves improved segmentation and 3D reconstruction by leveraging complementary information from different physical signals, reducing the impact of misleading data from transparent and reflecting objects and enhancing the accuracy of 3D models in complex scenes.
Implementation Method 1
Each sensor is configured for a respective acquisition of a physical signal to which a respective one of the plurality of images of the scene corresponds; the one or more sensors comprise a material property sensor and one or both of an RGB sensor and a depth sensor; the material property sensor is an infrared sensor
Implementation Method 2
Depth data represents, for each pixel, the distance from the sensor. Depth data can be captured using available devices such as the Microsoft Kinect®, Asus XtionTM or Google TangoTM
Data Source
AI summary
A computer-implemented method of computer vision in a scene that includes one or more transparent objects and/or one or more reflecting objects comprises obtaining a plurality of images of the scene, each image corresponding to a respective acquisition of a physical signal, the plurality of images including at least two images corresponding to different physical signals; and generating a segmented image of the scene based on the plurality of images. This improves the field of computer vision.


