Transparent Object Detection Using Visible and Thermal Image Discrepancies

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing imaging technologies struggle to detect and accurately determine the presence and depth of transparent objects such as windows, as they often appear as voids or inaccuracies in depth maps due to specular reflection and lossy reflection in IR and VL images.

Innovation Solution

A computing system utilizing a visible light camera, thermal camera, and processor to detect image discrepancies between VL and thermal images, employing machine learning classifiers to identify transparent objects and determine their depth by processing image data from both sources, generating point clouds and surface meshes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional imaging technologies (VL cameras, IR cameras, thermal cameras) are used to detect transparent objects, then the imaging process is simple and fast, but transparent objects cannot be accurately detected and depth values cannot be determined

Engineering Contradiction:
Improvedetection accuracy of transparent objectsVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple imaging modalities (visible light camera, thermal camera, and depth sensor) into a unified system. The processor integrates data from all three sources, using visible light images for general scene understanding, thermal images for detecting transparent objects through reflection patterns, and depth sensors for spatial information. This merging allows accurate detection of transparent objects that cannot be detected by any single modality alone.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary processing layer (the processor with machine learning classifiers) that mediates between the raw image data from multiple cameras and the final detection results. This intermediary layer analyzes discrepancies between visible light and thermal images, applies machine learning models trained on transparent object patterns, and synthesizes depth information from multiple sources to produce accurate depth maps with transparent objects properly identified.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple imaging sources (visible light camera, thermal camera, depth sensor) are used to detect transparent objects, then detection accuracy is improved, but processing time and computational load increase

Engineering Contradiction:
Improvedetection accuracy of transparent objectsVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the image processing task into distinct functional modules: (1) separate processing of visible light and thermal images to identify transparent object regions, (2) independent depth estimation for different scene regions, (3) specialized machine learning classification for transparent object detection, and (4) final synthesis of depth maps. This segmentation allows each module to process data independently and optimize for its specific function, reducing overall processing time while maintaining high accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-processing individual image streams (visible light and thermal) to identify potential transparent object regions before the main processing stage. Machine learning classifiers are pre-trained on transparent object patterns, and depth estimation algorithms are pre-configured with region-specific parameters. This preparation allows the system to quickly process new images without retraining or reconfiguration, significantly reducing real-time processing time.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If traditional single-modality imaging is used, then processing efficiency is maintained, but transparent objects appear as voids or inaccuracies in depth maps

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidaccuracy of depth map
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent creates a composite imaging system that combines multiple data types (visible light images, thermal images, and depth sensor data) into a unified representation. The machine learning classifiers analyze composite patterns across different modalities to identify transparent objects, while the depth estimation algorithm synthesizes composite depth information. This composite approach maintains processing efficiency by using optimized algorithms for each data type while achieving reliable depth map accuracy that single-modality systems cannot achieve.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentEP4004808B1Identification of transparent objects from image discrepancies
Publication Date: 2025.12.17 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4004808B1 patent drawingFigure 1
  • EP4004808B1 patent drawingFigure 2
  • EP4004808B1 patent drawingFigure 3

AI summary

A computing system is provided. The computing system includes a visible light camera, a thermal camera, and a processor with associated storage. The processor is configured to execute instructions stored in the storage to receive, from the visible light camera, a visible light image for a frame of a scene and receive, from the thermal camera, a thermal image for the frame of the scene. The processor is configured to detect image discrepancies between the visible light image and the thermal image and, based on the detected image discrepancies, determine a presence of a transparent object in the scene. The processor is configured to, based on the detected image discrepancies, output an identification of at least one location in the scene that is associated with the transparent object.