Thermal Enhanced Semantic Segmentation for Object Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Typical computing devices using two-dimensional color cameras struggle with accurately classifying objects in a scene, especially when objects with similar color appearances are blended, leading to reduced classification accuracy.
Innovation Solution
A computing device equipped with both an RGB camera and an infrared (IR) camera captures and enhances input image data by registering IR images with visible light images, using a pixelwise semantic classifier and Gabor jet convolution to generate thermal boundary saliency images, which are then processed with an artificial neural network for improved semantic segmentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If only two-dimensional color cameras are used for object classification, then device complexity is reduced, but classification accuracy deteriorates when objects have similar color appearances
Solution Approach 1:
The patent combines visible light images from RGB cameras with thermal images from infrared cameras to create a multi-modal imaging system. This merging of different sensing modalities allows the system to distinguish objects with similar color appearances by utilizing thermal signature differences, thereby improving classification accuracy without excessive complexity increase
Solution Approach 2:
The patent transitions from two-dimensional color information to three-dimensional information by adding the thermal dimension. This dimensional expansion allows objects to be differentiated not only by color but also by temperature characteristics, resolving the limitation of color-based blending
2Measurement precision
If thermal imaging is added to improve boundary detection, then classification accuracy improves, but device complexity and cost increase
Solution Approach 1:
The patent processes thermal images through a series of segmentation operations including boundary detection, saliency map generation, and region extraction. This segmented processing approach isolates the beneficial thermal boundary information while managing computational complexity through modular processing stages
Solution Approach 2:
The patent introduces intermediate processing layers including a thermal processing network and a visible light processing network that act as mediators between the raw thermal/visible inputs and the final classification. These intermediary networks transform thermal data into meaningful boundary representations that can be integrated with visible light information
3Measurement precision
If high-resolution IR cameras are used for improved thermal imaging, then measurement precision improves, but cost and device complexity increase
Solution Approach 1:
The patent applies preliminary thermal processing operations including boundary detection and saliency enhancement to thermal images from lower-resolution cameras. By performing these preprocessing operations, the system extracts meaningful thermal features that compensate for the lower input resolution, achieving effective high-resolution thermal information without requiring expensive high-resolution IR cameras
Data Source
AI summary
Technologies for thermal enhanced semantic segmentation include a computing device having a visible light camera and an infrared camera. The computing device receives a visible light image of a scene from the visible light camera and an infrared image of the scene from the infrared camera. The computing device registers the infrared image to the visible light image to generate a registered image. Registering the infrared image may include increasing resolution of the infrared image. The computing device generates a thermal boundary saliency image based on the registered infrared image. The computing device may generate the thermal boundary saliency image by applying a Gabor jet convolution to the registered infrared image. The computing device performs semantic segmentation on the visible light image, the registered infrared image, and the thermal boundary saliency image to generate a pixelwise semantic classification of the scene. Other embodiments are described and claimed.


