Invisibility Mask for Non-Visible Object Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automated vehicles face challenges in detecting obstacles in poor weather conditions and low light levels, leading to hazardous navigation conditions due to the limitations of color cameras.
Innovation Solution
A vehicle imaging system that generates a pixel-level invisibility mask for color images, allowing for the detection of non-visible objects and estimating the reliability of object detectors, by transferring mid-level features from color images to infrared images and using an invisibility map as a confidence map for sensor fusion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If color cameras are used for obstacle detection, then the system is simple and cost-effective, but detection reliability deteriorates in poor weather conditions and low light levels
Solution Approach 1:
The patent combines color camera and infrared camera sensors into a unified imaging system. The color camera captures visible light information while the infrared camera captures thermal radiation, and both images are fused through a visibility mask generation mechanism to produce enhanced obstacle detection results, particularly in low-light and poor-weather conditions.
Solution Approach 2:
The patent introduces a visibility mask as an intermediary component that mediates between color image and infrared image processing. The visibility mask, generated by comparing color and infrared image features, serves as a confidence map to guide sensor fusion and determine regions where infrared data should supplement or replace color data, thereby improving detection reliability without requiring complete system redesign.
2Reliability
If infrared sensors are added to improve detection in low light, then detection reliability improves, but device complexity increases
Solution Approach 1:
The patent uses the color image as a reference copy to guide infrared image processing. By transferring mid-level features from the color image domain to the infrared image domain, the system leverages the well-trained color object detectors to improve infrared detection performance, effectively copying successful detection patterns across modalities.
Solution Approach 2:
The patent changes the parameter space from raw pixel values to mid-level features extracted by neural networks. By operating in the feature space rather than direct pixel space, the system can transfer knowledge between color and infrared domains more effectively, improving detection reliability while managing computational complexity through efficient feature-level processing.
3Measurement precision
If manual labelling is used to create invisibility masks, then mask accuracy improves, but time consumption increases
Solution Approach 1:
The patent implements a self-service mechanism where the system automatically generates visibility masks by comparing color and infrared image features without requiring manual annotation. The automated mask generation uses neural network feature extraction and comparison to produce confidence maps in real-time, eliminating the time-consuming manual labelling process while maintaining acceptable accuracy for sensor fusion applications.
Data Source
AI summary
A computer-implemented method of detecting one or more objects in a driving environment located externally to a vehicle, and a vehicle imaging system configured to detect one or more objects. The computer-implemented method includes training a first neural network to detect objects in a color video stream, the first neural network having a plurality of mid-level color features at a plurality of scales, and training a second neural network, operatively coupled to color neural network and an infrared video stream, to match, at the plurality of scales, mid-level infrared features of the second neural network to mid-level color features of the first neural network. A pixel-level invisibility map is then generated from the color video stream and the infrared video stream by determining differences, at each of the plurality of scales, between mid-level color features at the first neural network and mid-level infrared features at the second infrared neural network, and coupling the result to a fusing function.


