Multi-Sensor Target Detection Using Depth Maps and Pseudo-Point Clouds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing target detection methods are sensitive to light and weather conditions, leading to poor performance in harsh environments, and face challenges in integrating multi-modal information from visible cameras, infrared cameras, and LiDARs.

Innovation Solution

A method that synchronizes and preprocesses data from visible cameras, infrared cameras, and LiDARs to generate depth maps, pseudo-point clouds, and aggregated point clouds, using a backbone network to extract multi-source features and fuse information for accurate target detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple sensors (visible camera, infrared camera, LiDAR) are introduced to improve robustness in different lighting conditions, then the robustness and detection accuracy improve, but the system complexity and data integration difficulty increase

Engineering Contradiction:
Improverobustness of target detectionVSAvoidcomplexity of multi-sensor integration
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines visible camera, infrared camera, and LiDAR sensors into a unified target detection system. The multi-sensor fusion architecture integrates data from all three sensors to achieve complementary advantages: visible camera for daytime/good lighting, infrared camera for nighttime/low-light, and LiDAR for distance and shape information. This merging approach resolves the contradiction by achieving improved robustness through sensor combination while managing complexity through systematic data fusion procedures.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a multi-functional detection system where each sensor type serves multiple purposes. The visible camera provides both color/texture information and depth estimation through monocular depth prediction. The infrared camera contributes thermal distribution characteristics and works in low-light conditions. The LiDAR provides accurate distance and shape information. This multi-functionality allows the system to maintain robustness across different lighting conditions while avoiding the need for separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple sensors are used to acquire multi-dimensional target information, then the target detection accuracy improves, but the difficulty of data fusion and synchronization increases

Engineering Contradiction:
Improvetarget detection accuracyVSAvoiddifficulty of data fusion
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent performs time synchronization and data preprocessing as preliminary actions before main detection. The system synchronizes timestamps from all sensors to ensure temporal alignment, performs extrinsic calibration to establish coordinate transformations between sensors, and generates depth maps from visible and infrared images before fusion. These preliminary actions resolve the data fusion difficulty by preparing standardized, synchronized inputs that can be efficiently integrated in the subsequent detection pipeline.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces depth maps as an intermediary representation that bridges image data from visible and infrared cameras with point cloud data from LiDAR. The depth prediction module generates depth information from 2D images, which then serves as a common spatial reference for fusing with LiDAR point clouds. This intermediary approach simplifies the fusion process by providing a unified spatial framework that accommodates different sensor modalities.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If visible camera, infrared camera, and LiDAR are combined to cover different lighting conditions, then the adaptability to varying environments improves, but the system cost and processing complexity increase

Engineering Contradiction:
Improveadaptability to lighting conditionsVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic sensor utilization where the system adapts its processing based on environmental conditions. The multi-sensor fusion architecture dynamically weights and integrates data from visible camera, infrared camera, and LiDAR according to lighting conditions. In good lighting, visible camera data is weighted higher; in low-light, infrared data becomes more prominent; LiDAR provides consistent geometric information across all conditions. This dynamic adaptation achieves environmental versatility while managing processing complexity through conditional data prioritization.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240355105A1Methods for target detection based on visible cameras, infrared cameras, and lidars
Publication Date: 2024.10.24 DONGHAI LAB
  • US20240355105A1 patent drawing
  • US20240355105A1 patent drawing
  • US20240355105A1 patent drawing

AI summary

A method for target detection based on a visible camera, an infrared camera, and a LiDAR is provided. The method designates visible light images, infrared images, and LiDAR point clouds, which are synchronously acquired, as inputs, and generates an input pseudo-point cloud using visible light images and infrared images, to realize alignment of multimodal information in a three-dimensional space and fusion feature extraction. Then the method adopts a cascade strategy to output more accurate target detection results step by step. In the present disclosure, different characteristics of multi-sensors are complemented, which improves and extends traditional target detection algorithms, improves the accuracy and robustness of target detection, and realizes multi-category target detection in a road scene.