Dichroic Mirror Optical Alignment for 2D-3D Sensor Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automated object detection systems face challenges in accurately recognizing objects in 3D point clouds due to low resolution, missing data, poor signal-to-noise ratio, and the lack of comprehensive training sets, while 2D and 3D imaging hardware calibration is computationally expensive and non-ideal.
Innovation Solution
Simultaneously capturing 2D images from a camera and 3D point clouds using a time-of-flight sensor with shared optics, aligning the images with a dichroic mirror, and enhancing the 2D images with 3D data to create an e-RGB image that includes true reflectivity, scale, and occlusion information, which can be used for object recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 3D point cloud data is used for object recognition, then true distance and reflectivity information is obtained, but resolution, frame rate, and signal-to-noise ratio deteriorate
Solution Approach 1:
The patent combines 2D camera images and 3D LIDAR point cloud data into a unified representation. The 2D image provides high-resolution visual information while the 3D data contributes depth and distance measurements. By merging these complementary data sources, the system achieves both high resolution and accurate distance measurement without the limitations of using either modality alone.
Solution Approach 2:
The patent creates a composite data structure that integrates 2D image pixels with 3D point cloud information. This composite representation allows the system to leverage the high resolution of 2D imaging and the accurate depth measurement of 3D sensing simultaneously, effectively creating a hybrid data format that combines the strengths of both sensing modalities.
2Adaptability or versatility
If separate 2D and 3D imaging hardware is used, then both 2D images and 3D point clouds are captured, but calibration complexity and computational cost increase
Solution Approach 1:
The patent implements a nested architecture where the 3D LIDAR system is positioned within the optical path of the 2D camera. The LIDAR shares the camera's optical elements and housing, creating a compact integrated unit. This nesting eliminates the need for separate mounting and alignment of independent 2D and 3D imaging systems, significantly reducing calibration complexity while maintaining dual imaging capability.
Solution Approach 2:
The patent designs the imaging system with universal optical components that serve both 2D and 3D sensing functions. The shared optics and processing architecture allow the same hardware to perform multiple functions (2D imaging, 3D ranging, depth mapping) without requiring separate dedicated systems, thereby reducing overall device complexity and calibration requirements.
3Extent of automation
If 3D point cloud data is processed by AI systems, then object recognition is performed, but computational cost increases and training sets are insufficient
Solution Approach 1:
The patent performs preliminary processing of the 3D point cloud data to generate depth maps and segmented region information before feeding data to the AI system. This preprocessing step organizes the raw 3D data into structured formats that are more suitable for AI processing, reducing the computational burden on the neural network while maintaining automated object recognition capability.
Solution Approach 2:
The patent introduces intermediate processing steps that act as mediators between the raw 3D point cloud data and the AI system. These intermediate representations (such as depth maps, normalized point clouds, and segmented regions) bridge the gap between the two data modalities, enabling more efficient AI processing by presenting data in a format that leverages both 2D and 3D information without requiring the AI system to process raw 3D point clouds directly.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The enhanced e-RGB images improve object recognition accuracy by providing true reflectivity, scale, and occlusion information, allowing existing AI algorithms trained on 2D images to identify objects more accurately without the need for additional training, and can be used in various applications like surveillance and autonomous vehicles.
Implementation Method 1
A dichroic mirror splits the incoming light towards both a 2D sensor and a 3D sensor
Implementation Method 2
by the use of LIDAR, which by timing the reflection of light, creates a set of 3D points, or voxels, called a point cloud
Implementation Method 3
by the use of LIDAR, which by timing the reflection of light
Data Source
AI summary
A device and method of object detection in a scene by combining traditional 2D visual light imaging such as pixels with 3D data such as a voxel map are described. A single lens directs image light from the scene to a dichroic mirror which then provides light to a both a 2D visible light image sensor and a 3D sensor, such as a time-of-flight sensor that uses a transmitted, modulated IR light beam, which is then synchronously demodulated to determine time of flight as well as 2D coordinates. 2D portions (non-distance) of 3D voxel image data are aligned with the 2D pixel image data such that each is responsive to the same portion of the scene. Embodiments determine true reflectivity, true scale, and image occlusion. 2D images may be enhanced by the 3D true reflectivity. Combined data may be used as training data for object detection and recognition.


