3D Scene Reconstruction Using Asynchronous Sensor Event Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for 3D scene reconstruction using asynchronous sensors, such as DVS and ATIS, face challenges due to the lack of standard 2D images, requiring artificial reconstruction and significant processing, which discards time-dependent visual information and is inefficient.
Innovation Solution
A method for 3D reconstruction using asynchronous sensors involves matching events from multiple sensors based on a cost function that includes luminance, movement, and geometric components, allowing precise sampling of time information without the need for recreating standard 2D images, using convolution cores and Gaussian variance for luminance signals and decreasing functions for movement components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If standard 2D image reconstruction methods are applied to asynchronous sensor data, then 3D reconstruction can be performed using existing algorithms, but time information is discretized and processing complexity increases
Solution Approach 1:
Instead of reconstructing 2D images first and then performing 3D reconstruction, the patent inverts the approach by directly performing 3D reconstruction on asynchronous event data without intermediate 2D image formation. This avoids the information loss inherent in discretizing continuous time events into frame-based images.
Solution Approach 2:
The patent extracts and utilizes the temporal information directly from asynchronous events by computing time differences between corresponding events in stereo pairs. This extracted time information is then used to calculate depth without needing to reconstruct standard 2D images, thereby preserving continuous time data.
2Adaptability or versatility
If asynchronous events are processed to create standard 2D images, then compatibility with existing 3D reconstruction methods is achieved, but processing time and computational resources increase
Solution Approach 1:
The patent replaces the mechanical process of constructing 2D images from asynchronous events with a direct computational approach that processes events in their native asynchronous format. By substituting the image reconstruction step with direct depth calculation from time differences, processing time is reduced while maintaining compatibility with 3D reconstruction goals.
3Measurement precision
If asynchronous sensor data is used directly for 3D reconstruction, then time information precision is improved, but the complexity of handling non-standard data increases
Solution Approach 1:
The patent changes the parameter representation from standard 2D image coordinates to asynchronous event parameters (pixel address and timestamp). By working directly with these parameters and computing time differences, the method achieves high time precision while managing complexity through a streamlined processing approach that avoids intermediate image construction.
4Device complexity
If standard camera sampling is used, then simple image processing is achieved, but time dynamics range and latency are limited
Solution Approach 1:
The patent embraces the dynamic nature of asynchronous sensors by processing events as they occur rather than waiting for fixed sampling intervals. This dynamic processing approach maintains simplicity by working directly with event streams while dramatically expanding the time dynamics range and reducing latency compared to standard camera sampling methods.
Data Source
AI summary
The present invention concerns a method for the 3D reconstruction of a scene comprising the matching (610) of a first event from among the first asynchronous successive events of a first sensor with a second event from among the second asynchronous successive events of a second sensor depending on a minimisation (609) of a cost function (E). The cost function comprises at least one component from: —a luminance component (Ei) that depends at least on a first luminance signal (Iu) convoluted with a convolution core (gσ(t)), the luminance of said pixel depending on a difference between maximums (te−,u,te+,u) of said first signal; and a second luminance signal (Iv) convoluted with said convolution core, the luminance of said pixel depending on a difference between maximums (te−,v,te+,v) of said second signal; —a movement component (EM) depending on at least time values relative to the occurrence of events located at a distance from a pixel of the first sensor and time values relative to the occurrence of events located at a distance from a pixel of the second sensor.


