Multi-Aperture Ranging with CNN Disparity Refinement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image and video-based ranging techniques are limited in deriving range and depth data, especially in the near field, due to dependency on overlapping fields of view and mechanical jitter, and monocular depth estimation methods rely on offline-trained models with limited generalization to new scenes.
Innovation Solution
A multi-aperture image processing system using a convolutional neural network (CNN) that computes and predicts disparity between subaperture images, incorporating physics-based computations and self-supervised learning to generate depth data throughout the entire field of view, independent of training data from different systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multi-aperture approaches are used to produce rich depth data at a single instance in time, then depth accuracy is improved, but the field of view is restricted to overlapping regions between subaperture images
Solution Approach 1:
The patent combines physics-based multi-aperture disparity computation (which provides accurate depth in overlapping regions) with CNN-based monocular depth estimation (which provides depth coverage across the entire field of view). The two methods are merged into a unified system where the CNN predictions are refined using physics-based constraints from multiple subaperture images, achieving both accurate depth measurement and complete field of view coverage.
2Measurement precision
If stereo or multi-view stereo systems are used to derive depth data, then depth information is obtained for points in overlapping regions, but the minimum range is limited by the cameras' focal lengths and baseline
Solution Approach 1:
The patent introduces CNN-based monocular depth estimation as an intermediary method to bridge the gap in the near field where traditional multi-aperture methods fail. The CNN provides depth predictions for all regions including those below the minimum ranging distance, while the physics-based multi-aperture computation serves as a constraint and refinement mechanism, creating a seamless depth map across all ranges.
3Area of stationary object
If monocular depth estimation using a single camera is used, then the entire field of view can be covered, but the methods rely on offline-trained models with limited generalization to new scenes
Solution Approach 1:
The patent implements a feedback mechanism where physics-based multi-aperture disparity computations provide ground truth constraints that guide and refine the CNN-based monocular depth estimation. The system uses the reliable depth measurements from overlapping regions (where physics-based methods work) as feedback to train and validate the CNN model, improving its generalization capability to new scenes and conditions.
4Measurement precision
If video or temporal-based approaches are used for depth estimation, then structure from motion can be derived, but unwanted changes such as camera mechanical jitter and attitude variability are difficult to predict and rectify
Solution Approach 1:
The patent segments the depth estimation problem into two independent components: (1) spatial segmentation using multi-aperture subimages captured simultaneously to avoid temporal jitter issues, and (2) region-based segmentation where CNN-based monocular estimation is applied to non-overlapping regions while physics-based methods handle overlapping regions. This segmentation eliminates dependency on temporal consistency and makes the system robust to camera jitter.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The system significantly improves the quality and quantity of range and depth data, providing accurate predictions across the entire field of view and adapting to new scenes without relying on aligned ground truth depth data, enhancing monocular depth estimation and overcoming field of view limitations.
Implementation Method 1
a multi-aperture optical component having optical elements optically coupled to the main lens and configured to create a multi-aperture image set that includes a plurality of subaperture images
Implementation Method 2
an array of sensing elements, the array of sensing elements being optically coupled to the multi-aperture optical component and configured to generate signals that correspond to the at least two subaperture images
Data Source
AI summary
Embodiments of systems and methods for multi-aperture ranging are disclosed. An embodiment of an image processing system includes at least one processor and memory configured to receive a multi-aperture image set that includes a high-resolution subaperture image and a low-resolution subaperture image, wherein the high-resolution subaperture image and the low-resolution subaperture image were captured simultaneously from a camera using dissimilar focal lengths, predict a high-resolution predicted disparity map from the high-resolution subaperture image using a neural network, predict a low-resolution predicted disparity map from the low-resolution subaperture image using the neural network, and generate an integrated range map from the high-resolution and low-resolution predicted disparity maps, wherein the integrated range map includes an array of range information that corresponds to the multi-aperture image set and that is generated by overlaying common points in both the high-resolution predicted disparity map and the low-resolution predicted disparity map.


