Unsupervised Multi-Scale Disparity and Optical Flow Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for dual camera stereo reconstruction face challenges in achieving accurate and dense disparity/optical flow maps, as they require ground truth disparity data for training, which is difficult to gather, and struggle with noise in high-resolution images and accuracy in low-resolution images.
Innovation Solution
An unsupervised multi-scale disparity/optical flow fusion method using a generative adversarial network (GAN) algorithm that creates stereo pairs from single images, predicts object motion, and generates fused disparity and optical flow maps without needing a ground truth dataset, improving density and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If disparity is calculated in high resolution images only, then accuracy is improved, but density deteriorates due to noise
Solution Approach 1:
The patent merges disparity maps generated from high-resolution images (for accuracy) and low-resolution images (for density) into a fused disparity map. The fusion process combines the complementary strengths of both resolutions, achieving both high accuracy and high density simultaneously, thereby resolving the contradiction between these two parameters.
Solution Approach 2:
The patent segments the image processing into multiple resolution levels, calculating disparity at both high and low resolutions separately, then fusing the results. This segmentation allows each resolution level to contribute its strengths to the final output, resolving the trade-off between accuracy and density.
2Quantity of substance
If disparity is calculated in low resolution images only, then density is improved, but accuracy deteriorates
Solution Approach 1:
The patent merges disparity maps generated from high-resolution images (for accuracy) and low-resolution images (for density) into a fused disparity map. The fusion process combines the complementary strengths of both resolutions, achieving both high accuracy and high density simultaneously, thereby resolving the contradiction between these two parameters.
Solution Approach 2:
The patent segments the image processing into multiple resolution levels, calculating disparity at both high and low resolutions separately, then fusing the results. This segmentation allows each resolution level to contribute its strengths to the final output, resolving the trade-off between accuracy and density.
3Ease of manufacture
If a normal CNN is used for fusion, then training is simplified, but ground truth data requirement makes it complex in practice
Solution Approach 1:
The patent employs a cycle-consistent adversarial network that generates synthetic ground truth data from real images through a cycle consistency mechanism. The network trains using only real image pairs, with the synthetic ground truth generated by the system itself, eliminating the need for external ground truth datasets and enabling self-service training.
Solution Approach 2:
The patent creates synthetic ground truth disparity maps by copying and transforming real image pairs through the cycle-consistent adversarial network. This copying process generates artificial ground truth data that mirrors real disparities, allowing the network to train without requiring actual ground truth measurements.
Data Source
AI summary
An apparatus includes an interface and a processor. The interface may be configured to receive pixel data from a capture device. The processor may be configured to (i) process the pixel data arranged as one or more video frames, (ii) extract features from the one or more video frames, (iii) generate fused maps for at least one of disparity and optical flow in response to the features extracted, (iv) generate regenerated image frames by performing warping on a first subset of the video frames based on (a) the fused maps and (b) first parameters, (v) perform a classification of a sample image frame based on second parameters, and (vi) update the first parameters and the second parameters in response to whether the classification is correct. The classification generally comprises indicating whether the sample image frame is one of a second subset of the video frames or one of the regenerated image frames.


