Multiview Multiscale MPI Generation for Variable View Geometry
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for generating multiplane image (MPI) representations from posed images are limited by a fixed number and resolution of views, and fail to accurately estimate scene geometry for variable numbers of views with arbitrary geometry, lacking scalability and generalization.
Innovation Solution
A method and device utilizing a processor to operate three modules: view feature extraction, scene feature extraction, and color score extraction, with a multiscale architecture that processes input images at different resolutions using shared parameters, enabling the generation of an MPI with an alpha component and normalized color component, and allowing for variable input views and resolutions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing systems use a fixed number and resolution of views for MPI generation, then the system structure is simple and predetermined, but the system lacks adaptability to variable numbers of views and arbitrary scene geometry
Solution Approach 1:
The system dynamically adjusts the number of views and their resolutions based on input requirements rather than being fixed. The MPI generation process accepts variable numbers of posed images with different resolutions, and the network architecture adapts its processing accordingly, allowing the system to handle arbitrary scene geometries and view configurations.
Solution Approach 2:
The proposed system serves multiple functions: it can process variable numbers of views, handle different resolutions, work with arbitrary scene geometries, and generate accurate MPI representations. The unified network architecture performs view feature extraction, scene feature extraction, and color score extraction in an integrated manner, making the system versatile across different input configurations.
2Adaptability or versatility
If existing systems use predetermined view configurations, then the system is easier to operate, but the system lacks generalization to arbitrary scene geometry
Solution Approach 1:
The system changes key parameters including the number of views, their resolutions, and scene geometry characteristics without requiring reconfiguration. The network processes variable input parameters (number of posed images, their resolutions) and adapts its feature extraction and MPI generation accordingly, enabling generalization to arbitrary geometries while maintaining ease of operation through automated parameter adaptation.
3Adaptability or versatility
If the system processes variable numbers of views at variable resolutions, then the adaptability improves, but the computational complexity increases
Solution Approach 1:
The processing is segmented into three distinct modules: view feature extraction, scene feature extraction, and color score extraction. Each module handles specific aspects of the variable input data independently, then combines their results. This segmentation allows the system to manage variable numbers of views and resolutions in a structured manner, reducing overall processing complexity through modular organization.
4Manufacturing precision
If existing systems use fixed resolution processing, then the processing speed is maintained, but the manufacturing precision of MPI representation deteriorates for variable resolutions
Solution Approach 1:
The system adds the dimension of variable resolution processing to the MPI generation pipeline. Instead of processing all views at a fixed resolution, the system accepts and processes images at their native resolutions, then integrates them in the feature extraction and MPI generation processes. This multi-resolution approach improves MPI representation accuracy for variable resolutions while maintaining processing efficiency through the streamlined three-module architecture.
Data Source
AI summary
The present disclosure relates to method and apparatus for the estimation of an accurate multiplane image from a set of posed images. The accuracy of an MPI representation is measured by the performance of the views generated by it through inverse homography sampling and alpha compositing. A variable number of posed images at a possibly different resolution representative of scenes with arbitrary geometry are used as input of the system. A first level of three modules, possibly trained based on deep learning methods, is built for extracting view and scene features and color scores. Numerous levels of the same three modules may be built and organized in a structure where a level deals with down-scaled version of the posed images.


