Multi-Sensor Facial Un-Distortion Using Predicted Depth Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Perspective distortion in selfie images due to arm-length capture distances results in unappealing facial appearances and unsatisfactory user experience, with existing techniques requiring additional depth sensors and complex algorithms.
Innovation Solution
A system and method using multiple imaging sensors and an end-to-end differentiable deep learning pipeline to correct perspective distortion without a pre-generated depth map, employing landmark alignment, depth map prediction, warp field generation, and inpainting to generate undistorted images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If multiple imaging sensors and deep learning algorithms are used for facial un-distortion, then facial distortion correction quality is improved, but device complexity and computational requirements increase
Solution Approach 1:
The patent creates virtual depth information by synthesizing a depth map from multiple 2D images captured by different imaging sensors. Instead of requiring physical depth sensors, the system copies spatial relationships from multiple image perspectives to reconstruct 3D depth information, which is then used for accurate facial un-distortion while keeping the device structure simple.
Solution Approach 2:
The patent replaces the need for physical depth sensing hardware (mechanical/optical systems) with computational algorithms. Multiple 2D images are processed through neural networks to generate depth maps and warp fields, substituting complex hardware depth sensors with software-based computational photography approaches.
2Measurement precision
If depth sensors and complex algorithms are used for perspective distortion correction, then distortion correction accuracy is improved, but ease of manufacture and device simplicity deteriorate
Solution Approach 1:
The patent makes existing imaging sensors perform multiple functions: they capture both 2D image data and serve as the basis for generating depth information through computational methods. The same camera modules used for standard photography are leveraged for depth estimation and facial un-distortion, eliminating the need for separate depth sensing hardware and simplifying manufacturing.
Solution Approach 2:
The system uses its own multiple imaging sensors to generate the depth information it needs for distortion correction, rather than relying on external or dedicated depth sensors. The imaging sensors themselves provide the raw data that is processed into depth maps, making the system self-sufficient and easier to manufacture.
3Reliability
If pre-generated depth maps are required for distortion correction, then correction reliability is improved, but device complexity and processing requirements increase
Solution Approach 1:
The patent performs preliminary depth map generation and warp field computation in advance, before the actual distortion correction is applied to the final image. By pre-processing the multiple images to create depth information and transformation fields, the system ensures reliable correction results while organizing the computational workflow to improve efficiency and reduce real-time processing requirements.
Data Source
AI summary
A method includes aligning landmark points between multiple distorted images to generate multiple aligned images, where the multiple distorted images exhibit perspective distortion in at least one face appearing in the multiple distorted images. The method also includes predicting a depth map using a disparity estimation neural network that receives the multiple aligned images as input. The method further includes generating a warp field using a selected one of the multiple aligned images. The method also includes performing a two-dimensional (2D) image projection on the selected aligned image using the depth map and the warp field to generate an undistorted image. In addition, the method includes filling in one or more missing pixels in the undistorted image using an inpainting neural network to generate a final undistorted image.


