Free Viewpoint Image Synthesis Using Residual Data Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for synthesizing free viewpoint images either result in distortion due to mismatch between the predefined projection surface and the actual three-dimensional structure, or require costly three-dimensional sensing devices like LiDAR to reduce distortion.
Innovation Solution
A computer-implemented image processing method that acquires multiple captured images from cameras, estimates projection surface residual data using machine learning, and maps these images onto a display projection surface to synthesize free viewpoint images, without relying on three-dimensional sensing devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a bowl-shaped predefined projection surface is used to synthesize free viewpoint images, then the synthesis process is simple and fast, but image distortion occurs due to mismatch between the predefined surface and actual three-dimensional structure
Solution Approach 1:
The system performs preliminary action by pre-defining a bowl-shaped projection surface and pre-training a neural network model with projection surface residual data captured from multiple viewpoints. This allows the system to quickly adapt to new scenes without real-time 3D sensing, maintaining fast synthesis speed while improving accuracy through pre-computed correction data.
Solution Approach 2:
The invention creates a copy of the actual projection surface characteristics by capturing projection surface residual data from multiple cameras at different viewpoints and storing this data in advance. This copied data is then used to correct image mapping without requiring real-time 3D sensing, thus maintaining simplicity while improving accuracy.
2Manufacturing precision
If three-dimensional sensing devices like LiDAR are used to calculate the projection surface, then image distortion is reduced, but system cost and complexity increase significantly
Solution Approach 1:
The invention introduces projection surface residual data as an intermediary that bridges the gap between simple predefined projection surfaces and accurate actual surfaces. This intermediary data, captured by standard cameras rather than expensive 3D sensors, enables accurate image mapping without requiring complex LiDAR devices or point cloud processing.
Solution Approach 2:
The system replaces the mechanical/optical 3D sensing mechanism (LiDAR) with a computational approach using standard cameras and neural networks. Instead of using active 3D sensing hardware to measure depth, the system uses passive imaging with multiple cameras and processes the data computationally through pre-trained models, eliminating the need for expensive 3D sensing devices.
3Manufacturing precision
If multiple cameras are used to capture images for free viewpoint synthesis, then image quality and coverage improve, but data processing complexity and time increase
Solution Approach 1:
The system performs preliminary action by pre-training neural network models using projection surface residual data captured from multiple cameras at various viewpoints. This pre-processing of multi-camera data into training sets allows the system to handle complex multi-camera inputs during inference without real-time processing bottlenecks, maintaining both high image quality and efficient processing.
Data Source
AI summary
A computer-implemented image processing method of synthesizing a free viewpoint image on a display projection surface from plurality of captured images, the method including: acquiring the plurality of captured images with a plurality of respective cameras; estimating projection surface residual data by machine learning using the plurality of captured images and viewpoint data as inputs, the projection surface residual data representing a difference between a bowl-shaped predefined projection surface and the display projection surface; and acquiring the free viewpoint image by mapping the plurality of captured images onto the display projection surface using information about the predefined projection surface, the projection surface residual data, and the viewpoint data.


