3D Model Generation via Point Cloud and Silhouette Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating Free-Viewpoint Video (FVV) using real-world content are limited by the need for complex and expensive camera setups, professional studios, and significant computing resources, often resulting in accuracy issues and restricted navigation ranges.
Innovation Solution
A method that combines object point clouds and shape volumes from silhouette information to generate accurate three-dimensional models using a sparse setup of image capturing devices, such as RGB cameras or mobile phones, without IR sensing, enabling dynamic and static 3D model creation for FVV, AR, and VR applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a sparse setup of image capturing devices is used, then device complexity and cost are reduced, but manufacturing precision and measurement precision of the 3D model deteriorate
Solution Approach 1:
The patent combines two different 3D reconstruction approaches (point cloud-based reconstruction and silhouette-based reconstruction) into a hybrid system. The point cloud method provides accurate geometric detail while the silhouette method ensures complete volume coverage, particularly for occluded regions. This merging allows the system to achieve high 3D model accuracy using a sparse camera setup without requiring complex professional studio equipment.
Solution Approach 2:
The patent creates a composite 3D model by integrating results from multiple reconstruction methods. The final 3D model combines the geometric precision from point cloud reconstruction with the volumetric completeness from silhouette-based reconstruction, producing a hybrid model that achieves high accuracy equivalent to dense camera setups while using only sparse imaging devices.
2Manufacturing precision
If professional high-end cameras and studio setups are used, then manufacturing precision of the 3D model is improved, but device complexity and cost increase
Solution Approach 1:
The patent enables consumer-grade cameras to achieve professional-quality 3D reconstruction results by implementing sophisticated software algorithms. The system copies the functional capability of professional studio setups through computational methods, including multi-view stereo reconstruction and silhouette-based volume synthesis, allowing standard RGB cameras to produce accurate 3D models without requiring expensive specialized equipment.
Solution Approach 2:
The patent replaces the mechanical/physical complexity of professional studio camera rigs with computational algorithms. Instead of relying on dense physical camera arrangements and specialized optical equipment, the system uses software-based reconstruction techniques including point cloud processing, silhouette extraction, and volumetric rendering to achieve equivalent or superior 3D model accuracy.
3Device complexity
If image-based techniques with limited cameras are used, then device complexity is reduced, but the navigation range in FVV is limited
Solution Approach 1:
The patent transitions from 2D image-based rendering to 3D volumetric rendering by constructing complete 3D models of the scene. By creating volumetric representations using silhouette-based reconstruction and point cloud integration, the system enables navigation in three-dimensional space rather than being constrained to fixed camera viewpoints, significantly expanding the navigation range in free-viewpoint video applications.
Solution Approach 2:
The patent creates a universal 3D model representation that can serve multiple functions: it enables free-viewpoint video navigation, supports virtual reality applications, allows arbitrary camera position rendering, and provides complete scene reconstruction. This multi-functional 3D model eliminates the navigation limitations of traditional image-based techniques while maintaining compatibility with various application requirements.
Data Source
AI summary
The method comprising providing a plurality of images of a scene captured by a plurality of image capturing devices (101); providing silhouette information of at least one object in the scene (102); generating a point cloud for the scene in 3D space using the plurality of images (103); extracting an object point cloud from the generated point cloud, the object point cloud being a point cloud associated with the at least one object in the scene (104); estimating a 3D shape volume of the at least one object from the silhouette information (105); and combining the object point cloud and the shape volume of the at least one object to generate a three-dimensional model (106). An apparatus for generating a 3D model, and a computer readable medium for generating the 3D model.


