Stylized 3D Mesh Generation via Surface Normal Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods struggle to accurately generate a 3D mesh from stylized face images, particularly due to the ill-posed nature of 3D shape reconstruction and the lack of ground-truth 3D stylized face data, which limits scalability and accuracy.
Innovation Solution
The proposed solution involves a device and method that utilize a first neural network to generate a per-pixel feature vector from a 2D input image and a target style, and a second neural network to generate a surface normal map corresponding to a 3D mesh, allowing for the creation of a stylized 3D mesh that faithfully represents the characteristics of the input image.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If self-supervised approach is used to obtain 3D shape from GAN-generated images, then 3D synthesis with consistency is achieved, but accurate 3D reconstruction is not accomplished and surfaces become smooth or noisy around facial components
Solution Approach 1:
The method performs preliminary 3D mesh generation from the input image before applying stylization. By first establishing an accurate 3D geometric foundation using the 3D-aware GAN, and then applying style transfer in the latent space, the system preserves geometric accuracy while achieving stylization, avoiding the smooth/noisy surface problems of direct self-supervised approaches
Solution Approach 2:
The system separates the 3D reconstruction task from the stylization task. The 3D-aware GAN component handles accurate geometric reconstruction, while the style transfer component handles aesthetic transformation. This segmentation allows each component to optimize for its specific function without compromising the other
2Manufacturing precision
If manual construction of 3D caricature dataset by skilled artists is used, then accurate stylized 3D shapes are obtained, but the method is not scalable
Solution Approach 1:
The system employs self-supervised learning where the 3D-aware GAN automatically learns to generate accurate 3D shapes from 2D images without requiring manually annotated 3D ground truth data. The network uses its own generated 3D meshes as training targets, enabling automatic dataset generation at scale without skilled artists
Solution Approach 2:
The patent replaces the mechanical process of manual sculpting by artists with an automated neural network-based 3D reconstruction system. The 3D-aware GAN automatically infers 3D geometry from 2D images, substituting human artistic labor with machine learning, thereby achieving both accuracy and scalability
3Adaptability or versatility
If standard 3D reconstruction methods are applied to stylized face images, then 3D mesh generation is attempted, but the ill-posed nature of reconstruction and lack of ground-truth data prevent accurate results
Solution Approach 1:
The system changes the parameter space by operating in the latent space of the GAN rather than directly in pixel or mesh space. By learning the mapping between 2D image features and 3D mesh parameters in the latent space, the system can accurately reconstruct 3D shapes from stylized images that would be ill-posed in traditional reconstruction approaches
Solution Approach 2:
The patent introduces the GAN latent space as an intermediary between 2D images and 3D meshes. This intermediate representation captures both geometric and stylistic information, enabling accurate 3D reconstruction from stylized images by serving as a bridge that resolves the ill-posed nature of direct reconstruction
Data Source
AI summary
A stylized 3D mesh generation device may comprise: a memory for storing at least one or more instructions; a processor for executing the at least one or more instructions; a first neural network for generating a per-pixel feature vector based on a 2D input image and a target style; and a second neural network for generating a surface normal map corresponding to a 3D mesh of the 2D input image based on the per-pixel feature vector, wherein the processor generates a stylized 3D mesh of the 2D input image based on the surface normal map, and the second neural network generates and outputs the surface normal map applied with the target style based on the per-pixel feature vector.


