3D Morphable Model Facial Image Generation via Optical Flow
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing face image generation technologies relying on generative adversarial networks suffer from high model complexity, over-fitting, and inability to produce natural and realistic images, limiting personalized synthesis.
Innovation Solution
The method employs a 3D morphable model (3DMM) to generate an initial optical flow map, which is then processed using a convolutional neural network to retain the contour and pose/expression of a target face image, allowing for personalized and realistic synthesis without relying on a single large network.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a generative adversarial network is used for face image generation, then the synthesized face image can be generated, but the model complexity is high and the network parameters are large
Solution Approach 1:
The patent segments the face image generation task into multiple components: 3D morphable model for facial structure, optical flow network for motion estimation, and refinement network for detail enhancement. This divides the complex generative adversarial network into smaller, specialized modules that can be trained and processed independently, reducing overall model complexity while maintaining generation capability.
Solution Approach 2:
The patent introduces an optical flow map as an intermediary element between the input face image and the synthesized output. The optical flow network generates this intermediate representation that captures motion and deformation patterns, which then guides the refinement network. This intermediary approach avoids the need for a single complex direct mapping network.
2Reliability
If a generative adversarial network is used for face image generation, then the synthesized face image can be generated, but the training effect is poor and over-fitting occurs
Solution Approach 1:
The patent changes the parameter representation from direct pixel-to-pixel mapping in a large GAN to a structured representation using 3D morphable model parameters and optical flow fields. This parameterization provides better inductive bias and regularization, reducing over-fitting while improving generalization to new face images and expressions.
Solution Approach 2:
The patent performs preliminary decomposition of the face image into structural components using the 3D morphable model before generating the synthesized image. By pre-processing the input to extract meaningful facial parameters and optical flow patterns, the system avoids learning these representations from scratch during training, improving training efficiency and reducing over-fitting.
3Reliability
If a generative adversarial network is used for face image generation, then the synthesized face image can be generated, but the synthesized face image is not natural and realistic enough
Solution Approach 1:
The patent employs dynamic optical flow computation to capture temporal variations and motion patterns in face images. The optical flow network continuously estimates pixel-level motion fields that adapt to different expressions and poses, enabling the synthesis of natural dynamic facial movements rather than static transformations. This dynamic approach significantly improves the realism of synthesized face sequences.
Solution Approach 2:
The patent applies local refinement through the refinement network that operates on specific regions of the face image guided by optical flow patterns. Different parts of the face (eyes, mouth, eyebrows) are refined with locally-adapted transformations that preserve fine details and natural textures. This local quality enhancement ensures that synthesized images maintain realistic facial features rather than appearing globally smoothed or distorted.
4Reliability
If a generative adversarial network is used for face image generation, then the synthesized face image can be generated, but personalized face image synthesis cannot be achieved
Solution Approach 1:
The patent creates a universal framework based on 3D morphable models that can represent diverse facial characteristics, expressions, and identities through a unified parameter space. The same optical flow network and refinement pipeline can process any input face image regardless of identity or expression type, enabling personalized synthesis by simply changing the input parameters rather than requiring identity-specific models. This multi-functional approach achieves adaptability across different persons and conditions.
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
Disclosed is a facial image generation method, comprising: according to a first facial image in a first reference element, determining a three-dimensional face morphable model, corresponding to the first facial image, to be a first model; according to a second reference element, determining a three-dimensional face morphable model, corresponding thereto, to be a second model; according to the first model and the second model, determining an initial optical flow diagram corresponding to the first facial image, and deforming the first facial image according to the initial optical flow diagram to obtain an initial deformation diagram; according to the first facial image, and the initial optical flow diagram and initial deformation diagram corresponding thereto, obtaining an optical flow increment graph and a visibility probability graph by means of a convolutional neural network; and generating a target facial image according to the first facial image, and the initial optical flow diagram, optical flow increment graph and visibility probability graph corresponding thereto. The method realizes parameterization control, and also retains detail information of an original image based on an optical flow, such that a generated image is realistic and natural. Further disclosed are a corresponding apparatus, a device and a medium.