3D Face Reconstruction Model Training via Stylized Map Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current three-dimensional face reconstruction technologies face challenges in accurately reconstructing faces at a low cost, particularly in cross-style scenarios where rich personalized features are required, and existing methods often necessitate costly sample labeling.

Innovation Solution

A method and apparatus for training a three-dimensional face reconstruction model by acquiring sample face images and their stylized maps, transforming them into a camera coordinate system, and rendering them to optimize the model training process, eliminating the need for face keypoint labeling and reducing costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional three-dimensional face reconstruction methods are used, then face reconstruction can be achieved, but costly sample labeling (face keypoint labeling) is required

Engineering Contradiction:
Improveface reconstruction accuracyVSAvoidlabeling cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent extracts and removes the face keypoint labeling step from the traditional reconstruction pipeline. By using stylized face maps as alternative supervision signals, the method eliminates the need for expensive manual face keypoint annotation while maintaining reconstruction accuracy through the rendered map comparison approach.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a rendered copy of the three-dimensional stylized face image and compares it with the actual stylized face map. This copying approach allows the model to learn from the rendered visualization without requiring direct access to or labeling of ground truth face keypoints, thereby reducing annotation costs.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If cross-style three-dimensional face reconstruction is implemented to meet diverse needs, then personalized features are enriched, but the complexity of accurate reconstruction increases

Engineering Contradiction:
Improvecross-style reconstruction capabilityVSAvoidreconstruction complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent employs a universal training framework that handles multiple styles simultaneously. The model is trained on diverse stylized face maps representing different styles (cartoon, anime, realistic, etc.), enabling it to generalize across styles without requiring style-specific reconstruction pipelines, thus managing complexity while maintaining versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent leverages parameter changes in the rendering process by adjusting camera coordinates and transformation parameters to generate different views and styles. By controlling rendering parameters rather than model architecture complexity, the system achieves cross-style capability with manageable computational requirements.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If three-dimensional stylized face images are constructed accurately without face keypoint labeling, then labeling costs are reduced, but new methods for model training are required

Engineering Contradiction:
Improvelabeling costVSAvoidtraining process complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by pre-processing sample face images into stylized face maps and pre-defining rendering configurations before model training. This preparation work creates a ready-to-use training dataset that eliminates the need for face keypoint labeling during the actual training process, reducing labeling costs while organizing the complexity upfront.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a rendered map as an intermediary between the three-dimensional model and the stylized face map supervision signal. This intermediary allows the model to be trained without direct access to face keypoints, as the rendered map serves as a bridge that connects the 3D reconstruction output with the 2D stylized target for comparison and loss computation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12260492B2Method and apparatus for training a three-dimensional face reconstruction model and method and apparatus for generating a three-dimensional face image
Publication Date: 2025.03.25 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12260492B2 patent drawing
  • US12260492B2 patent drawing
  • US12260492B2 patent drawing

AI summary

A method for training a three-dimensional face reconstruction model includes inputting an acquired sample face image into a three-dimensional face reconstruction model to obtain a coordinate transformation parameter and a face parameter of the sample face image; determining the three-dimensional stylized face image of the sample face image according to the face parameter of the sample face image and the acquired stylized face map of the sample face image; transforming the three-dimensional stylized face image of the sample face image into a camera coordinate system based on the coordinate transformation parameter, and rendering the transformed three-dimensional stylized face image to obtain a rendered map; and training the three-dimensional face reconstruction model according to the rendered map and the stylized face map of the sample face image.