Text-to-3D Avatar Generation via Stylized Image Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-image synthesis models primarily generate two-dimensional images and lack the capability to produce high-quality three-dimensional avatars efficiently.

Innovation Solution

The method involves stylizing a dataset of images based on a user-input text prompt using a Stable Diffusion model, and then utilizing an efficient geometry-aware 3D generative adversarial network (EG3D) model to produce three-dimensional avatars.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If current text-to-image synthesis models are used, then two-dimensional images can be generated, but three-dimensional avatar generation capability is lacking

Engineering Contradiction:
Improve3D avatar generation capabilityVSAvoidgeneration quality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent extends 2D image generation capabilities to 3D avatar generation by introducing a novel architecture that processes images in three-dimensional space. The system uses a 3D variational autoencoder to transform 2D input images into 3D representations, enabling the model to generate avatars with depth, volume, and spatial structure rather than flat two-dimensional outputs.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces a pose estimator and a 3D pose encoder as intermediary components between the input image and the final 3D avatar generation. The pose estimator extracts pose information from input images, and the 3D pose encoder transforms this pose data into a format suitable for guiding 3D avatar synthesis, serving as a bridge that enables accurate 3D reconstruction from 2D inputs.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If traditional 3D generation methods are used, then three-dimensional structure can be achieved, but efficiency and quality are insufficient

Engineering Contradiction:
Improvegeneration efficiencyVSAvoidavatar quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent performs preliminary pose estimation and feature extraction from input images before proceeding to 3D avatar generation. By pre-processing the input data to extract pose information and key features, the system prepares optimized input representations that accelerate the subsequent 3D synthesis process while maintaining high output quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the input image data into different parameter spaces, converting 2D image pixels into 3D spatial parameters and pose parameters. This parameter transformation enables the model to work more efficiently in the 3D domain, where geometric relationships can be exploited to improve both generation speed and avatar quality.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12340480B2Text-to-3D avatars
Publication Date: 2025.06.24 LEMON INC(GB)
  • US12340480B2 patent drawing
  • US12340480B2 patent drawing
  • US12340480B2 patent drawing

AI summary

Three-dimensional (3D) avatars may be produced by stylizing a dataset of images based on a user-input text prompt input to a stable diffusion model, and using the output stylized dataset of images to train an efficient geometry-aware 3D generative adversarial network (EG3D) model.