Text-Guided 3D Texture Generation With Multi-View Consistency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating 3D textures suffer from inconsistencies and artifacts when combining 2D textures from different views of a 3D model, leading to issues like over-saturation and noticeable seams.

Innovation Solution

A method and system that utilize a pre-trained 2D image generation diffusion model to progressively generate a 3D texture directly, using attention-guided sampling and multi-conditioned classifier-free guidance to ensure consistency across views, refining noise estimation at each step to produce a high-quality, view-consistent texture map.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing methods combine 2D textures from different views to generate 3D textures, then the 3D texture can be created from multiple perspectives, but inconsistencies and artifacts (over-saturation, visible seams) occur in the generated texture

Engineering Contradiction:
Improvemulti-view texture generationVSAvoidtexture consistency
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediary consistency verification and adjustment mechanism that mediates between different 2D view textures. The system processes textures from multiple views through a unified framework that verifies and adjusts consistency across views, preventing artifacts and seams while maintaining the ability to generate textures from multiple perspectives.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements a feedback mechanism where the generated 3D texture is continuously verified against consistency criteria across different views. The system provides feedback loops that detect and correct inconsistencies, over-saturation, and seam artifacts by adjusting the texture generation process based on multi-view validation results.

Inventive Principle:
Principle #23Feedback

2Productivity

If traditional 3D texture generation methods are used, then the process is simpler and faster, but the generated textures contain artifacts and lack realism

Engineering Contradiction:
Improvetexture generation speedVSAvoidtexture quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent performs preliminary consistency verification and artifact detection during the texture generation process itself, rather than as a separate post-processing step. By embedding consistency checks and correction mechanisms within the generation workflow, the system maintains high productivity while ensuring texture quality and realism are achieved during generation.

Inventive Principle:
Principle #10Preliminary action

3Shape

If existing approaches focus on geometric components of 3D assets, then the 3D structure is well-defined, but the texture components receive less attention and quality

Engineering Contradiction:
Improve3D geometric structureVSAvoidtexture quality
Core Design Contradiction:
ShapeVSManufacturing precision

Solution Approach 1:

The patent merges the treatment of geometric components and texture components into a unified 3D asset generation framework. The system processes both geometry and texture together, ensuring that texture quality receives the same level of attention and precision as geometric structure, while maintaining their respective qualities through integrated processing.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250371790A1Methods and systems for text-guided 3D texture generation
Publication Date: 2025.12.04 HUAWEI TECH CO LTD
  • US20250371790A1 patent drawing
  • US20250371790A1 patent drawing
  • US20250371790A1 patent drawing

AI summary

System, method, and computer readable medium for generating a 3D texture for a 3D object are disclosed. A 3D mesh and a text prompt for a desired texture are obtained. A sequence of texture sampling steps is performed, where each given texture sampling step includes iterating over a plurality of 2D views of the 3D mesh to generate an intermediate texture map. For a given iteration, a given 2D view and the text prompt are processed using a pre-trained 2D image generation diffusion model to fill in a portion of an intermediate texture map based on the given 2D view. A noise estimation generated by the diffusion model is refined, adding the intermediate texture map as guidance, to generate a latent variable to be inputted to a subsequent texture sampling step, enabling generation of a 3D texture, based on a text prompt, with fewer artifacts.