3D-Aware Text-to-Image Generation Using Depth Map Guidance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image processing systems struggle to generate accurate and aesthetically pleasing images from text prompts due to the difficulty in conveying desired visual details through text, particularly in depicting scenes with specific 3D features.

Innovation Solution

An image processing system that combines 3D modeling with text-guided image generation, using a depth map generated from user-provided 3D scene geometry and text prompts to create output images that adhere to the intended scene geometry and textual guidance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If text-guided image generation is used, then versatility and ease of operation are improved, but manufacturing precision and reliability deteriorate due to difficulty in conveying visual details through text

Engineering Contradiction:
Improveease of image generationVSAvoidaccuracy of scene geometry
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediate representation (depth map or canonical view image) that mediates between the text prompt and the final image generation. This intermediary structure encodes the 3D geometric constraints, allowing the text-guided model to operate on a structured representation rather than raw text, thereby improving both ease of operation and geometric accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions from 2D image generation to 3D-aware image generation by introducing depth maps or canonical views as intermediate representations. This dimensional enhancement allows the system to capture and enforce 3D geometric constraints while maintaining the flexibility of text-guided generation, resolving the contradiction between ease of use and geometric precision.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Manufacturing precision

If 3D modeling is used to ensure geometric precision, then manufacturing precision is improved, but device complexity and ease of operation worsen due to requiring extensive expertise

Engineering Contradiction:
Improveaccuracy of scene geometryVSAvoidcomplexity of 3D modeling process
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by automatically generating depth maps or canonical view images from the 3D scene representation before the actual image generation process. This preliminary structuring of geometric information simplifies the subsequent generation step, allowing users to achieve precise 3D-aware images without manually complex 3D modeling operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system performs self-service by automatically converting the 3D scene representation into depth maps or canonical views without requiring user intervention. This automation eliminates the need for users to manually create complex 3D models, reducing the perceived complexity while maintaining geometric precision through the structured intermediate representation.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If conventional text-guided generation is used, then ease of operation is improved, but measurement precision and reliability worsen due to inability to accurately depict 3D features

Engineering Contradiction:
Improvesimplicity of text inputVSAvoidaccuracy of visual details
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary representation (depth map or canonical view) that bridges the gap between text input and image generation. This intermediary structure preserves the simplicity of text input while ensuring accurate depiction of 3D features by encoding geometric constraints in a structured format that the generation model can reliably interpret.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12592030B2Interactive three-dimension aware text-to-image generation
Publication Date: 2026.03.31 ADOBE INC
  • US12592030B2 patent drawing
  • US12592030B2 patent drawing
  • US12592030B2 patent drawing

AI summary

An image processing system is configured to receive a three-dimensional (3D) model and a text prompt that describes a scene corresponding to the 3D model. The system may then generate a depth map of the 3D model and generate an output image based on the depth map and the text prompt. The output image may depicts a view of the scene that includes textures described by the text prompt. The output image may be generated using an image generation model.