Text-Guided 3D Scene Generation With Sparse Point Clouds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating high-quality 3D scene backgrounds are inefficient, lack customization, and are not versatile enough for commercial mixed reality platforms.

Innovation Solution

A method involving obtaining a target text, generating a panoramic image, performing depth estimation to determine a sparse point cloud, and constructing a 3D scene model using multi-view information and the sparse point cloud, enhanced by a pre-trained diffusion model and 3D reconstruction techniques like NeRF and NeuS.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional methods are used for generating 3D scene backgrounds, then the process is simple, but the quality and realism of the generated scenes are insufficient

Engineering Contradiction:
Improvequality of 3D scene generationVSAvoidcomplexity of generation process
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the 3D scene generation process into distinct modules: text-to-panoramic-image generation, panoramic image to multi-view image conversion, depth estimation, and point cloud generation. Each module handles a specific task, improving overall quality while managing complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from 2D panoramic images to 3D point cloud representations by introducing depth estimation. This dimensional transformation enables realistic 3D scene generation while maintaining controllable complexity through the structured pipeline.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If manual 3D scene creation methods are used, then customization is possible, but the efficiency and speed are low

Engineering Contradiction:
Improveefficiency of 3D scene generationVSAvoidtime required for scene generation
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical 3D scene creation processes with an automated AI-based system. The text-to-image generation model and depth estimation algorithms automatically produce 3D scenes from text descriptions, dramatically improving efficiency and reducing time loss.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service 3D scene generation where users provide text descriptions and the system automatically generates complete 3D scenes without requiring manual modeling or complex operations. This automation significantly boosts productivity while minimizing time investment from users.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If existing AIGC methods are used for 3D scene generation, then speed is improved, but customization and versatility are insufficient

Engineering Contradiction:
Improvecustomization capabilityVSAvoidcomplexity of AIGC system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal AIGC system that handles multiple tasks: text-to-panoramic-image generation, panoramic-to-multi-view conversion, depth estimation, and point cloud generation. This multi-functional system provides extensive customization capability while managing complexity through integrated architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary actions by pre-training diffusion models for text-to-panoramic-image generation and pre-processing panoramic images into multi-view formats before depth estimation. These preliminary steps enable flexible customization while organizing complexity into manageable sequential operations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250299431A1Method, apparatus, and electronic device for three-dimensional scene generation
Publication Date: 2025.09.25 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250299431A1 patent drawing
  • US20250299431A1 patent drawing
  • US20250299431A1 patent drawing

AI summary

Embodiments of the present application disclose a method and an apparatus, and an electronic device for three-dimensional scene generation. A specific implementation of the method includes: obtaining a target text, and generating a panoramic image described by the target text; obtaining multi-view information in a plurality of preset views, and generating a multi-view image in the plurality of views with the panoramic image; performing depth estimation on the panoramic image to determine a sparse point cloud corresponding to the panoramic image; and generating, based on the multi-view image, the multi-view information, and the sparse point cloud, a three-dimensional scene model described by the target text.