Text-guided scene generation method and device based on scene semantic graph

By using a text-guided scene generation method based on scene semantic graphs, object relationships and attributes are explicitly modeled. A two-stage strategy is adopted to handle discrete and continuous attributes, which solves the problem of insufficient interpretability and controllability of existing 3D scene generation technologies and achieves higher quality 3D scene generation.

CN122391473APending Publication Date: 2026-07-14TSINGHUA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2026-04-02
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing methods for generating 3D indoor scenes suffer from limited text instruction parsing capabilities, lack of explicit object relationship modeling, and imprecise control over object attributes, resulting in deficiencies in interpretability, controllability, and layout rationality of the generated results.

Method used

We adopt a text-guided scene generation method based on scene semantic graphs. By constructing a structured semantic graph that includes object categories, quantities, and semantic and spatial relationships between objects, and combining discrete and continuous diffusion models, we can explicitly express object relationships and global layout structure. We use a two-stage strategy to handle discrete and continuous attributes respectively, which reduces the optimization difficulty and improves training stability.

Benefits of technology

It improves the interpretability, controllability, and layout rationality of the generated results, enhances the semantic consistency and spatial rationality of the generated results, has strong generalization ability, and is suitable for instruction-driven zero-shot generation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391473A_ABST
    Figure CN122391473A_ABST
Patent Text Reader

Abstract

The application provides a text-guided scene generation method and device based on a scene semantic graph, and is applied to the technical field of text-guided scene generation. The method comprises the following steps: based on a received scene description text, a trained target scene semantic graph construction encoder is called to construct a scene semantic graph corresponding to the scene description text, and the scene semantic graph construction encoder comprises a discrete diffusion model; based on the scene semantic graph, a trained target vector quantization variational autoencoder and a trained target scene layout decoder are called to generate a target scene corresponding to the scene description text, and the scene layout decoder comprises a continuous diffusion model. The application converts the text into an intermediate representation to explicitly express object relationships and global layout structures, the continuous diffusion model performs noise adding and noise removing in a semantic graph latent space, learns a reasonable distribution, guarantees the performance in semantic consistency, spatial reasonableness and functional constraints, and generates a target scene in combination with the semantic graph, so that the practicability and reasonableness of the generated result can be improved.
Need to check novelty before this filing date? Find Prior Art