The invention belongs to the technical field of three-dimensional scene modeling, and discloses a dual-
branch diffusion three-dimensional scene generation method based on a multi-
modal semantic graph, which comprises the following steps of: firstly, receiving multi-
modal data such as sketches, texts, automatic completion instructions and scene general knowledge, extracting features and fusing the features into a unified multi-
modal semantic graph; utilizing a graph neural network and an attention mechanism to enhance semantic graph features, and optimizing physical constraints through a physical engine; complementing the missing
visual modality and graph structure relationship; performing quality scoring on the scene based on
semantics, spatial relationships and physical constraints; and finally, respectively generating a spatial
layout and a geometric shape through a double-
branch diffusion model, and ensuring the coordination of the
layout and the shape. The method has the advantages of multi-modal
information fusion, physical rationality guarantee, high structure
complementation capability, high-
quality score optimization and efficient
generation process, and is suitable for three-dimensional scene modeling requirements in the fields of
virtual reality,
augmented reality, robots and the like.