The invention belongs to the crossing field of
computer vision and city design, and provides a multi-mode cooperative driving DiT city
layout generation method, so as to solve the defects of an existing generation model in the aspects of spatial rationality and planning
controllability. According to the method, a
transformer-based
diffusion generation model is innovatively constructed, and two additional control signals are integrated in a DiT model to realize controlled generation. On the aspect of
model architecture design, a DiT framework with multi-
modal condition fusion is constructed, and a road network sketch (image control) and planning
semantics (text control) dual guide mechanism is integrated. In the
text mode, a dynamic semantic fusion module based on a pre-training
language model is designed, and text information is deeply embedded into the
generation process. In the aspect of image control, a mixed attention regulation and control mechanism is put forward, multi-scale fusion of road network structural features is realized in combination with cross attention and AdaLN technologies, and the problem of spatial
layout distortion is solved. Experiments show that the method can effectively integrate text and image information, and the generated city
layout image is superior to that generated by an existing method.