DiT city layout generation method based on multi-mode cooperative driving
Through the multimodal collaboration-driven DiT urban layout generation method, the DiffusionTransformer model and multi-scale data processing are used to solve the problems of insufficient data sets and the inability of model to capture the diversity of urban styles in the existing technology, and the generation of high-quality, diverse and reasonable structured urban layouts is achieved.
Patent Information
- Application Number
- CN202510579359.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing urban layout generation technology faces the problems of insufficient data set resources, inability to effectively capture urban style diversity, and poor spatial coherence of the generation results.
A multimodal collaboration-driven DiT urban layout generation method is proposed. Through the DiffusionTransformer model, a diversified, controllable style and reasonable structure are generated through the DiffusionTransformer model combined with multi-scale data processing and semantic information fusion.
The quality and efficiency of urban layout generation have been improved, and the generated urban layout is more structurally and visually more reasonable and coherent, which can better capture the layout characteristics of different cities.
Smart Images

Figure CN120107491A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the intersection of computer vision and urban design technology, and in particular relates to a multimodal collaboratively driven DiT urban layout generation method, which aims to use advanced model architecture to achieve diversified, style-controllable and realistic urban layout generation. Background Art
[0002] Urban layout generation has become a key research frontier in computer vision and computational design, and it has made significant contributions to the advancement of urban morphological pattern learning and urban layout planning. At present, the research ideas on urban layout morphological patterns are mainly concentrated in three technical paths, each of which solves different challenges encountered in urban layout generation research. (1) Geospatial reconstruction and deep segmentation networks using satellite and aerial imagery, but face difficulties such as insufficient image resolution, content occlusion and segmentation accuracy. (2) Procedural modeling frameworks generate hierarchical urban elements through shape grammars and L-systems to achieve reasonable road networks and vegetation distribution. However, they rely on specially formulated rules, such as space grammar parameters, which makes it too complicated to generate coordinated layouts. (3) Inverse procedural modeling attempts to model urban layout using data-driven methods, but often oversimplifies urban complexity into geometric priors and is unable to generate stylized urban layouts.
[0003] With the continuous enrichment of geospatial data and the development of generative artificial intelligence, new ideas have been provided for learning different urban morphological patterns and layout generation. However, existing models often have difficulty capturing the diversity of layout styles inherent in global urban structures. This diversity requires the next generation of generative models to be able to learn the layout characteristics of different cities.
[0004] In addition, on the technical path, three fundamental challenges hinder the application of new generative models to the task of urban layout generation. First, due to the high cost of professional measurement, high-quality urban layout datasets are still scarce. Open source platforms such as OpenStreetMap (OSM) have geospatial asymmetry: developed areas provide high accuracy, while emerging cities are often represented using incomplete or schematic data, which limits the ability of existing models to capture real urban morphological patterns. Second, the mainstream model architecture has a convolutional inductive bias, which prioritizes local texture consistency rather than global spatial consistency, which will lead to discontinuities in the generated roads and color drift of plots. Third, the existing model frameworks often learn urban layout styles that tend to be homogenized, fail to capture the diversity between different urban layouts, and do not support the generation of layout data guided by multimodal information. Summary of the invention
[0005] In order to solve the problems of insufficient data set resources, inability of model architecture to effectively capture the diversity of urban styles, and poor spatial coherence of generated results in existing urban layout generation technologies, this paper proposes a multimodal collaboratively driven DiT urban layout generation method, which aims to learn cross-city layout patterns from limited real data sets through innovative data processing methods and model architecture design, generate diversified, style-controllable and structurally reasonable urban layouts, improve the quality and efficiency of urban layout generation, and fill the gaps in stable diffusion models and Transformer models in urban planning, design and related research fields.
[0006] We treat urban layout styles as learnable multimodal annotation information and embed them into the DiffusionTransformer (DiT) model. This enables our framework to capture the structural elements and style nuances of different urban layouts. To train our model, we curate a large-scale dataset of real urban layouts, which contains urban layout images and their style descriptions. We propose a comprehensive multi-scale data processing workflow and corresponding model training framework, which leverages the diffusion model with the Transformer architecture to generate different urban layouts.
[0007] To achieve the above object, the present invention adopts the following technical solution:
[0008] A multi-modal collaboratively driven DiT urban layout generation method comprises the following steps:
[0009] Step S101, training data acquisition and preprocessing, obtaining training sample data from the open source data platform OSM, and performing data preprocessing and enhancement to build a spatial database;
[0010] Step S102, constructing a three-level spatial pyramid to perform multi-scale visualization on the data acquired in step S101;
[0011] Step S103, classifying the city layout and annotating the semantic information of the visualization result in step S102;
[0012] Step S104, introducing a VAE (variational autoencoder) perceptual compression model trained on a large-scale dataset as a pre-training model;
[0013] Step S105, using the pre-trained model introduced in step S104 to compress the sample instance obtained in step S102 in the feature space, transferring the original sample instance from the image space to the latent space, and obtaining a training sample in the latent space;
[0014] Step S106, inputting the sample semantic annotation information in step S103 into a semantic information fusion module to extract semantic features;
[0015] Step S107, introducing the large visual model DINOv2 trained on a large-scale data set as an image feature encoding module, obtaining the encoded image guidance information features as a strong conditional signal to guide the model generation process;
[0016] Step S108, adding Gaussian noise to the training sample in step S105 through a forward diffusion process to obtain a noisy training sample;
[0017] Step S109, integrating the semantic feature information, the image guidance information features, the noisy training samples and the temporal dynamic characteristics of the diffusion process, using the Diffusion Transformer-based DiT generation model to train the diffusion model, and outputting the predicted noise;
[0018] Step S110, inverse diffusion is used to generate a predicted noise image, a final generated layout image is obtained, and a mean square error between the predicted noise and the actually added noise is calculated, and a weighted average is performed on the losses of multiple samples.
[0019] The technical solution is further optimized in that a three-level spatial pyramid is constructed in step S102, and the layout data obtained in step S101 is layered and saved as image sample data according to a macro-scale radius of 600m, a meso-scale radius of 400m, and a micro-scale radius of 200m.
[0020] In a further optimization of the technical solution, the VAE architecture is adopted in step S105 to achieve eight times space compression.
[0021] The technical solution is further optimized. In terms of the optimization of the semantic information fusion mechanism, step S106 adopts the Flan-T5-base model as the semantic encoding core, maps the text model semantic information into a 1024-dimensional text feature vector, and performs linear mapping and GeLU activation processing.
[0022] This technical solution is further optimized. In step S107, the self-supervised pre-trained DINOv2 visual large model is introduced into the urban layout generation task for the first time, and a strong semantic guidance is constructed through multi-scale feature extraction and hierarchical fusion strategy.
[0023] The technical solution is further optimized, and the step S108 is specifically as follows: performing a forward diffusion process on N training samples in the latent space. By mixing linear interpolation with random noise, the noisy latent variable sequence is generated as The maximum number of diffusion steps; specifically, each training sample is independently sampled time step , add noise according to the following formula:
[0024]
[0025] in, , diffuse noise scheduling coefficient , . Loop variable Indicates from 1 to Each step is used to calculate the cumulative coefficient The index of the intermediate time steps. Represents the time from time step 1 to the current step every intermediate step. represents a random Gaussian noise, Indicates that the mean is 0 and the covariance is the unit matrix The standard normal distribution of Represents the random Gaussian noise added Characteristics of data that follow a standard normal distribution. represents a decreasing noise schedule sequence, represent The amount of noise increased by one step.
[0026] The technical solution is further optimized. The DiT generation model in step S109 forms a deep network architecture by stacking 28 Transformer modules. Each Transformer module gradually refines the multi-level features of the image through the self-attention mechanism and the feedforward network, and performs adaptive instance normalization AdaLN-Zero on the semantic feature information in step S106.
[0027] In the further optimization of this technical solution, during the training process of the multimodal DiT city layout generation model, the probability Randomly discard conditional information and use the classifier to freely bootstrap the loss The constrained multimodal DiT urban layout generation model simultaneously learns conditional and unconditional generation capabilities. The specific calculation method is as follows:
[0028]
[0029] in, is a binary indicator function, Indicates conditional information. Indicates that there is no condition information. hour, =1, when the condition information is discarded, is the noisy output predicted by the DiT model, is random noise sampled from a standard Gaussian distribution, Represents a variable Take expectations. Indicates that the model is under given conditions In the case of generate The noise added when the conditional generation loss The smaller it is, the better the model is at recovering the target image from noise under conditional guidance. Represents the prediction noise of the model without any conditions, unconditional generation loss The smaller the loss, the better the model is at restoring the image even without any hints.
[0030] In a further optimization of the technical solution, a multi-step iterative denoising mechanism is used in step S110 to generate the final layout image and calculate the mean square error loss as follows:
[0031]
[0032] Among them, the expected symbol Represents the time step is a uniform distribution on the interval [1,T] Randomly sampled from Represents the original sample From the data distribution Obtained from Indicates that the noise is obtained from a standard Gaussian distribution. is the first The noisy samples of the step, , is the diffusion noise scheduling coefficient, It is the Flan-T5 semantic fusion vector and the DINOv2 visual guidance vector. represents standard Gaussian noise, is the model prediction output.
[0033] Different from the prior art, the above technical solution has the following beneficial effects:
[0034] Data advantages: The CityStyle-OSM dataset is large in scale, covers a wide range of cities, and has rich annotations. It provides sufficient and diverse data support for the urban layout generation model, helps the model learn more comprehensive and realistic urban layout patterns, and improves the accuracy and diversity of the generated results.
[0035] Previous similar methods were all based on the GAN framework. This invention, for the first time, introduced the Diffusion model and the Transformer architecture into the urban layout image generation work, and integrated the text information fusion module and the visual information attention guidance module into the DiT generation model. The generated urban layout is more structurally reasonable and visually coherent, and performs better than existing methods in terms of structural and visual consistency. This is the first time that the potential of the DiT model in urban layout generation tasks has been demonstrated.
[0036] Application value: The diversified and style-controllable urban layout generated by this invention can provide rich design inspiration and reference solutions for urban planners and designers, assisting them in carrying out urban planning and design work more efficiently; at the same time, it provides a powerful tool for urban-related research, which is helpful to promote research progress in the fields of urban development and urban morphology analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 Flowchart of the DiT city layout generation method driven by multimodal collaboration. DETAILED DESCRIPTION
[0038] In order to explain the technical content, structural features, achieved objectives and effects of the technical solution in detail, the following is a detailed description in conjunction with specific embodiments and accompanying drawings.
[0039] like Figure 1 , which is a schematic diagram of the process flow of the DiT city layout generation method driven by multi-modal collaboration. The city layout generation method of this embodiment specifically includes the following steps:
[0040] Step S101: acquiring and preprocessing training data.
[0041] Get geographic information spatial features within a specified set of longitude and latitude from the OpenStreetMap (OSM) platform, including road data nodes, building shapes, longitude and latitude coordinate information, and geographic name information.
[0042] Step S102, construct a three-level spatial pyramid to visualize the data obtained in step S101 at multiple scales. Visualize the geographic information features obtained within the latitude and longitude range according to the ranges of macro scale (radius 600m), meso scale (radius 400m), and micro scale (radius 200m), and save them as a set of N image data (1024×1024 resolution).
[0043] Step S103, semantic information annotation is performed on the sample data, and the visualization result in step S102 is classified into urban layout and semantic information annotation is performed. According to the information such as the number of buildings, road length, building type, and geographical location obtained in step S101, the image sample in step S102 is classified into urban layout and semantic information annotation is performed.
[0044] Step S104, introducing the VAE perceptual compression model trained on a large-scale dataset as a pre-training model.
[0045] Data-driven deep learning technology requires a large amount of training data to build large-scale models with strong generalization and robustness. A pre-trained variational autoencoder that learns basic feature representations on the ImageNet-21K general vision dataset is introduced as a plug-and-play perceptual compression module. The model is trained on the LAION-5B ultra-large-scale multimodal dataset (containing 5.85 billion image-text pairs) and achieves 8x compression from the original image to the latent space through an 8-layer convolutional downsampling structure (each layer with a step size of 2). The residual block structure and self-attention mechanism of its encoder and decoder can accurately capture key geometric features such as road topology and building outlines in urban layout.
[0046] Step S105, using the pre-trained model introduced in step S104 to compress the sample instance obtained in step S102 in the feature space, transferring the original sample instance from the image space to the latent space, and obtaining the training sample in the latent space.
[0047] The VAE architecture is used to achieve eight-fold spatial compression. Taking the original sample image of 512 pixels as an example, the encoder uses four convolutional layers with a stride of 2 to achieve spatial downsampling, compressing the original 512×512 pixels to a 64×64 latent space representation, and the number of channels is increased to 512 dimensions. At the same time, residual connections and convolutional layers are introduced in the decoder to maintain reconstruction accuracy.
[0048] Based on the introduced pre-trained variational autoencoder, the original N city layout images (resolution 1024×1024 pixels, RGB three channels) are compressed into the latent space through four levels of convolution downsampling operations with a step size of 2, generating N 128×128 dimensional four-channel latent variable representations. Specifically, the input image Via encoder The latent variable is obtained by nonlinear mapping At the same time, the multi-scale feature extraction capability obtained through pre-training is used to retain key spatial semantic information such as road connectivity, plot shape, and building orientation, making it possible to efficiently train the subsequent DiT model in the latent space.
[0049] Step S106, extracting features from the semantic information of the training samples, and inputting the sample semantic annotation information in step S103 into a semantic information fusion module to extract semantic features.
[0050] A semantic encoder is constructed based on the Flan-T5-base pre-trained language model to convert the N text semantic information (including unstructured descriptions such as building types, geographical names, and road grades) annotated in step S103 into a dense vector representation that can be fused. In terms of optimizing the semantic information fusion mechanism, the Flan-T5-base model is used as the semantic encoding core to map the text model semantic information into a 1024-dimensional text feature vector, and perform linear mapping and GeLU activation processing.
[0051] Step S107, introduce the large visual model DINOv2 trained on a large-scale data set as an image feature encoding module, and obtain the encoded image guidance information features as a strong conditional signal to guide the model generation process.
[0052] For the first time, the self-supervised pre-trained DINOv2 visual model was introduced into the urban layout generation task, and strong semantic guidance was constructed through multi-scale feature extraction and hierarchical fusion strategy. Low-level texture features with a resolution of 128×128 (number of channels: 256) and high-level semantic features with a resolution of 32×32 (number of channels: 1536) were extracted from the guidance information image, and after feature compression using global average pooling and maximum pooling operations, a 1024-dimensional guidance vector was generated through linear transformation.
[0053] Input road network image as visual guidance condition. Use DINOv2-giant pre-trained visual model as image guidance encoder and input reference image After 14×14 block processing in the ViT architecture, the image information features are extracted:
[0054]
[0055] Global average pooling preserves spatial distribution. Global maximum pooling to highlight salient areas. It is a channel splicing operation. Represents low-level detail features such as road network texture and building edge details. Represents high-level semantic features such as spatial topological relationships and layout area distribution. Finally, the encoded visual guidance feature vector is obtained , To initialize the learnable weight matrix, the input features are linearly transformed. is the bias vector, which adds a constant offset to the result of the linear transformation.
[0056] Step S108, adding noise to the training samples through a forward diffusion process, that is, adding Gaussian noise to the training samples in step S105 through a forward diffusion process to obtain noisy training samples.
[0057] Perform the forward diffusion process on N training samples in the latent space. Generate a noisy latent variable sequence by mixing linear interpolation with random noise ( =1000 is the maximum number of diffusion steps). Specifically, each training sample is independently sampled with time steps , add noise according to the following formula:
[0058]
[0059] in, , diffuse noise scheduling coefficient . Loop variable Indicates from 1 to Each step is used to calculate the cumulative coefficient The index of the intermediate time steps. Represents the time from time step 1 to the current step every intermediate step. represents a random Gaussian noise, Indicates that the mean is 0 and the covariance is the unit matrix The standard normal distribution of Represents the random Gaussian noise added Characteristics of data that follow a standard normal distribution. represents a decreasing noise schedule sequence, represent The amount of noise increased by one step.
[0060] Step S109, constructing a DiT generation model for training. The noisy samples, semantic feature information, image guidance information and the time dynamic characteristics of the diffusion process are integrated, and the Diffusion Transformer-based DiT generation model is used to train the diffusion model and output the predicted noise.
[0061] The DiT generation model in step S109 forms a deep network architecture by stacking 28 Transformer modules, and each Transformer module gradually refines the multi-level features of the image through the self-attention mechanism and the feedforward network. Adaptive instance normalization AdaLN-Zero is performed on the semantic feature information in step S106.
[0062] A cross-attention layer information fusion layer is introduced to fuse the visual guidance vector generated in step S108, and multimodal conditional injection is achieved through a key-value projection matrix.
[0063] During the training process of the multimodal DiT urban layout generation model, the probability Randomly discard conditional information and use the classifier to freely bootstrap the loss The constrained multimodal DiT urban layout generation model simultaneously learns conditional and unconditional generation capabilities. The specific calculation method is as follows:
[0064]
[0065] in, is a binary indicator function, Indicates conditional information. Indicates that there is no condition information. hour, =1, when the condition information is discarded, is the noisy output predicted by the DiT model, is random noise sampled from a standard Gaussian distribution, Represents a variable Take expectations. Indicates that the model is under given conditions In the case of generate The noise added when the conditional generation loss The smaller it is, the better the model is at recovering the target image from noise under conditional guidance. Represents the prediction noise of the model without any conditions, unconditional generation loss The smaller the loss, the better the model is at restoring the image even without any hints.
[0066] By stacking 28 Transformer modules, a DiT deep network architecture is formed. During the training phase, the noise latent variable in step S108 is , the semantic information vector in step S106 With time step Input DiT network to predict target noise and with real noise Calculating MSE loss , using the AdamW optimizer ( ) to update the parameters, and the learning rate is kept at 10 -4 .
[0067] In the In the layer Transformer module, the latent space tensor after adding noise obtained in step S108 is input , after the Patchify operation, it is divided into a 16 × 16 block sequence (a total of 128 / 16 × 128 / 16 = 4096), and each block is mapped to a 1152-dimensional feature through linear embedding to obtain the initial sequence , and superimposes learnable position encodings. Each Transformer module gradually refines the multi-level features of the image through multi-head self-attention and feedforward networks, and performs adaptive instance normalization and cross-attention fusion operations.
[0068] In the In the layer Transformer module, adaptive instance normalization AdaLN-Zero is performed on the semantic feature information in step S106, and the fusion guide semantic information enters the training framework. The specific operations are as follows:
[0069]
[0070] Then perform conditional normalization and self-attention:
[0071]
[0072]
[0073] in, It is 8-head self-attention.
[0074] In the In the layer Transformer module, the visual guidance vector generated in step 108 is Using the cross-attention layer information fusion layer, the visual guidance vector , after the key-value projection matrix , mapped to, , visual conditional injection DiT generation model is realized by interacting with the current feature through cross attention. Calculate cross attention and superimpose residuals:
[0075]
[0076] Among them, the cross attention calculation method is:
[0077]
[0078] Key-value projection matrix Map the visual conditions to a 128-dimensional subspace, query the matrix Process the current feature.
[0079] In the In the layer Transformer module, the cross-attention output is subjected to secondary normalization and MLP operation.
[0080] Step S110, generate the layout image and calculate the mean square error loss The predicted noise image is generated by reverse diffusion, and the final generated layout image is obtained. The mean square error between the predicted noise and the actual added noise is calculated, and the loss of multiple samples is weighted averaged.
[0081] A multi-step iterative denoising mechanism is used to generate the final layout image and calculate the mean square error loss as follows:
[0082]
[0083] Among them, the expected symbol Represents the time step is a uniform distribution on the interval [1,T] Randomly sampled from Represents the original sample From the data distribution In the Indicates that the noise is obtained from a standard Gaussian distribution. is the first The noisy samples of the step, , is the diffusion noise scheduling coefficient, It is the Flan-T5 semantic fusion vector and the DINOv2 visual guidance vector. represents standard Gaussian noise, is the model prediction output.
[0084] Output at layer 28 After linear projection, it is restored to , the generated city layout image is obtained through the VAE decoder:
[0085]
[0086] And calculate the prediction noise With real noise The mean square error :
[0087]
[0088] Under the conditions of generating images with different resolutions, the algorithm proposed in this invention has achieved better comprehensive results in seven common evaluation indicators compared with the mainstream framework based on generative adversarial networks (GAN) and stable diffusion (SDXL) in the current generative model. These seven evaluation indicators are: FID (Fréchet Inception Distance), KID (Kernel Inception Distance), MMD (Maximum Mean Discrepancy), MSE (Mean Square Error), LPIPS (Learned Perceptual Image Patch Similarity), PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index Measure). The specific results are shown in the table below:
[0089] Among them, FID is used to measure the distance between the generated image and the real image in the feature space. The lower the value, the closer the generated image is to the real image. KID evaluates the distribution difference. The lower the value, the better the effect. MMD is also a measure of distribution difference. A smaller value means that the distribution of the generated image is more similar to that of the real image. MSE calculates the mean square value of the pixel error of the image, which can intuitively reflect the degree of distortion of the image. The smaller the value, the closer the image is to the original image. LPIPS measures the image difference from a perceptual perspective and is more in line with human visual perception. The lower the value, the more similar the generated image is to the real image in perception. PSNR evaluates image quality based on the peak signal-to-noise ratio. The higher the value, the better the image quality. SSIM evaluates images from the perspective of structural similarity. The closer it is to 1, the more similar the structure of the generated image is to the real image, and the higher the quality. Under the comprehensive consideration of these evaluation indicators, the algorithm of the present invention shows significant advantages over the mainstream frameworks of GAN and SDXL.
[0090] Among the currently known solutions to the urban layout generation task, the present invention introduces the embedding idea of multimodal guidance conditions for the first time, supports visual information and text information to guide urban layout generation, and integrates the text information fusion module and the visual information attention guidance module to form a DiT generation model framework that supports multimodal information guidance. For the first time, the Diffusion model and the Transformer architecture are introduced into the urban layout image generation work, which provides a new method for solving the task of regenerating urban layout from geospatial data and demonstrates the potential of the DiT model in urban layout generation tasks. In addition, the present invention supports the generation of layout image data of different resolutions and has good generalization.
[0091] The invention has broad application prospects and high practical value, for example: these generated urban geospatial layout data can actually be used for spatial analysis, rather than just providing layout images. The model of the invention can generate realistic urban forms in areas lacking architectural data. The invention can also help planners and designers with initial needs for complex program modeling, such as simulating urban expansion or new area development.
[0092] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "include..." or "comprise..." do not exclude the existence of other elements in the process, method, article or terminal device including the elements. In addition, in this article, "greater than", "less than", "exceed" and the like are understood to exclude the number itself; "above", "below", "within" and the like are understood to include the number itself.
[0093] Although the above embodiments have been described, once those skilled in the art know the basic creative concepts, they can make additional changes and modifications to these embodiments. Therefore, the above description is only an embodiment of the present invention and does not limit the patent protection scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the specification and drawings of the present invention, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A multi-modal collaboratively driven DiT city layout generation method, characterized in that: The following steps are involved: Step S101, training data acquisition and preprocessing, obtaining training sample data from the open source data platform OSM, and performing data preprocessing and enhancement to build a spatial database; Step S102, constructing a three-level spatial pyramid to perform multi-scale visualization on the data acquired in step S101; Step S103, classifying the city layout and annotating the semantic information of the visualization result in step S102; Step S104, introducing a variational autoencoder perceptual compression model trained on a large-scale data set as a pre-training model; Step S105, using the pre-trained model introduced in step S104 to compress the sample instance obtained in step S102 in the feature space, transferring the original sample instance from the image space to the latent space, and obtaining a training sample in the latent space; Step S106, inputting the sample semantic information annotation in step S103 into a semantic information fusion module for semantic feature extraction; Step S107, introducing the large visual model DINOv2 trained on a large-scale data set as an image feature encoding module, obtaining the encoded image guidance information features as a strong conditional signal to guide the model generation process; Step S108, adding Gaussian noise to the training sample in step S105 through a forward diffusion process to obtain a noisy training sample; Step S109, integrating the semantic feature information, the image guidance information features, the noisy training samples and the temporal dynamic characteristics of the diffusion process, using the Diffusion Transformer-based DiT generation model to train the diffusion model, and outputting the predicted noise; Step S110, inverse diffusion is used to generate a predicted noise image, a final generated layout image is obtained, and a mean square error between the predicted noise and the actually added noise is calculated, and a weighted average is performed on the losses of multiple samples.
2. The multi-modal collaboratively driven DiT city layout generation method according to claim 1, characterized in that: In step S102, a three-level spatial pyramid is constructed, and the layout data obtained in step S101 is layered and saved as image sample data according to a macro scale radius of 600m, a meso scale radius of 400m, and a micro scale radius of 200m.
3. The multi-modal collaboratively driven DiT city layout generation method according to claim 1, characterized in that: In step S105, VAE architecture is used to achieve eight times space compression.
4. The multi-modal collaboratively driven DiT city layout generation method according to claim 1, characterized in that: In terms of optimizing the semantic information fusion mechanism, step S106 adopts the Flan-T5-base model as the semantic encoding core, maps the text model semantic information into a 1024-dimensional text feature vector, and performs linear mapping and GeLU activation processing.
5. The multi-modal collaboratively driven DiT city layout generation method according to claim 1, characterized in that: In step S107, the self-supervised pre-trained DINOv2 visual large model is introduced into the urban layout generation task for the first time, and a strong semantic guidance is constructed through multi-scale feature extraction and hierarchical fusion strategy.
6. The multi-modal collaboratively driven DiT city layout generation method according to claim 1, characterized in that: The step S108 is specifically as follows: performing a forward diffusion process on N training samples in the latent space, transforming the original latent variables Generate a noisy latent variable sequence by mixing linear interpolation with random noise , = 1000 is the maximum number of diffusion steps; specifically, each training sample is independently sampled at the time step , add noise according to the following formula: in, , diffuse noise scheduling coefficient , loop variable Indicates from 1 to Each step is used to calculate the cumulative coefficient The intermediate time step index of Represents the time from time step 1 to the current step At each intermediate step, represents a random Gaussian noise, Indicates that the mean is 0 and the covariance is the unit matrix The standard normal distribution of Represents the random Gaussian noise added The characteristics of data that obey the standard normal distribution are: represents a decreasing noise schedule sequence, represent The amount of noise increased by one step.
7. The multi-modal collaboratively driven DiT city layout generation method according to claim 1, characterized in that: The DiT generation model in step S109 forms a deep network architecture by stacking 28 Transformer modules. Each Transformer module gradually refines the multi-level features of the image through the self-attention mechanism and the feedforward network, and performs adaptive instance normalization AdaLN-Zero on the semantic feature information in step S106.
8. The multi-modal collaboratively driven DiT city layout generation method according to claim 1, characterized in that: During the training process of the multimodal DiT urban layout generation model, the probability Randomly discard conditional information and use the classifier to freely bootstrap the loss The constrained multimodal DiT urban layout generation model simultaneously learns conditional and unconditional generation capabilities. The specific calculation method is as follows: in, is a binary indicator function, Indicates conditional information. Indicates that there is no conditional information. hour, , when the condition information is discarded, is the noisy output predicted by the DiT model, is random noise sampled from a standard Gaussian distribution, Represents a variable Take expectations, Indicates that the model is under given conditions In the case of generate The noise added when the conditional generation loss The smaller it is, the better the model is at recovering the target image from noise under conditional guidance. Represents the prediction noise of the model without any conditions, unconditional generation loss The smaller the loss, the better the model is at restoring the image even without any hints.
9. The multi-modal collaboratively driven DiT city layout generation method according to claim 1, characterized in that: In step S110, a multi-step iterative denoising mechanism is used to generate the final layout image and calculate the mean square error loss as follows: Among them, the expected symbol Represents the time step is a uniform distribution on the interval [1,T] Randomly sampled from Represents the original sample From the data distribution In the Indicates that the noise is obtained from a standard Gaussian distribution, is the first The noisy samples of the step, , is the diffusion noise scheduling coefficient, It is the Flan-T5 semantic fusion vector and the DINOv2 visual guidance vector. represents standard Gaussian noise, is the model prediction output.
Citation Information
Patent Citations
Multi-modal data driven urban road layout design automation method
CN115544613A
Construction method of image description model, image description method and equipment
CN119295770A
Building planning image generation method and system based on potential diffusion model
CN119557955A
Method and equipment for generating ultrahigh-resolution city layout by using design drawing, and medium
CN119850786A
Text and color-guided layout control with a diffusion model
US20240169604A1
Cited By
Urban three-dimensional modeling and dynamic simulation environment generation method and device based on multi-modal large model
CN120747380A
Urban safety abnormal event image generation method, system, equipment and medium
CN121392471A
Methods, systems, equipment and media for generating images of urban safety anomalies
CN121392471B
Remote sensing image semantic segmentation method and device and storage medium
CN121616837A