Urban road layout diagram conditional expression generation method based on potential diffusion model
Through the conditional coding, denoising and feature mapping modules of the potential diffusion model, the problem of insufficient generation quality and diversity in the prior art is solved, and the efficient generation of urban road layout design drawings that conform to geographical information is achieved.
Patent Information
- Application Number
- CN202510742740.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-05
AI Technical Summary
The existing urban road layout design automation methods are difficult to balance between generating high quality and diversity. The road layout generated by the method based on the generative adversarial network is insufficient, while the method training and inference efficiency of the diffusion model based on pixel space is low.
The potential diffusion model is adopted, and the population density, topographic elevation and land use data are encoded through the conditional encoding module, and the urban road layout feature map is generated by combining the denoising module, and the road layout design map that conforms to geographical information is decoded through the feature mapping module.
It improves the generation quality and diversity of urban road layout design, reduces the calculation complexity and calculation time, enhances the stability and adaptability of the generation process, and improves the level of intelligence.
Smart Images

Figure CN120429932A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of deep learning and computational vision, and in particular relates to a conditional generation method for urban road layout diagrams based on a potential diffusion model. Background Art
[0002] Urban road systems are the organizational support framework for a city's functional structure. Urban road layout design has a wide range of applications, including urban planning, game design, and autonomous driving simulation. A well-planned urban road layout can effectively improve a city's transportation capacity, meet the growing travel needs of urban residents, and enhance their well-being. In the gaming and film industries, well-designed urban road layouts enhance user experience and create a sense of immersion. In the field of autonomous driving simulation, urban road designs that mirror real-world conditions are essential to accommodate diverse testing scenarios.
[0003] Existing automated methods for urban road layout design can be divided into two categories. The first category is traditional methods based on procedural modeling. These methods can automatically generate road layouts that meet requirements using predefined rules with low storage and computational overhead, but they require extensive user intervention and are extremely inflexible. The second category is automated generation methods based on deep learning technology. These methods use deep learning models to extract structural and texture features from real-world data to automatically generate road layouts. Existing deep learning-based methods mainly include those based on generative adversarial networks and those based on pixel-space diffusion models. GAN-based methods use adversarial learning techniques to train generative models, which can reduce user intervention and enable more efficient road layout modeling. However, the road layouts generated by these methods lack diversity and are of low quality. Pixel-space diffusion models generate high-quality road layout images by learning a noise inverse process, while also taking into account the diversity of generated roads. However, these methods suffer from low training and inference efficiency.
[0004] Therefore, it is necessary to design a new deep learning-based automated method for urban road layout design that takes into account generation efficiency while ensuring high generation quality. Summary of the Invention
[0005] The technical problem to be solved by the present invention is how to introduce multimodal spatial geographic information into the potential diffusion model to generate high-quality urban road layout design drawings that meet the corresponding input conditions, and provide a conditional generation method for urban road layout drawings based on the potential diffusion model.
[0006] In order to achieve the above-mentioned object of the invention, the present invention specifically adopts the following technical solutions:
[0007] A conditional generation method for an urban road layout diagram based on a potential diffusion model comprises the following steps:
[0008] S1. Obtain population density raster data, terrain elevation raster data, and land use raster data for the target area, and construct multimodal input data from these three raster data. Superimpose the different modal input data using the feature channel dimension, and input the superimposed input data into a pre-trained conditional coding module to achieve data compression and dimensionality reduction, thereby generating conditional coding maps at different levels.
[0009] S2. Randomly sample a Gaussian noise from a standard Gaussian distribution and use it as the initial input of the pre-trained denoising module. At the same time, use the conditional coding maps of each level generated by the conditional coding module as conditional information. The denoising module is used to generate a noise image through iterative denoising in successive rounds, ultimately obtaining the corresponding urban road layout feature map.
[0010] S3. Input the urban road layout feature map output by the denoising module into a pre-trained feature mapping module for decoding, and finally generate an urban road layout design map.
[0011] Based on the above solution, each step can be implemented in the following preferred specific manner.
[0012] Preferably, in step S1, the conditional encoding module includes a first input module and four downsampling modules; the first input module is composed of three 3*3 convolutional layers cascaded in sequence, and the four downsampling modules all adopt the ResNet architecture; in the first downsampling module, the superimposed input data is processed by two 3*3 convolutional layers in sequence to obtain the first feature, and the superimposed input data is added to the first feature through a residual connection to obtain the second feature and input it into the 3*3 downsampling convolutional layer to obtain the output of the first downsampling module; each 3*3 convolutional layer is provided with a normalization layer and a Swish activation function.
[0013] Preferably, in step S2, the denoising module adopts a U-Net structure. In the denoising module, the Gaussian noise initially input is processed by the second input module to generate a feature map with a target number of channels and passes through four encoder modules in sequence to generate feature maps of different sizes; the feature map finally output by the fourth encoder module is input to the middle layer module, and the feature map output by the middle layer module is used as the input of the fourth decoder module, and each decoder module cross-layer connects its own input with the feature map of the same size generated by the corresponding encoder module, and at the same time, the output of each decoder module is feature fused with the corresponding size conditional coding map output by the conditional coding module through the conditional control module, and the fused feature map is used as the input of the next decoder module; finally, the output of the first decoder module is processed by the output module to generate a feature map of the urban road layout.
[0014] Preferably, the first two encoder modules of the denoising module have the same structure, and the last two encoder modules have the same structure; in the first encoder module, the input features are sequentially processed by two 3*3 convolutional layers to obtain the third feature, the input features are added to the third feature through residual connection to obtain the fourth feature, the fourth feature is sequentially processed by two 3*3 convolutional layers to obtain the fifth feature, the fourth feature is added to the fifth feature through residual connection to obtain the sixth feature, the sixth feature is sequentially processed by two 3*3 convolutional layers to obtain the seventh feature, the sixth feature is added to the seventh feature through residual connection to obtain the eighth feature and the eighth feature is processed by a 3*3 downsampling convolution layer to obtain the output of the first encoder module; in the third encoder module, the input features are sequentially processed by two 3* After processing by 3 convolution layers, the ninth feature is obtained. The input feature is added to the ninth feature through residual connection to obtain the tenth feature. After the tenth feature is processed by the self-attention layer, the first self-attention feature is obtained. The first self-attention feature is processed by two 3*3 convolution layers in sequence to obtain the eleventh feature. The first self-attention feature is added to the eleventh feature through residual connection to obtain the twelfth feature. After the twelfth feature is processed by the self-attention layer, the second self-attention feature is obtained. After the second self-attention feature is processed by two 3*3 convolution layers in sequence, the thirteenth feature is obtained. The second self-attention feature is added to the thirteenth feature through residual connection to obtain the fourteenth feature. The fourteenth feature is processed by the self-attention layer and the 3*3 downsampling convolution layer in sequence to obtain the output of the third encoder module.
[0015] Preferably, the 3*3 downsampling convolution layer of the encoder module in the denoising module is replaced by a 4*4 upsampling deconvolution layer to obtain a decoder module corresponding to the encoder module; the second input module and the output module are both a 3*3 convolution layer; in the intermediate layer module, the input feature is processed by two 3*3 convolution layers in sequence to obtain the fifteenth feature, the input feature is added to the fifteenth feature through a residual connection to obtain the sixteenth feature, the sixteenth feature is processed by two 3*3 convolution layers in sequence to obtain the seventeenth feature, the sixteenth feature is added to the seventeenth feature through a residual connection to obtain the output of the intermediate layer module.
[0016] Preferably, in the conditional control module, the conditional coding map output by the conditional coding module is processed by the first convolution layer to obtain a bias parameter, and the conditional coding map output by the conditional coding module is processed by the second convolution layer to obtain a scale parameter. The scale parameter is multiplied by the feature map output by the decoder module to obtain a weighted feature map, and the weighted feature map is added to the bias parameter to obtain the output of the conditional control module.
[0017] Preferably, in step S3, the feature mapping module adopts an autoencoder structure, which is composed of an encoder and a decoder; in the pre-training stage of the feature mapping module, the original urban road layout design drawing is first processed by the encoder to generate an encoded feature map and input it into the decoder to reconstruct the original urban road layout design drawing; in the inference stage of the feature mapping module, the urban road layout feature map generated by the denoising module is used as the input of the feature mapping module decoder, and the final generated urban road layout design drawing is output after decoding.
[0018] Preferably, the conditional coding module and the denoising module adopt a joint training method; in the joint training process, a training data set consisting of real-world urban road layout maps is first obtained, and an urban road layout map is randomly selected from the training data set as a training image, the training image is encoded into a target feature map using the encoder of the feature mapping module, and a current iteration time step t is randomly selected from the interval [0, T], T represents the maximum number of iteration steps, a Gaussian noise is randomly sampled from the standard Gaussian distribution, and the target feature map is denoised through a forward diffusion process to generate a noisy map corresponding to time step t, and then the time step t and the corresponding noisy map are used as inputs of the denoising module, and the conditional coding map output by the conditional coding module is used as conditional information, and finally the denoising module outputs an estimated noise image, and the mean square error loss between the noise image generated by the denoising module and the randomly sampled Gaussian noise is calculated, and the mean square error loss is optimized to train the conditional coding module and the denoising module.
[0019] Preferably, the time step t input to the denoising module must first be processed by the time encoding module to generate the corresponding time code, and then the time code is input to each encoder module and decoder module, and added to the output features of the first 3*3 convolutional layer in the encoder module and decoder module to achieve deep fusion of time information and feature information; wherein, the time encoding module is composed of two linear layers cascaded in sequence.
[0020] Preferably, in step S3, during the pre-training stage of the feature mapping module, the mean absolute error between the urban road layout design map reconstructed by the decoder and the original urban road layout design map is used as the loss function.
[0021] Compared with the prior art, the present invention has the following beneficial effects:
[0022] The present invention introduces a potential diffusion model in the task of urban road layout design to enhance the intelligent generation capability of road layout. Specifically, the model consists of a conditional coding module, a denoising module and a feature mapping module. First, the conditional coding module encodes three kinds of conditional information, namely population density, terrain elevation and land use, to generate a conditional coding map. Subsequently, the denoising module uses Gaussian noise and the conditional coding map as input to generate a characteristic map of urban road layout that meets the conditions. Finally, the feature mapping module decodes the characteristic map of urban road layout to generate an urban road layout design map that meets the requirements of population distribution, topography and land type. Compared with traditional manual design methods, the present invention does not need to rely on professional knowledge in related fields, nor is it limited to preset road templates. It can efficiently generate urban road layouts with rich structures and diverse features. In addition, compared with the previous diffusion model method based on pixel space, the present invention adopts a latent diffusion model with better effect in the field of image generation, and performs denoising in the latent space, so that the model converges more stably, has higher computational efficiency and lower computational complexity during the generation process. At the same time, it improves the generation quality and diversity of road layout images, and can effectively ensure that the generation results meet the relevant geographic information conditions, thereby improving the intelligence level and adaptability of urban road layout design. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 Design a latent diffusion model for conditional generation of urban road layout diagrams;
[0024] Figure 2 Schematic diagram of the training process of the feature mapping module;
[0025] Figure 3 It is a structural diagram of the denoising module;
[0026] Figure 4 Schematic diagram of the encoder module structure with attention layer in the denoising module;
[0027] Figure 5 This is a structural diagram of the conditional control module in the denoising module;
[0028] Figure 6 Schematic diagram of the joint training process of the denoising module and the conditional encoding module;
[0029] Figure 7 This is a structural diagram of the conditional coding module;
[0030] Figure 8 Schematic diagram of the middle layer module structure in the denoising module;
[0031] Figure 9 Schematic diagram of the reconstruction result of the pre-trained feature mapping module; (a) is a schematic diagram of the original image, and (b) is a schematic diagram of the reconstructed image;
[0032] Figure 10 Generate the corresponding urban road layout map for the potential diffusion model; among them, (a) is the generated road network layout map, (b) is the real road network layout map, (c) is the population density map, (d) is the terrain elevation map, and (e) is the land use map. DETAILED DESCRIPTION
[0033] In order to make the above-mentioned objects, features and advantages of the present invention more clearly understood, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined accordingly without conflicting with each other.
[0034] In the description of the present invention, it should be understood that the terms "first" and "second" are used solely for descriptive purposes and are not to be construed as indicating or implying relative importance or implicitly specifying the number of technical features being described. Therefore, features defined as "first" or "second" may explicitly or implicitly include at least one of such features.
[0035] One of the core aspects of this invention lies in the fusion of multimodal data to improve the accuracy and adaptability of conditional formula generation for urban road layouts. Multimodal data refers to data types such as images, speech, text, and sensor data that come from different sources and have different representations and feature dimensions. In deep learning, the fusion of multimodal data can provide more comprehensive, accurate, and rich information, thereby improving model performance and generalization capabilities.
[0036] In this paper, a multimodal data fusion solution is constructed using population density raster data, terrain elevation raster data, and land use data as examples. The conditional encoding module is used to address multimodal data fusion. It encodes information of different data types, extracts key features, and provides constraints to guide the generation of road layouts. It should be noted that the specific structure of the conditional encoding module is not limited; different model architectures can be used to adapt to different task requirements.
[0037] Another key innovation of this invention is the introduction of a conditional control module into the latent diffusion model, which more efficiently integrates conditional information and provides precise conditional control. This improvement effectively integrates geographic information provided by multimodal data, thereby improving the rationality of generated urban road layouts and enhancing the controllability and adaptability of the model.
[0038] The specific design method and implementation of the present invention will be described in detail below.
[0039] In a preferred embodiment of the present invention, a method for conditional generation of an urban road layout diagram based on a potential diffusion model is provided. The method comprises three steps:
[0040] S1. Obtain population density raster data, terrain elevation raster data, and land use raster data for the target area, and construct multimodal input data from these three raster data. Superimpose the different modal input data using the feature channel dimension, and input the superimposed input data into a pre-trained conditional coding module to achieve data compression and dimensionality reduction, thereby generating conditional coding maps at different levels.
[0041] S2. Randomly sample a Gaussian noise from a standard Gaussian distribution and use it as the initial input of the pre-trained denoising module. At the same time, use the conditional coding maps of each level generated by the conditional coding module as conditional information. The denoising module is used to generate a noise image through iterative denoising in successive rounds, ultimately obtaining the corresponding urban road layout feature map.
[0042] S3. Input the urban road layout feature map output by the denoising module into a pre-trained feature mapping module for decoding, and finally generate an urban road layout design map.
[0043] The above S1, S2 and S3 jointly design a complete potential diffusion model for conditional generation of urban road layout diagrams. The overall structure of the model is as follows: Figure 1 .from Figure 1 It can be seen that the model contains a conditional encoding module, a denoising module and a feature mapping module.
[0044] It should be noted that, in step S1 of the present invention, Figure 7 As shown in FIG, the conditional encoding module includes a first input module and four downsampling modules; wherein the first input module is composed of three cascaded 3*3 convolutional layers, and the four downsampling modules all adopt the ResNet architecture. In the first downsampling module, the superimposed input data is processed by two 3*3 convolutional layers in sequence to obtain the first feature. The superimposed input data is added to the first feature through the residual connection to obtain the second feature and input it into the 3*3 downsampling convolutional layer (i.e. Figure 7 The output of the first downsampling module is obtained by adding the downsampling layer in the convolutional layer. Each 3*3 convolutional layer is equipped with a normalization layer and a Swish activation function.
[0045] In an embodiment of the present invention, the input of the above-mentioned conditional coding module is data of different modalities. The present invention uses three modal raster data, namely population density raster data, terrain elevation raster data and land use raster data. These raster data are stored in the form of single-channel 512*512 size rasterized images and are superimposed in the feature channel dimension to form multimodal input data that is finally input to the conditional coding module. Then, the raster data of the three different modalities are fused in the conditional coding module and finally generate 4 different levels of coding maps, whose sizes are 64*64, 32*32, 16*16 and 8*8 respectively. These 4 different levels of coding maps are used to perform feature fusion with the output of the corresponding size of the decoder module in the denoising module. While fusing multimodal semantic information, this module can effectively reduce computational complexity and avoid the high computing power requirement problem caused by the excessive amount of model parameters.
[0046] In the denoising module of step S2 of the present invention, the Gaussian noise of the initial input is processed by the second input module to generate a feature map with a target number of channels and pass through four encoder modules in sequence to generate feature maps of sizes 64*64, 32*32, 16*16, and 8*8 respectively; the feature map finally output by the fourth encoder module is input to the middle layer module to further extract and fuse features to generate a new 8*8 feature map; then, the feature map output by the middle layer module is used as the input of the fourth decoder module, and each decoder module cross-layer connects its own input with the feature map of the same size generated by the corresponding encoder module, and at the same time, the output of each decoder module is feature-fused with the corresponding size conditional coding map output by the conditional coding module through the conditional control module, and the fused feature map is used as the input of the next decoder module; finally, the output of the first decoder module is processed by the output module to generate a 64*64 feature map, i.e., the urban road layout feature map.
[0047] It should be noted that in step S2 of the embodiment of the present invention, the denoising module adopts a U-Net structure, which includes a second input module, four encoder modules, an intermediate layer module, four decoder modules, four condition control modules and an output module. Figure 3As shown in the figure. The second input module is a 3*3 convolutional layer. The first two encoder modules have the same structure, both using the ResNet architecture, and each contains six 3*3 convolutional layers and one 3*3 downsampling convolutional layer. According to the ResNet design, a residual connection is set between every two consecutive 3*3 convolutional layers. The structure of the last two encoder modules is the same and similar to the first two encoder modules. The difference is that the last two encoder modules have an additional self-attention layer connected after the second, fourth, and sixth 3*3 convolutional layers. The intermediate layer module adopts the ResNet architecture and includes four 3*3 convolutional layers with normalization layers and Swish activation functions. The four decoder modules are connected to the corresponding four encoder modules using a skip-connection structure. The structure of each decoder module is the same as the corresponding encoder module, but the 3*3 downsampling convolutional layer is replaced by a 4*4 upsampling deconvolution layer. The output module is a 3*3 convolutional layer with a normalization layer and Swish activation function.
[0048] Specifically, in the first encoder module, the input features are processed by two 3*3 convolution layers in sequence to obtain the third feature, the input features are added to the third feature through residual connection to obtain the fourth feature, the fourth feature is processed by two 3*3 convolution layers in sequence to obtain the fifth feature, the fourth feature is added to the fifth feature through residual connection to obtain the sixth feature, the sixth feature is processed by two 3*3 convolution layers in sequence to obtain the seventh feature, the sixth feature is added to the seventh feature through residual connection to obtain the eighth feature and it is processed by a 3*3 downsampling convolution layer to obtain the output of the first encoder module. Figure 4 As shown in the figure, in the third encoder module, the input features are processed by two 3*3 convolution layers in sequence to obtain the ninth feature, and the input features are added to the ninth feature through residual connection to obtain the tenth feature. The tenth feature is processed by the self-attention layer to obtain the first self-attention feature. The first self-attention feature is processed by two 3*3 convolution layers in sequence to obtain the eleventh feature. The first self-attention feature is added to the eleventh feature through residual connection to obtain the twelfth feature. The twelfth feature is processed by the self-attention layer to obtain the second self-attention feature. The second self-attention feature is processed by two 3*3 convolution layers in sequence to obtain the thirteenth feature. The second self-attention feature is added to the thirteenth feature through residual connection to obtain the fourteenth feature. The fourteenth feature is processed by the self-attention layer and the 3*3 downsampling convolution layer in sequence to obtain the output of the third encoder module.
[0049] like Figure 8As shown in the figure, in the middle layer module, the input feature is processed by two 3*3 convolution layers in sequence to obtain the fifteenth feature, and the input feature is added to the fifteenth feature through residual connection to obtain the sixteenth feature. The sixteenth feature is processed by two 3*3 convolution layers in sequence to obtain the seventeenth feature, and the sixteenth feature is added to the seventeenth feature through residual connection to obtain the output of the middle layer module.
[0050] like Figure 5 As shown in the figure, the conditional control module uses spatial modulation to fuse the feature map output by the decoder module with the conditional coding map output by the conditional coding module. Specifically, first, the conditional coding map output by the conditional coding module is processed by the first convolutional layer to obtain the bias parameter β(f), and the conditional coding map output by the conditional coding module is processed by the second convolutional layer to obtain the scale parameter γ(f). These two parameters are then used to modulate the feature map output by the decoder module, that is, the scale parameter is multiplied by the feature map output by the decoder module to obtain a weighted feature map, and the weighted feature map is added to the bias parameter to obtain the output of the conditional control module:
[0051] h=γ(f)·g+β(f)
[0052] Among them, f is the conditional coding map output by the conditional coding module; g is the feature map output by the decoder module; h is the fused feature map after spatial modulation, that is, the output of the conditional control module.
[0053] It should be noted that, in the present invention, the above-mentioned conditional coding module and denoising module need to be jointly trained before being used for actual reasoning. Among them, the denoising module includes a forward diffusion process and a reverse denoising process. During the training process, a training data set consisting of a real-world urban road layout map is first obtained, and an urban road layout map is randomly selected from the training data set as a training image. Subsequently, the training image is encoded into a target feature map using the encoder of the feature mapping module, and a time step t is randomly taken from the interval [0, T], where T represents the maximum number of iteration steps (generally set to 1000). At the same time, a Gaussian noise is randomly sampled from the standard normal distribution (Gaussian distribution), and noise processing is performed through the forward diffusion process, that is, a series of noise weights {β1,...,β T}The Gaussian noise is gradually added to the target feature map to generate the noise map z corresponding to time step t t, then the time step t and the corresponding noise image are used as the input of the denoising module, and the conditional encoding image output by the conditional encoding module is used as the conditional information. Finally, the denoising module outputs the estimated noise image. The mean square error (MSE) loss between the noise image generated by the denoising module and the randomly sampled Gaussian noise is calculated, and the conditional encoding module and denoising module are trained by optimizing this loss. The specific training process is as follows Figure 6 shown.
[0054]
[0055] α t =1-β t
[0056]
[0057] Among them, z0 is the training image; ∈ is random Gaussian noise; β t is the noise weight corresponding to time step t; and α i All represent intermediate variables, D is the denoising module, c is the conditional information, Indicates expectation.
[0058] In this embodiment, for the first time step, the input noisy image is randomly sampled Gaussian noise, and in subsequent iterative steps, the input noisy image is the noise image after the previous round of denoising.
[0059] It is important to note that the time step t input to the denoising module must first be processed by the time encoding module to generate the corresponding time code. The time code is then input to each encoder module and decoder module and added to the output features of the first 3*3 convolutional layer in the encoder module and decoder module to achieve a deep fusion of time information and feature information, thereby improving the model's adaptability and expressiveness for different time steps. In this embodiment, the time encoding module is composed of two linear layers cascaded in sequence.
[0060] During the inference process of the conditional encoding module and the denoising module, a maximum number of iterations, T, is first determined (usually the same as the maximum number of iterations set during training, typically 1000). Next, a Gaussian noise is randomly sampled from a standard normal distribution as the original input. Next, the required conditional information, such as population density raster data, terrain elevation raster data, and land use raster data, is obtained and input into the conditional encoding module to generate conditional encoding maps of varying sizes. Subsequently, a reverse denoising process is performed step by step. In each iteration, the current noise map, the corresponding conditional encoding map, and the time step t are used as input to generate the corresponding denoised noise. This noise is then removed using a specific denoising algorithm to generate the next noise map. This noise map is then used as input for the next iteration, and this process is repeated. After T iterations, a feature map of the urban road layout is generated. Finally, this feature map is passed as input to the decoder of the feature mapping module, where the final urban road layout design map is generated through the decoding process.
[0061] It should be noted that in step S3 of the present invention, the feature mapping module employs an autoencoder structure, consisting of an encoder and a decoder. During the pre-training phase of the feature mapping module, the original urban road layout design drawing is first processed by the encoder to generate an encoded feature map, which is then input into the decoder to reconstruct the original urban road layout design drawing. During the inference phase of the feature mapping module, the urban road layout feature map generated by the denoising module serves as input to the feature mapping module decoder, which decodes the map and outputs the final urban road layout design drawing.
[0062] It should be noted that the encoder is only used in the training phase. Its main function is to assist the decoder in training so that the urban road layout design diagram reconstructed by the decoder is as close as possible to the original urban road layout design diagram. It does not actually participate in the reasoning process in the reasoning phase.
[0063] In the embodiment of the present invention, during the pre-training phase of the feature mapping module, a 512*512 sized urban road layout design map is used as input. The map is first processed by the encoder to generate a 64*64 sized encoded feature map. Then, the decoder decodes the feature map to reconstruct a 512*512 sized urban road layout design map. The training process of the feature mapping module is as follows: Figure 2 In the inference phase, the 64*64 urban road layout feature map generated by the denoising module is used as the input to the decoder of the feature mapping module. After decoding, the final urban road layout design map is output.
[0064] It should be noted that in the pre-training stage of the feature mapping module, the mean absolute error (MAE) between the urban road layout design map reconstructed by the decoder and the original urban road layout design map is used as the loss function.
[0065] In the embodiment of the present invention, the feature mapping module uses the mean absolute error (MAE) as the loss function in the pre-training stage. The urban road layout design diagram reconstructed by the optimized decoder (i.e. Figure 2 The reconstructed image in the figure) and the original urban road layout design map (i.e. Figure 2 The calculation formula is as follows:
[0066]
[0067] Among them, x is the input original urban road layout design map, A is the autoencoder; represents the expectation; ‖·‖1 represents the L1 norm.
[0068] The conditional generation method of the urban road layout diagram based on the potential diffusion model described in S1 to S3 of the above embodiments is applied to a specific case to demonstrate the technical effects that can be achieved, and some specific details are not repeated here.
[0069] Example
[0070] The overall process in this embodiment can be divided into four stages: data preprocessing, feature mapping module training, joint training of denoising module and conditional coding module, and image generation. The specific structures of the feature mapping module, denoising module and conditional coding module are as described above and will not be repeated here.
[0071] Step 1: Data preprocessing
[0072] Step 1.1: In this embodiment, the raw data is recorded in raster form. Each set of raw data includes four raster images: population density data, terrain elevation data, land use data, and road layout data. The first three are input data, and the last road layout data is the label, i.e., the actual road layout map. After preliminary preprocessing, each raw data set is aligned at its central longitude and latitude. The population density data, terrain elevation data, and land use data are all single-channel grayscale images of 512*512, while the road layout data is a three-channel color image of 512*512.
[0073] Step 1.2: Filter the raw data to remove meaningless or sparsely populated data sets, such as data with zero population, high altitude, or sparse roads. Finally, dilate the road layout data to amplify the features.
[0074] Step 2: Feature Mapping Module Training
[0075] Step 2.1: Divide the training dataset and the test dataset into batches in a ratio of 7:3, and divide the training dataset into batches according to a fixed batch size (in this embodiment, the batch size is 4), with a total batch size of N.
[0076] Step 2.2: Sequentially select a batch of training samples indexed as i from the training dataset, where i∈{0,1,…,N}. Use each batch of training samples to train the feature mapping module, and use the Adam optimizer for model optimization with a learning rate parameter set to 0.00002. During training, calculate the loss of each batch of training samples. To adjust all network parameters in the model. After 100 rounds of training, the model is optimized. Figure 9 This is a schematic diagram of the training results of the feature mapping module. It can be seen that the image reconstructed by this module is almost the same as the original input image.
[0077] Step 3: Joint training of denoising module and conditional encoding module
[0078] Sequentially select a batch of training samples with index i from the training dataset, where i∈{0,1,…,N}. Use each batch of training samples to jointly train the denoising module and the conditional encoding module. Use the Adam optimizer for model optimization, and set the learning rate parameter to 0.00002. During the training process, calculate the loss of each batch of training samples. To adjust all network parameters in the model. After 100 rounds of training, the model is optimized.
[0079] Step 4: Image Generation
[0080] A set of population density maps, terrain elevation maps, and land use maps are selected from the test dataset and input into the conditional coding module to generate the corresponding conditional coding maps. At the same time, a Gaussian noise is randomly sampled from the standard normal distribution and compared with the generated conditional coding maps. Figure 1 The model is fed into the denoising module and generates a city road layout feature map through the reverse denoising process. This feature map is then fed into the feature mapping module, which ultimately outputs a city road layout design map. The results show that even for unlearned test datasets, the model can still generate a city road layout design map that meets actual planning requirements based on the input population density, terrain elevation, and land use information. Figure 10 shown.
[0081] The embodiment described above is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Persons skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent substitution or equivalent transformation falls within the scope of protection of the present invention.
Claims
1. A conditional generation method for urban road layout graph based on potential diffusion model, characterized in that: The following steps are involved: S1. Obtain population density raster data, terrain elevation raster data, and land use raster data for the target area, and construct multimodal input data from these three raster data. Superimpose the different modal input data using the feature channel dimension, and input the superimposed input data into a pre-trained conditional coding module to achieve data compression and dimensionality reduction, thereby generating conditional coding maps at different levels. S2. Randomly sample a Gaussian noise from a standard Gaussian distribution and use it as the initial input of the pre-trained denoising module. At the same time, use the conditional coding maps of each level generated by the conditional coding module as conditional information. The denoising module is used to generate a noise image through iterative denoising in successive rounds, ultimately obtaining the corresponding urban road layout feature map. S3. Input the urban road layout feature map output by the denoising module into a pre-trained feature mapping module for decoding, and finally generate an urban road layout design map.
2. The method for generating a conditional formula for an urban road layout diagram based on a potential diffusion model according to claim 1, wherein: In step S1, the conditional coding module includes a first input module and four downsampling modules; the first input module is composed of three 3*3 convolutional layers cascaded in sequence, and the four downsampling modules all adopt the ResNet architecture; in the first downsampling module, the superimposed input data is processed by two 3*3 convolutional layers in sequence to obtain the first feature, and the superimposed input data is added to the first feature through a residual connection to obtain the second feature and input it into the 3*3 downsampling convolutional layer to obtain the output of the first downsampling module; each 3*3 convolutional layer is provided with a normalization layer and a Swish activation function.
3. The method for generating a conditional formula for an urban road layout diagram based on a potential diffusion model according to claim 2, wherein: In step S2, the denoising module adopts a U-Net structure. In the denoising module, the initial input Gaussian noise is processed by the second input module to generate a feature map with a target number of channels and passes through four encoder modules in sequence to generate feature maps of different sizes. The feature map finally output by the fourth encoder module is input to the middle layer module, and then the feature map output by the middle layer module is used as the input of the fourth decoder module. Each decoder module cross-layer connects its input with the feature map of the same size generated by the corresponding encoder module. At the same time, the output of each decoder module is fused with the conditional coding map of the corresponding size output by the conditional coding module through the conditional control module, and the fused feature map is used as the input of the next decoder module. Finally, the output of the first decoder module is processed by the output module to generate a feature map of the urban road layout.
4. The method for generating a conditional formula for an urban road layout diagram based on a potential diffusion model according to claim 3, wherein: The first two encoder modules of the denoising module have the same structure, and the last two encoder modules have the same structure; in the first encoder module, the input features are processed by two 3*3 convolution layers in sequence to obtain the third feature, and the input features are added to the third feature through residual connection to obtain the fourth feature. The fourth feature is processed by two 3*3 convolution layers in sequence to obtain the fifth feature, and the fourth feature is added to the fifth feature through residual connection to obtain the sixth feature. The sixth feature is processed by two 3*3 convolution layers in sequence to obtain the seventh feature, and the sixth feature is added to the seventh feature through residual connection to obtain the eighth feature and processed by the 3*3 downsampling convolution layer to obtain the output of the first encoder module; in the third encoder module, the input features are processed by two 3*3 convolution layers in sequence to obtain the seventh feature, and the sixth feature is added to the seventh feature through residual connection to obtain the eighth feature and processed by the 3*3 downsampling convolution layer to obtain the output of the first encoder module. After layer processing, the ninth feature is obtained. The input feature is added to the ninth feature through the residual connection to obtain the tenth feature. After the tenth feature is processed by the self-attention layer, the first self-attention feature is obtained. The first self-attention feature is processed by two 3*3 convolutional layers in sequence to obtain the eleventh feature. The first self-attention feature is added to the eleventh feature through the residual connection to obtain the twelfth feature. After the twelfth feature is processed by the self-attention layer, the second self-attention feature is obtained. After the second self-attention feature is processed by two 3*3 convolutional layers in sequence, the thirteenth feature is obtained. The second self-attention feature is added to the thirteenth feature through the residual connection to obtain the fourteenth feature. The fourteenth feature is processed by the self-attention layer and the 3*3 downsampling convolution layer in sequence to obtain the output of the third encoder module.
5. The method for generating a conditional formula for an urban road layout diagram based on a potential diffusion model according to claim 4, wherein: The 3*3 downsampling convolution layer of the encoder module in the denoising module is replaced with a 4*4 upsampling deconvolution layer to obtain a decoder module corresponding to the encoder module; the second input module and the output module are both a 3*3 convolution layer; in the intermediate layer module, the input feature is processed by two 3*3 convolution layers in sequence to obtain the fifteenth feature, the input feature is added to the fifteenth feature through a residual connection to obtain the sixteenth feature, the sixteenth feature is processed by two 3*3 convolution layers in sequence to obtain the seventeenth feature, the sixteenth feature is added to the seventeenth feature through a residual connection to obtain the output of the intermediate layer module.
6. The method for generating a conditional formula for an urban road layout diagram based on a potential diffusion model according to claim 5, wherein: In the conditional control module, the conditional coding map output by the conditional coding module is processed by the first convolution layer to obtain the bias parameter. The conditional coding map output by the conditional coding module is processed by the second convolution layer to obtain the scale parameter. The scale parameter is multiplied by the feature map output by the decoder module to obtain a weighted feature map. The weighted feature map is added to the bias parameter to obtain the output of the conditional control module.
7. The method for generating a conditional formula for an urban road layout diagram based on a potential diffusion model according to claim 6, wherein: In step S3, the feature mapping module adopts an autoencoder structure, which is composed of an encoder and a decoder. During the pre-training stage of the feature mapping module, the original urban road layout design drawing is first processed by the encoder to generate an encoded feature map, which is then input into the decoder to reconstruct the original urban road layout design drawing. In the inference stage of the feature mapping module, the urban road layout feature map generated by the denoising module is used as the input of the feature mapping module decoder, and after decoding, the final urban road layout design map is output.
8. The method for generating a conditional formula for an urban road layout diagram based on a potential diffusion model according to claim 7, wherein: The conditional coding module and the denoising module are jointly trained. During the joint training process, a training dataset consisting of real-world urban road layout maps is first obtained, and a city road layout map is randomly selected from the training dataset as a training image. The training image is encoded into a target feature map using the encoder of the feature mapping module, and a current iteration time step t is randomly selected from the interval [0, T], where T represents the maximum number of iteration steps. A Gaussian noise is randomly sampled from the standard Gaussian distribution, and the target feature map is denoised through a forward diffusion process to generate a noisy map corresponding to time step t. Time step t and the corresponding noisy map are then used as inputs to the denoising module, and the conditional coding map output by the conditional coding module is used as conditional information. Finally, the denoising module outputs an estimated noise image, and the mean square error loss between the noise image generated by the denoising module and the randomly sampled Gaussian noise is calculated. The mean square error loss is optimized to train the conditional coding module and the denoising module.
9. The method for generating a conditional formula for an urban road layout diagram based on a potential diffusion model according to claim 8, wherein: The time step t input to the denoising module must first be processed by the time encoding module to generate the corresponding time code, and then the time code is input to each encoder module and decoder module and added to the output features of the first 3*3 convolutional layer in the encoder module and decoder module to achieve deep fusion of time information and feature information; among them, the time encoding module is composed of two linear layers cascaded in sequence.
10. The method for generating a conditional formula for an urban road layout diagram based on a potential diffusion model according to claim 7, wherein: In step S3, during the pre-training stage of the feature mapping module, the mean absolute error between the urban road layout design map reconstructed by the decoder and the original urban road layout design map is used as the loss function.