An automated method for urban road layout design driven by multimodal data
Through the combination of multimodal data fusion and conditional generation adversarial network CGAN, the efficiency and quality problems in the urban road layout design automation method in the prior art are solved, and a high-quality road layout design diagram that meets geographic information constraints are generated.
Patent Information
- Application Number
- CN202211138806.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-09-19
AI Technical Summary
The existing urban road layout design automation method is difficult to meet the constraints of spatial geographic information while ensuring high efficiency and generation quality. The generated road layout quality is low and the connectivity is insufficient.
The multimodal data fusion module is used to fuse population density, terrain elevation and land use data, and the urban road layout design drawing is generated through the conditional generation adversarial network CGAN, and the U-Net structure generator and the PatchGAN structure discriminator are used for training and optimization.
The generated urban road layout design drawings have higher visual and structural quality, meet the constraints of spatial geographic information, and do not require professional knowledge in related fields, and can efficiently generate road layouts with diverse characteristics.
Smart Images

Figure CN115544613B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and computer vision, and particularly to an automated method for designing urban road layouts driven by multi-modal data. Background Art
[0002] In recent years, with the rapid development of deep learning technology and the field of computer vision, more and more traditional tasks have burst out with new power by means of deep learning technology, and the urban road layout modeling task is one of them. In urban planning, the planned roads need to have a certain connectivity and ensure that the roads can meet the needs of high traffic volume. In addition, road layout modeling is also applied to the game industry. Game players pursue rich and diverse game scenes, and the instantaneously generated virtual environment can enhance the game player experience. Road layout planning also plays an important role in autonomous driving vehicles, and virtual roads of various urban streets are created to test autonomous driving vehicles.
[0003] The existing automated methods for designing urban road layouts can be divided into two categories: First, procedural modeling. Procedural modeling designs urban road layouts based on a rule set defined by technicians. This manual design and modeling method is very time-consuming and inefficient. Although it can ensure that the roads meet the specified constraint rules of users, it is still not flexible enough. Second, automated generation algorithms based on deep learning technology. The generative adversarial network (GAN) is used to learn a low-dimensional generative model from a high-dimensional model containing shape features, and then a road layout map is generated through sampling points in the low-dimensional space, so as to perform road layout modeling faster and more efficiently. Although the urban road layouts obtained by this method are rich and diverse, they are often of low quality, unable to guarantee road connectivity, and unable to design roads by combining spatial geographical information and user requirements.
[0004] Therefore, it is necessary to design a new automated generation algorithm based on deep learning technology, which can not only ensure high efficiency and generation quality, but also constrain the road layout to meet the corresponding spatial geographical information. Summary of the Invention
[0005] The technical problem to be solved by the present invention is how to introduce multi-modal spatial geographical information into GAN to generate a high-quality urban road layout design drawing that meets the corresponding input conditions. The present invention provides an automated method for designing urban road layouts driven by multi-modal data.
[0006] The specific technical solutions adopted by the present invention are as follows:
[0007] A multi-modal data-driven automated method for urban road layout design, which is implemented as follows: Obtain multi-modal data consisting of population density raster data, terrain elevation raster data, and land use raster data of the target area. Input the multi-modal data into a pre-trained multi-modal data fusion module. The multi-modal data fusion module compresses and reduces the dimension of different modal data and outputs an encoded map. Then, a pre-trained Conditional Generative Adversarial Networks (CGAN) module captures the geospatial features and road texture features in the encoded map, and finally outputs the corresponding urban road layout design map;
[0008] The multi-modal data fusion module is an autoencoder, consisting of an encoder and a decoder; the input different modal data are stacked along the feature channel dimension and then input into the encoder, and the encoder outputs the fused encoded map; in the pre-training stage, the encoded map output by the encoder is input into the decoder, and the decoder outputs the reconstructed modal data. Through training, the reconstructed modal data is made to approach the original modal data; in the inference stage, the encoded map output by the encoder is directly used as the output of the multi-modal data fusion module;
[0009] The conditional generative adversarial network module consists of a generator and a discriminator; the generator takes the encoded map output by the multi-modal data fusion module as input and outputs the road layout design map; in the pre-training stage, the road layout design map output by the generator is input into the discriminator. The discriminator takes the real road network layout map and the encoded map as the other two inputs at the same time, and outputs patches of three scales of 128*128, 32*32, and 8*8, and discriminates the authenticity of the input map from three perspectives of shallow, middle, and deep features, and then alternately optimizes the discriminator and the generator; in the inference stage, the road layout design map output by the generator is directly used as the output of the conditional generative adversarial network module.
[0010] Preferably, both the encoder and the decoder in the multi-modal data fusion module adopt the U-Net model as the baseline model.
[0011] Preferably, in the pre-training stage of the multi-modal data fusion module, training is carried out by optimizing the mean absolute error between the reconstructed modal data and the original modal data.
[0012] Preferably, in the multi-modal data fusion module, the input population density raster data, terrain elevation raster data, and land use raster data are all rasterized images with a size of 512*512, and the size of the encoded map formed by fusing the three different modal rasterized images is also 512*512.
[0013] Preferably, the input of the generator is the encoded graph output by the encoder. The generator adopts a U-Net structure, which includes 8 downsampling modules, 7 upsampling modules and 1 output module; the first 4 downsampling modules are 4*4 convolutional layers with a normalization layer and a LeakyRelu activation function, and the last 4 downsampling modules include a 4*4 convolutional layer with a normalization layer and a LeakyRelu activation function and a Dropout layer; after passing through each downsampling module, the size of the feature map is halved; the first 4 upsampling modules include a 4*4 transposed convolutional layer with a normalization layer and a Relu activation function and a Dropout layer, and the last 3 upsampling modules are 4*4 transposed convolutional layers with a normalization layer and a Relu activation function; the output module includes an upsampling layer and a 3*3 convolutional layer with a Tanh activation function, and the upsampling layer uses nearest neighbor interpolation to double the image size; the 8 upsampling modules and the 7 downsampling modules are connected by a cross-layer connection (Skip-connection) structure.
[0014] Preferably, there are 3 discriminators, all of which adopt the PatchGAN structure. The input of each discriminator is the encoded graph output by the encoder, the real road network layout graph, and the road layout design graph output by the generator. The sizes of the three inputs of the discriminator are all 512*512, and the sizes of the output patches of the three discriminators are 128*128, 32*32, and 8*8 respectively, so as to discriminate the comprehensive authenticity of the input images at three feature sizes: shallow, middle, and deep.
[0015] Preferably, the three discriminators are specifically as follows:
[0016] The first discriminator consists of 2 downsampling modules and 1 output layer. The downsampling module includes a 4*4 convolutional layer with a normalization layer and a Relu activation function, and the output layer is a 3*3 convolutional layer. The output image is a single-channel patch of size 128*128.
[0017] The second discriminator consists of 4 downsampling modules and 1 output layer. The downsampling module includes a 4*4 convolutional layer with a normalization layer and a Relu activation function, and the output layer is a 3*3 convolutional layer. The output image is a single-channel patch of size 32*32.
[0018] The third discriminator consists of 6 downsampling modules and 1 output layer. The downsampling module includes a 4*4 convolutional layer with a normalization layer and a Relu activation function, and the output layer is a 3*3 convolutional layer. The output image is a single-channel patch of size 8*8.
[0019] Preferably, the loss function used for training the conditional generative adversarial network module is the sum of the losses of the three discriminators, and the loss function form of each discriminator is the MSE loss.
[0020] The present invention has the following beneficial effects compared with the prior art:
[0021] In the urban road network layout design task, the present invention introduces a multi-modal data fusion module to fuse three types of data: population density, terrain elevation, and land use, and inputs them into the CGAN module to generate road layouts that conform to the population, terrain, and land types. Compared with traditional manual design methods, this method does not require professional knowledge in related fields, is not limited by road templates, and can efficiently generate urban road layouts with diverse features. Compared with the method that only uses GAN, the method of the present invention introduces multi-modal geospatial data, compresses and reduces its dimension, avoids the problem of high computing power requirements caused by a large number of parameters, and highlights the key features of the data, making the generated urban road layouts of higher quality visually and structurally. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a structural diagram of a multi-modal data-driven urban road layout design model;
[0023] Figure 2 It is a schematic diagram of the generator structure;
[0024] Figure 3 It is a schematic diagram of the discriminator structure;
[0025] Figure 4 It is a schematic diagram of the generation result of the pre-trained multi-modal data fusion module for 200 rounds;
[0026] Figure 5 It is a schematic diagram of the optimization process of the road layout design drawing generated by the generator;
[0027] Figure 6 It is a schematic diagram of the test result of the road layout design drawing generation. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined correspondingly without conflict.
[0029] A modality refers to the way things occur or exist, and can also refer to the source or manifestation form of information. Multimodality refers to the combination of two or more modalities. Just as a person's understanding of an unfamiliar concept is based on the joint participation of vision, smell, hearing, and other tactile sensations, a model also needs to rely on data from different modalities to capture the comprehensive semantic information of things. Different modalities have different manifestation forms and can capture different features of the same thing. There is information redundancy among multimodal data, and they can also be combined with each other to generate richer semantic feature information. This is the redundancy and complementarity of multimodal data. The fusion of multimodal data is very necessary. However, due to the different manifestation forms of different modality information, the fusion of multimodal data is also very difficult. Currently, the mainstream approach is to use deep learning-based methods to handle the problem of multimodal data fusion.
[0030] One of the cores of the present invention is precisely the processing of multimodal data fusion. As an example, population density data, terrain elevation data, and land use data are used as three types of multimodal data, and an autoencoder is used to handle the problem of multimodal data fusion. It should be noted that the specific form of this multimodal data processing module is not limited, and other models with an encoder-decoder structure can be used.
[0031] Another core of the present invention is the improvement of the Conditional Generative Adversarial Networks (CGAN) and the embedded combination of the CGAN module and the multimodal data fusion module. The generator in the CGAN module adopts a U-Net structure, the discriminator adopts a PatchGAN structure and the Multi-scale Discriminator idea is used.
[0032] In a preferred embodiment of the present invention, an automated method for urban road layout design driven by multimodal data is provided. The overall structure of the model of this method is as Figure 1 . From Figure 1 it can be seen that this model is divided into two parts, the upper part is the multimodal data fusion module, which is an autoencoder and is divided into an encoder and a decoder. The input of this autoencoder is data from different modalities. In the present invention, three modalities of raster data are used, namely population density data, terrain elevation data, and land use data. The three modalities of data are all single-channel rasterized images with a size of 512*512. After being stacked along the image feature channel dimension, they are used as the final input of this multimodal data fusion module. In the multimodal data fusion module, the input image is fused by the encoder into an encoded image containing the semantic features of the three modalities. The size of this encoded image is a single-channel image with a size of 512*512. This encoded image has multimodal semantic information and at the same time avoids the problem of high computing power requirements brought by a large number of parameters. In addition, this encoded image can be restored to the original rasterized image through the decoder.
[0033] However, it should be noted that the decoder is mainly used in the model training stage. Its role is to assist the encoder in training. However, in the inference stage, the decoder does not participate in the inference process. Therefore, in the pre-training stage, the encoded map output by the encoder is input to the decoder, and the decoder outputs the reconstructed multi-modal data. Through training, the reconstructed multi-modal data is made to be close to the original multi-modal data. In the inference stage, the encoded map output by the encoder is directly used as the output of the multi-modal data fusion module. In this embodiment, in the pre-training stage of the multi-modal data fusion module, training can be performed by optimizing the mean absolute error between the reconstructed multi-modal data and the original multi-modal data, that is, by optimizing the L1 loss between the input image and the restored image output by the decoder to train this module, so that the reconstructed multi-modal data is close to the original multi-modal data.
[0034] In this embodiment, the objective function of the multi-modal data fusion model can be expressed as: A is an autoencoder, and x is the input after stacking all modal data in the feature channel dimension.
[0035] Such as Figure 1, the lower part is a CGAN module composed of a generator G and a discriminator D. The input of the generator is a single-channel encoded map with a size of 512*512 output by the encoder, and the output is a single-channel road network layout design map with a size of 512*512. There are 3 discriminators D, all of which adopt the PatchGAN structure, but output three different-sized patches respectively. This structure is called the Multi-scale Discriminator. The three discriminators discriminate the comprehensive authenticity of the input image at small, medium, and large sizes. Specifically, the input of the three discriminators is two sets of data. One set is a single-channel encoded map with a size of 512*512 and a real road layout map, and the other set is a single-channel encoded map with a size of 512*512 and a road layout design map. These two sets of data are downsampled by different discriminators to generate three patches of different sizes. In the present invention, the three discriminators are downsampled to three sizes of 128*128, 32*32, and 8*8 respectively, and these three sizes represent three kinds of features of the shallow layer, middle layer, and deep layer. Similarly, it should be particularly noted that the discriminator D is mainly used in the model training stage, and its role is to assist the generator G in training, but in the inference stage, the discriminator D does not participate in the inference process. Therefore, in the pre-training stage of the CGAN module, the road layout design map output by the generator is input into the discriminator, and the discriminator takes the real road network layout map and the encoded map as the other two inputs at the same time, and outputs patches of three scales of 128*128, 32*32, and 8*8, and discriminates the authenticity of the input map from three perspectives of shallow layer, middle layer, and deep layer features, and then alternately optimizes the discriminator and the generator; in the inference stage, the road layout design map output by the generator is directly used as the output of the conditional generative adversarial network module, that is, the final output result.
[0036] The training process of the CGAN module is the same as the traditional training method of the conditional generative adversarial network. The network can be trained in an alternating training method, that is, first train the discriminator network D and then train the generator network G. The loss function of the discriminator is in the form of MSE loss, that is, the Patch generated by the real road layout map calculates the MSE Loss with the all-ones matrix of the same size, and the Patch generated by the road layout design map calculates the MSE Loss with the all-zeros matrix of the same size. However, since the Multi-scale Discriminator structure uses three discriminators, the loss function used when calculating the loss needs to be the sum of the losses of the three discriminators.
[0037] In this embodiment, the specific generator and discriminator network structures and parameters in the CGAN module are optimized according to the actual urban road layout design scenario.
[0038] Specifically, the generator structure is as Figure 2As shown in the figure. The generator adopts the U-Net structure, which includes 8 downsampling modules (d1-d8), 7 upsampling modules (u1-u7), and 1 output module (u8). The downsampling module is a 4*4 convolutional layer (with a normalization layer and a LeakyRelu activation function). The upsampling module includes a 4*4 transposed convolutional layer (with a normalization layer and a Relu activation function). The use of a Dropout layer can be selected. In this example, the downsampling blocks numbered d5-d8 and the upsampling blocks numbered u1-u1 use the Dropout layer. The output layer includes an upsampling layer (using nearest neighbor interpolation) and a 3*3 convolutional layer (with a Tanh activation function). The upsampling module and the downsampling module are also connected using the Skip-connection structure. For example, the output of d2 will be input to d3, and at the same time, it will be stacked with the output of u6 in the channel dimension as the final output. Compared with the traditional encoder-decoder structure, the Skip-connection structure in U-Net can bring richer semantic information to the decoder during the upsampling process. Let d(x,y) represent the downsampling block with an input channel number of x and an output channel number of y. The parameter settings of the 8 downsampling blocks are: d(1,64)-d(64,128)-d(128,256)-d(256,512)-d(512,512)-d(512,512)-d(512,512)-d(512,512). Let u(x,y) represent the upsampling block with an input channel number of x and an output channel number of y. The parameter settings of the 7 upsampling blocks are: u(512,512)-u(1024,512)-u(1024,512)-u(1024,512)-u(1024,256)-u(512,128)-u(256,64).
[0039] In addition, the discriminator adopts the Multi-scale discriminator structure. In this embodiment, 3 discriminators are used, and the specific number can be adjusted according to the situation. The number of downsampling blocks of different discriminators is different, and all adopt the PatchGAN structure. As Figure 3 Figure 5 is the overall structure diagram of the discriminator. There are 6 downsampling modules in the schematic diagram, but this is only an exemplary illustration. The number of downsampling modules in the specific discriminator needs to be adjusted according to their respective parameters. The three discriminators discriminate the comprehensive authenticity of the input image at three sizes: small, medium, and large. The three sizes represent three types of features: shallow, middle, and deep.
[0040] The first discriminator consists of 2 downsampling modules and 1 output layer. The downsampling module includes a 4*4 convolutional layer (with a normalization layer and a Relu activation function). The output layer is a 3*3 convolutional layer. The output image is a single-channel Patch of size 128*128. This discriminator is used to evaluate the authenticity of the shallow features of the input image.
[0041] The second discriminator consists of 4 downsampling modules and 1 output layer. The downsampling module includes a 4×4 convolutional layer (with a normalization layer and a ReLU activation function). The output layer is a 3×3 convolutional layer, and the output image is a single-channel Patch with a size of 32×32. This discriminator is used to evaluate the authenticity of the intermediate-level features in the input image.
[0042] The third discriminator consists of 6 downsampling modules and 1 output layer. The downsampling module includes a 4×4 convolutional layer (with a normalization layer and a ReLU activation function). The output layer is a 3×3 convolutional layer, and the output image is a single-channel Patch with a size of 8×8. This discriminator is used to evaluate the authenticity of the deep-level features in the input image.
[0043] In this embodiment, the objective function of the CGAN can be expressed as: where G is the generator, D is the discriminator, x is the real road network layout map, and c is the corresponding encoded map. It should be particularly noted that since there are three discriminators, the total loss function needs to calculate the sum of the loss terms of the three discriminators. The process of training with the loss function in the CGAN module is the same as that of the traditional conditional generative adversarial network training method, so it will not be elaborated here.
[0044] After the multi-modal data fusion model and the CGAN model are trained, they can be applied in the inference stage. The specific method is as follows: Obtain the multi-modal data composed of the population density grid data, terrain elevation grid data, and land use grid data of the target area. Input the multi-modal data into the pre-trained multi-modal data fusion module. The multi-modal data fusion module compresses and dimensionality-reduces the different modal data and outputs the encoded map. Then, the pre-trained CGAN module captures the geospatial features and road texture features in the encoded map, and finally outputs the corresponding urban road layout design map.
[0045] Next, the multi-modal data-driven urban road layout design automation method in the above embodiment will be applied to a specific case to demonstrate the technical effects it can achieve.
[0046] Embodiment
[0047] The overall process in this embodiment can be divided into three stages: data preprocessing, multi-modal data fusion model training, CGAN model training, and image generation. The specific structures of the multi-modal data fusion model and the CGAN model are as described above, so they will not be elaborated here.
[0048] Step 1: Data preprocessing stage
[0049] Step 1.1: The original data in this case are population density data, terrain elevation data, land use data, and road layout data recorded in raster form. Four raster images form a set of sample data. The first three are used as input data of the sample, and the last road layout data is used as the label of the sample, that is, the real road layout map. Preliminary preprocessing is performed on each original raster image in the sample to ensure that the central longitude and latitude are aligned, and the size is a single-channel grayscale image of 512*512.
[0050] Step 1.2: The images are screened to remove data groups that are meaningless or have too few samples, such as data with a population of 0, too high altitude, and sparse roads. Finally, dilation processing is performed on the road layout map to magnify the features.
[0051] Step 2: Training of the multi-modal data fusion model
[0052] Step 2.1: The training data set and the test data set are divided in a ratio of 7:3, and the training data set is batched according to a fixed batch size, with a total of N.
[0053] Step 2.2: Sequentially select a batch of training samples with index i from the training data set, where i ∈ {0, 1, …, N}. Use each batch of training samples to train the multi-modal data fusion model. This model adopts an autoencoder structure. During the training process, calculate the of each training sample and use the gradient descent method to train. Adjust the network parameters in the entire model according to the total loss of all training samples in the batch. The learning rate parameter is set to 0.0002 and reduced to one-tenth of the previous value every 100 rounds. Figure 4 is a schematic diagram of the training results after 200 rounds. The image reconstructed and restored by this autoencoder is almost identical to the input image.
[0054] Step 3: Training of the CGAN model
[0055] Embed the encoder in the pre-trained multi-modal data fusion module in Step 2 into the CGAN model. This encoder can encode the original input data into a single-channel encoded map of 512*512. This encoded map is used as the input of the generator in the CGAN, and the generator generates the corresponding road layout design map according to this input. The discriminator will learn to distinguish between the real road layout map and the road layout design map. The training of the generator and the discriminator is a process of mutual game and continuous iterative optimization. The goal of the generator is to learn to deceive the discriminator and generate a road layout design map that looks real. The discriminator needs to distinguish between true and false. During the training process, calculate the and of each training sample and use the gradient descent method to train. Adjust the network parameters in the entire model according to the total loss Adjust the network parameters in the entire model. The learning rate parameters of the generator and discriminator are set to 0.0002 and reduced to one-tenth of the previous value every 100 rounds. Figure 5 For the optimization process of the generator to generate road layout design drawings, it can be seen that as the discriminator is continuously optimized, the generator is also constantly evolving. It can select and generate corresponding roads based on three types of data: terrain, population, and land use. After continuous optimization, the connectivity of the roads gradually improves, and the phenomenon of broken and fragmented roads significantly decreases.
[0056] 4. Image Generation
[0057] Input the test data set into the generator through the same process as above. It can be seen that for the unlearned test data set, the model can still generate corresponding road layout design drawings based on the input population density, terrain elevation, and land use data, as Figure 6 .
[0058] The embodiments described above are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by adopting equivalent substitution or equivalent transformation methods fall within the protection scope of the present invention.
Claims
1. A multi-modal data-driven automated method for urban road layout design, characterized in that, Obtain the multi-modal data composed of the population density raster data, terrain elevation raster data, and land use raster data of the target area, input the multi-modal data into the pre-trained multi-modal data fusion module. The multi-modal data fusion module compresses and reduces the dimension of different modal data and outputs an encoded map. Then, the pre-trained Conditional Generative Adversarial Networks (CGAN) module captures the geospatial features and road texture features in the encoded map, and finally outputs the corresponding urban road layout design map; The multi-modal data fusion module is an autoencoder, consisting of an encoder and a decoder; the input different modal data are stacked along the feature channel dimension and then input to the encoder, and the encoder outputs the fused encoded map; in the pre-training stage, the encoded map output by the encoder is input to the decoder, and the decoder outputs the reconstructed modal data. Through training, the reconstructed modal data is made to approach the original modal data; in the inference stage, the encoded map output by the encoder is directly used as the output of the multi-modal data fusion module; The conditional generative adversarial network module consists of a generator and a discriminator; the generator takes the encoded map output by the multi-modal data fusion module as input and outputs the road layout design map; in the pre-training stage, the road layout design map output by the generator is input to the discriminator. The discriminator takes the real road network layout map and the encoded map as the other two inputs at the same time, and outputs patches of three scales, 128*128, 32*32, and 8*8, to judge the authenticity of the input map from three perspectives of shallow, middle, and deep features, and then alternately optimizes the discriminator and the generator; in the inference stage, the road layout design map output by the generator is directly used as the output of the conditional generative adversarial network module.
2. The multi-modal data-driven automated method for urban road layout design according to claim 1, wherein, Both the encoder and decoder in the multi-modal data fusion module use the U-Net model as the baseline model.
3. The multimodal data-driven automated urban road layout design method according to claim 2, wherein In the pre-training stage, the multi-modal data fusion module is trained by optimizing the mean absolute error between the reconstructed modal data and the original modal data.
4. The multimodal data-driven automated method for urban road layout design according to claim 1, wherein In the multi-modal data fusion module, the input population density raster data, terrain elevation raster data, and land use raster data are all rasterized images with a size of 512*512, and the size of the encoded map formed by fusing the three rasterized images of different modalities is also 512*512.
5. The multimodal data-driven automated method for urban road layout design according to claim 1, characterized in that The input of the generator is the encoded graph output by the encoder. The generator adopts a U-Net structure, which includes 8 downsampling modules, 7 upsampling modules and 1 output module. The first 4 downsampling modules are 4*4 convolutional layers with a normalization layer and a LeakyRelu activation function. The last 4 downsampling modules include a 4*4 convolutional layer with a normalization layer and a LeakyRelu activation function, as well as a Dropout layer. After passing through each downsampling module, the size of the feature map is halved. The first 4 upsampling modules include a 4*4 transposed convolutional layer with a normalization layer and a Relu activation function, as well as a Dropout layer. The last 3 upsampling modules are 4*4 transposed convolutional layers with a normalization layer and a Relu activation function. The output module includes an upsampling layer and a 3*3 convolutional layer with a Tanh activation function. The upsampling layer uses nearest neighbor interpolation to double the image size. The 8 upsampling modules and 7 downsampling modules are connected by a cross-layer connection (Skip-connection) structure.
6. The multi-modal data-driven automated urban road layout design method according to claim 4, characterized in that There are 3 discriminators, all of which adopt the PatchGAN structure. The input of each discriminator is the encoded graph output by the encoder, the real road network layout graph, and the road layout design graph output by the generator. The sizes of the three inputs of the discriminator are all 512*512. The output patch sizes of the three discriminators are 128*128, 32*32, and 8*8 respectively, so as to discriminate the comprehensive authenticity of the input image at three feature sizes: shallow, middle, and deep.
7. The multimodal data-driven automated urban road layout design method according to claim 6, characterized in that The specific structures of the three discriminators are as follows: The first discriminator consists of 2 downsampling modules and 1 output layer. The downsampling module includes a 4*4 convolutional layer with a normalization layer and a Relu activation function. The output layer is a 3*3 convolutional layer, and the output image is a single-channel patch of size 128*128. The second discriminator consists of 4 downsampling modules and 1 output layer. The downsampling module includes a 4*4 convolutional layer with a normalization layer and a Relu activation function. The output layer is a 3*3 convolutional layer, and the output image is a single-channel patch of size 32*32. The third discriminator consists of 6 downsampling modules and 1 output layer. The downsampling module includes a 4*4 convolutional layer with a normalization layer and a Relu activation function. The output layer is a 3*3 convolutional layer, and the output image is a single-channel patch of size 8*8.
8. The multimodal data-driven automated urban road layout design method according to claim 1, characterized in that The loss function used in the training of the conditional generative adversarial network module is the sum of the losses of the three discriminators. The loss function form of each discriminator is the MSE loss.
Citation Information
Patent Citations
Multi-modal human activity recognition method based on generative adversarial network
CN110309861A
Path planning algorithm based on multi-modal multi-objective optimization algorithm
CN113704370A