A 3D scene controllable generation method and system for autonomous driving simulation
Through the graph dynamic attention codec based on conditional variation learning and CLIP model, the problems of insufficient diversity of generation results and poor visual quality are solved, and controllable three-dimensional scene generation in autonomous driving simulation is realized, which is suitable for scene boundaries of actual road network shapes.
Patent Information
- Application Number
- CN202510303775.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-14
AI Technical Summary
When generating autonomous driving simulation scenarios, the generation results are insufficient, the visual quality is poor, and it is difficult to generate controllable based on given boundaries and attributes. The existing methods cannot be applied to any shape scene boundaries divided by road networks in actual situations.
A graph dynamic attention codec based on conditional variation learning is adopted, combined with the scene asset matching strategy of the CLIP model, and controllable generation is achieved by obtaining the road network topology diagram, building the geometric minimum ring, generating graph structure data, and optimizing the matching three-dimensional model.
It improves the diversity and visual quality of generated three-dimensional scenes, and can generate controllable three-dimensional scenes that meet a given boundary and attributes, which are suitable for autonomous driving simulation testing.
Smart Images

Figure CN119810364B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer technology and autonomous driving, and particularly relates to a method and system for controllably generating a three-dimensional scene for autonomous driving simulation. Background Art
[0002] Scene generation technology provides a large number of three-dimensional scenes for sensing perception, supports the exploration of the impact of scene elements on the performance of perception algorithms, and is of great significance for model training and perception algorithm optimization. In a multi-vehicle autonomous driving system, intelligent vehicles often perceive the driving environment based on sensors and process and analyze the perception results based on artificial intelligence algorithms. Autonomous driving algorithms based on artificial intelligence require a large amount of data for training and testing. Furthermore, researchers often need to answer the question of "what kind of scenes are exactly suitable for our autonomous driving algorithms". Building simulation scenes in the virtual world is a very promising solution. However, most of the existing methods still rely on manual construction, and the number and diversity of scenes are limited. Therefore, this paper studies the automated generation of a large number of three-dimensional scene models with different characteristics for sensing perception, which has important practical significance for the training of autonomous driving perception models, the mutual adaptation sensitivity of scene elements and perception algorithms, and the optimization of perception algorithms.
[0003] In recent years, researchers have begun to design methods based on generative artificial intelligence to synthesize a large number of scenes. However, the existing image-based methods learn pixel-level relationships in building layout data rather than modeling the overall building instances, and the generated results often have relatively blurred boundaries. The research based on vector graphics directly models building instances to directly generate vector building outlines. However, such methods often generate based on rectangular boundaries and cannot be applied to the arbitrarily shaped scene boundaries divided by road networks in actual situations. In addition, the results generated by these methods are single and all rectangular, lacking diversity. Therefore, this paper models building outlines into various shapes and proposes a conditional variational learning model based on graph dynamic attention to achieve conditional generation based on arbitrary boundaries and given overall layout attributes. Aiming at the problem of inconsistent synthetic images in the three-dimensional scene generation method based on rendering, this paper proposes an asset matching strategy to directly generate a three-dimensional urban scene model, and then generates a large number of diverse three-dimensional scenes for sensing perception for autonomous driving simulation. Summary of the Invention
[0004] To solve the problems in the prior art, the present invention proposes a method and system for controllably generating a three-dimensional scene for autonomous driving simulation.
[0005] The technical solution adopted by the present invention is as follows:
[0006] In the first aspect, the present invention discloses a method for controllably generating a three-dimensional scene for autonomous driving simulation, including:
[0007] 1) Obtain the road network topology graph to be used for 3D scene generation, and construct a geometric minimum cycle set based on the road network topology graph;
[0008] 2) Divide the scene plane corresponding to the road network topology graph with the geometric minimum cycle as the boundary, obtain multiple sub-regions, assign attribute features to each sub-region, and then generate binary images corresponding to each sub-region; Obtain the trained convolutional encoder and use it to extract the encoded vectors of the binary images;
[0009] 3) Obtain the trained graph dynamic attention decoder, and use it to decode the attribute features, encoded vectors, and hidden layer vectors sampled from the standard normal distribution to obtain the graph structure data of each sub-region;
[0010] 4) Convert the graph structure data into a 2D layout, and splice all the 2D layouts to obtain the scene plane layout;
[0011] 5) Construct a scene information table based on the scene plane layout, and obtain a 3D model library; Traverse the floor plan outlines of each building in the scene information table, and screen out 3D models with a contour matching degree higher than the threshold from the 3D model library to construct a candidate set;
[0012] 6) Perform optimization matching based on the positional relationship and feature similarity of the buildings in the scene plane layout, and finally perform a 3D model replacement operation to generate the final 3D scene.
[0013] In a second aspect, the present invention discloses a 3D scene controllable generation system for autonomous driving simulation for implementing the above method, including:
[0014] An encoded vector and attribute feature acquisition module, which is used to obtain the road network topology graph to be used for 3D scene generation, and construct a geometric minimum cycle set based on the road network topology graph;
[0015] Divide the scene plane corresponding to the road network topology graph with the geometric minimum cycle as the boundary, obtain multiple sub-regions, assign attribute features to each sub-region, and then generate binary images corresponding to each sub-region; Obtain the trained convolutional encoder and use it to extract the encoded vectors of the binary images;
[0016] A graph structure data acquisition module, which is used to obtain the trained graph dynamic attention decoder, and use it to decode the attribute features, encoded vectors, and hidden layer vectors sampled from the standard normal distribution to obtain the graph structure data of each sub-region; A scene plane layout acquisition module, which is used to convert the graph structure data into a 2D layout, and splice all the 2D layouts to obtain the scene plane layout;
[0017] A candidate set acquisition module, which is used to construct a scene information table based on the scene plane layout and obtain a 3D model library; traverse the floor plan outlines of each building in the scene information table, and screen 3D models with a contour matching degree higher than a threshold value from the 3D model library to construct a candidate set;
[0018] A three-dimensional scene generation module, which is used to perform optimized matching based on the positional relationship and feature similarity of buildings in the scene plane layout, and finally perform a 3D model replacement operation to generate a final three-dimensional scene.
[0019] Compared with the prior art, the present invention has the following beneficial technical effects:
[0020] (1) The present invention proposes a method for controllably generating a three-dimensional scene for autonomous driving simulation, aiming to controllably generate a large number of three-dimensional scenes with different characteristics and available for sensing perception, and apply them to autonomous driving simulation tests. Aiming at the problems of weak diversity of generation results and poor visual quality existing in the prior art, the present invention proposes a graph dynamic attention encoder-decoder trained based on conditional variational learning. The trained graph dynamic attention decoder decodes the attribute features and encoded vectors obtained from the road network topology graph for three-dimensional scene generation to be performed, as well as the hidden layer vectors sampled from the standard normal distribution, to obtain graph structure data, improving the diversity and visual quality of the generated three-dimensional scene results. Aiming at the problem that it is difficult to perform controllable generation based on given boundaries and attributes in the prior art, when training the graph dynamic attention encoder-decoder, the present invention introduces virtual nodes into the graph structure data input to the graph dynamic attention encoder to enhance the intensity of attribute features in conditional control, so that the generated scene can meet the limitations of the input boundaries and scene attributes. In addition, the present invention proposes a scene asset matching scheme based on the CLIP model, that is, when optimizing the matching of buildings and 3D models, the CLIP image encoder is used to obtain the material features of the 3D models, and a three-dimensional scene available for sensor perception is generated.
[0021] (2) The present invention proposes a system for controllably generating a three-dimensional scene for autonomous driving simulation. The included scene plane layout data acquisition and processing module can efficiently and comprehensively collect the training data required by the scene plane layout controllable generation method; the included scene quality evaluation module can evaluate the quality of the generated three-dimensional scene in both qualitative and quantitative ways, and verify the effectiveness of the three-dimensional scene controllable generation method. Description of the Drawings
[0022] Figure 1 It is a structural diagram of the system for controllably generating a three-dimensional scene for autonomous driving simulation proposed by the present invention.
[0023] Figure 2It is the structural diagram of the convolutional-based encoder-decoder and the generated result diagram of the encoder-decoder for the three-dimensional scene local controllable generation method proposed by the present invention.
[0024] Figure 3 It is the training flow chart of the graph dynamic attention encoder-decoder for the three-dimensional scene local controllable generation method proposed by the present invention.
[0025] Figure 4 It is the schematic diagram of the process of adding virtual nodes to the graph dynamic attention encoder in the three-dimensional scene controllable generation method proposed by the present invention.
[0026] Figure 5 It is the flow chart of the matching between buildings and three-dimensional models in the three-dimensional scene controllable generation method proposed by the present invention.
[0027] Figure 6 It is the flow chart of the three-dimensional model optimization matching based on CLIP in the three-dimensional scene generation method proposed by the present invention.
[0028] Figure 7 It is the result diagram of the scene plane layout obtained by the scene plane layout generation method proposed by the present invention and the comparison diagram with the results of the existing pix2pix model.
[0029] Figure 8 It is the result diagram of the three-dimensional scene obtained by the three-dimensional scene generation method proposed by the present invention and the comparison diagram with the results of the existing SGAM model. Specific embodiments
[0030] The present invention will be further elaborated and described below in conjunction with specific embodiments. The described embodiments are only examples of the present disclosure and do not delimit the scope of limitation. The technical features of each embodiment of the present invention can be combined correspondingly without conflict.
[0031] Based on the generative artificial intelligence model, the present invention proposes a three-dimensional scene controllable generation method and system for autonomous driving simulation, aiming to controllably generate a large number of three-dimensional scene models with different characteristics for sensing perception and apply them to autonomous driving simulation tests.
[0032] In the first aspect, the present invention provides a three-dimensional scene controllable generation method for autonomous driving simulation, including the following steps:
[0033] S1: Obtain the road network topology graph to be used for three-dimensional scene generation, and construct a geometric minimum loop set based on the road network topology graph.
[0034] The process of finding the scene plane layout enclosed by the road network layout can be modeled as a problem of searching for the geometric minimum cycle from a road network topology graph. The present invention defines the geometric minimum cycle as the smallest geometric closed figure that does not contain any circular structure under the restriction of the plane coordinate system. The present invention proposes a geometric minimum cycle search method, which analyzes the simple paths from a random node to all its adjacent nodes and iteratively deletes the remaining shortest-length cyclic paths based on these paths.
[0035] When obtaining the geometric minimum cycle, first obtain the road network topology graph to be used for three-dimensional scene generation. Among them, the road network topology graph includes nodes and edges. Nodes represent intersections or road bends, and edges represent the roads (i.e., vehicle-passable areas) between two nodes. The node features of the nodes include the positions of the nodes (represented by horizontal and vertical coordinates). Then remove all nodes with a degree of 1 in the road network topology graph (these nodes cannot form a cycle and belong to isolated chains) to simplify the road network topology graph. Next, traverse all the remaining nodes v with a degree of 2 in the road network topology graph, and perform the following operations on each node v: find its two neighbor nodes v1 and v2, and search for all simple paths (i.e., paths without repeating nodes) from node v to neighbor node v1. Then, check whether these paths contain the other neighbor node v2. If there is a path that contains neighbor node v2, connect this path with the current node v to form a cycle, and find the shortest-length cycle among them. Subsequently, take the shortest-length cycle as the geometric minimum cycle, and remove the current node v from the road network topology graph to avoid repeated calculation. Then remove all nodes with a degree of 1 in the road network topology graph at the current moment, then traverse all the remaining nodes with a degree of 2 in the road network topology graph, and find the geometric minimum cycle, and then remove the current node from the road network topology graph. Finally, repeat the above steps until there are no nodes remaining in the road network topology graph. At this time, all the obtained geometric minimum cycles form the geometric minimum cycle set C.
[0036] S2: Divide the scene plane corresponding to the road network topology graph with the geometric minimum cycle as the boundary, construct a binary image based on the geometric minimum cycle, then obtain the trained convolutional encoder, and use the convolutional encoder to encode the binary image to obtain the boundary encoding vector; and assign attribute features to each sub-region. Among them, the boundary encoding vector is the encoding vector; the attribute features include the density of the building group and the height of the building group in the sub-region. The density of the building group is divided into high density and low density. When the density of the building group in the sub-region is greater than or equal to 0.1 (the proportion of the building area in the sub-region area), the density of the building group is high density, otherwise it is low density; the height is divided into high-rise and low-rise. When the average height of the building group is greater than or equal to 20m, it is high-rise, otherwise it is low-rise.
[0037] In a specific embodiment of the present invention, the scene plane corresponding to the road network topology map is divided into sub-regions by the geometric minimum loop, and the geometric minimum loop is used as the boundary of the sub-regions of the scene plane layout.
[0038] Step S2 uses the trained convolutional encoder-decoder to model the boundaries of the sub-regions of the scene plane to obtain boundary encoding vectors, so as to generate the building plane layout under subsequent specified boundary conditions. As shown in the appendix Figure 2 As shown, the core structure of the convolutional encoder-decoder consists of a convolutional encoder and a convolutional decoder which are composed of two parts; the convolutional encoder consists of 4 convolutional layers, and the convolutional decoder is a mirror image of the convolutional encoder. When obtaining the boundary encoding vector, first obtain the coordinates of each node of the geometric minimum loop, then create a blank image with the same size as the scene plane, and the initial pixel values are all 0. Then draw the geometric minimum loop in the image based on the node coordinates of the geometric minimum loop, and set the pixel values of the area inside the geometric minimum loop to 1, and the pixel values outside the geometric minimum loop remain 0. Finally, a binary plane boundary image of the scene plane with the area inside the geometric minimum loop being 1 and the area outside the loop being 0 is generated. The scene plane binary plane boundary image is segmented based on the boundary defined by the geometric minimum loop to obtain the binary image corresponding to each geometric minimum loop itself. At the same time, when inputting the binary image into the trained convolutional encoder, it is necessary to unify the size of the binary image. For example, in a specific embodiment of the present invention, the size of the binary image can be 256×256. The trained convolutional encoder compresses the input binary image into a boundary encoding vector . During training, the convolutional decoder takes the boundary encoding vector compressed by the convolutional encoder as the input, outputs the restored binary image, and uses the L1 distance between the restored binary image of the convolutional decoder and the original binary image as the loss function to update the parameters of the convolutional encoder-decoder, realizing self-supervised training.
[0039] Before training the convolutional codec, it is first necessary to prepare the training set. In this embodiment, the road network topology map in the real scenario and the corresponding scene plane layout vector map are collected from an online website. Then, based on the road network topology map, the scene plane corresponding to the scene plane layout vector map is divided into several sub-regions. Specifically, the scene plane is divided into several sub-regions according to the geometric minimum loop search method proposed in step S1 of the present invention. Then, the scene plane layout vector map of the sub-region is rasterized into image data. Among them, the method steps for rasterizing the scene plane layout of the sub-region are specifically as follows: First, select the geographic information coordinate system and determine the scaling scale of the image corresponding to the scene plane layout vector map of the sub-region (that is, the actual distance represented by one pixel in the image). Then, set the value inside the building polygon in the scene plane layout vector map of the sub-region to the building height, and set the rest of the region to 0. Then, this embodiment completes the missing building height information in the collected data. For the building layout samples, this embodiment observes that the building outlines extracted from the online website often lack height information and there are partial missing building samples. Therefore, the present invention combines it with the publicly available dataset containing building height information.
[0040] Specifically, for the Chinese region, this embodiment uses the height information in the CNBH dataset [WU W B, MA J, BANZHAF E, et al. A first Chinese building height estimate at 10 m resolution (CNBH-10m) using multi-source earth observations and machine learning [J]. Remote Sensing of Environment, 2023, 291: 113578-113591.]; for foreign regions such as the United States, Australia, and Europe, this embodiment uses the dataset provided by Microsoft [https: / / github.com / microsoft / GlobalMLBuildingFootprints]. The key to the processing process is the data alignment method and the data merging rule. Therefore, this embodiment standardizes the coordinate reference system of all building contour spaces to the pseudo-Mercator projection (EPSG:3857) to align the data from the three different sources in terms of spatial coordinates. Then, height information is assigned to each building in the original scene plane. When merging the two sets of building contours, taking the CNBH dataset as an example, this article follows the following rules: using the OSM (OpenStreetMap) data as a benchmark, indexing the corresponding area in the CNBH dataset, and calculating the average height value of the effective pixels. Finally, the average height value is assigned to the corresponding building as height information. Specifically: for the areas where the building height values in the sub-regions are missing, this embodiment uses the nearby buildings within a radius of 300 meters to estimate the height, updates the scene plane layout vector map of the sub-region, and finally the sub-regions of the entire scene plane and the updated scene plane layout vector map of the sub-region constitute the dataset, which is used as the training data for the graph dynamic attention encoder-decoder for obtaining the trained graph dynamic attention decoder. Finally, this embodiment calculates the average height of the buildings in the sub-region and the density of the building groups, and then assigns attribute features to each sub-region based on the calculated results. The attribute features include the density of the building groups and the height of the building groups. The density of the building groups is divided into high density and low density, and the height of the building groups is divided into high-rise and low-rise. Finally, the division of the scene plane sub-regions is completed and corresponding attribute features are assigned to each sub-region.
[0041] When training the convolutional encoder-decoder, as shown in the appendix Figure 2As shown, each sub-region in the dataset is converted into a binary image, and the sizes of the obtained binary images are unified. Then, all the unified binary images form a training set for training the convolutional autoencoder, and the convolutional autoencoder is trained. After obtaining the trained convolutional autoencoder, the trained convolutional encoder is taken to perform the operation in step S2.
[0042] S3: Obtain the trained graph dynamic attention decoder. Use the attribute features and boundary encoding vectors of the sub-regions and the hidden layer vectors sampled from the standard normal distribution as the input of the graph dynamic attention decoder to obtain the graph structure data of each sub-region.
[0043] Specifically, step S3 uses the boundary encoding vectors of the sub-regions corresponding to the geometric minimum rings obtained in step S2, the attribute features of the sub-regions in step S2, and the hidden layer vectors sampled from the standard normal distribution as the input of the trained graph dynamic attention decoder to obtain the graph structure data of each sub-region, that is, generate the scene plane layout results of the buildings in each sub-region.
[0044] The trained graph dynamic attention decoder is obtained by training the graph dynamic attention autoencoder. When training the graph dynamic attention autoencoder, the preparation of training data can be carried out first. That is, during training, the dataset for training the graph dynamic attention autoencoder is prepared according to the above steps. Among them, the scene plane layout vector map of the sub-region is the building layout vector map inside the sub-region. Then, the scene plane layout vector map of each sub-region in the dataset is normalized and expressed, and converted into graph structure data.
[0045] In this embodiment, in order to convert the scene plane layout into structured data, that is, into graph-structured data, to facilitate the processing and learning of deep learning models, the present invention respectively performs a standardized expression on the building monomers in the sub-region and the overall layout of the sub-region. Combining with the sample distribution in the building layout dataset used in this embodiment, the present invention divides the existing buildings into six categories: rectangular, "X"-shaped, "L"-shaped, "I"-shaped, "U"-shaped, and oval-shaped. The key to standardizing the expression of existing building monomers is to determine the type of the building monomer and solve the corresponding parameter set. Since this problem belongs to an unconstrained non-linear optimization problem and the derivative is difficult to calculate during the solution process, this embodiment uses Powell's method [VASSILIADIS V S, CONEJEROS R. Powell methodPowell Method [M] / / FLOUDAS C A, PARDALOS P M. Encyclopedia of Optimization. Boston; Springer. 2001: 2001-2003.] to solve it, obtaining the graph-structured data of the scene plane layout, that is, the graph-structured data. The nodes in the graph-structured data represent the standardized building monomers, and the edges represent the connection relationships between the nodes, that is, the relative position relationships between the building monomers.
[0046] Then, a graph dynamic attention encoder is constructed. As an important model for graph representation learning, as shown in the appendix Figure 3 In this embodiment, a learnable dynamic attention mechanism is introduced into the graph neural network to form the graph dynamic attention encoder, so as to improve the encoding performance of the graph-structured data of the scene plane.
[0047] In the controllable generation task of the scene plane layout of the present invention, the overall attributes of the scene plane layout are the key points of concern. Although the graph dynamic attention encoder enhances the expression ability of the model, it does not change the fact that the original graph neural network only focuses on the local features of neighboring nodes. Therefore, the graph dynamic attention encoder still lacks attention to the overall features of the graph. Therefore, this embodiment adds virtual nodes to the graph-structured data of the scene plane layout to strengthen the graph dynamic attention encoder's attention to the overall features of the graph-structured data of the scene plane layout. Specifically, as shown in the appendix Figure 4 In this embodiment, on the basis of the graph-structured data of the scene plane layout obtained after the standardized expression and conversion, a virtual node connected to all the nodes in the original graph-structured data of the scene plane layout is added to obtain the first graph-structured data, so as to better retain the global structure and focus on the overall attributes of the scene plane during the graph representation learning process, and realize better conditional control of the overall attributes of the scene plane layout during the generation process.
[0048] This embodiment defines each node in the first graph structure data whose node feature is , and the node feature represents each key attribute of the building monomer. As shown in the formula:
[0049]
[0050] wherein, represents whether there is actually a building monomer at node ; represents the type of the building monomer at node ; represents the abscissa of the center position of the building monomer at node ; represents the ordinate of the center position of the building monomer at node ; represents the height of the building monomer at node ; represents the length of the minimum rotated bounding box of the floor plan contour of the building monomer at node ; represents the width of the minimum rotated bounding box of the floor plan contour of the building monomer at node ; represents the proportion of the floor plan contour of the building monomer at node in the minimum rotated bounding box.
[0051] Experimental observations show that the calculation process of the attention weights of the graph neural network is static, that is, for different query nodes, the sorting of the importance degrees of the obtained nodes is determined and independent of the query nodes. Therefore, this embodiment adopts dynamic attention to improve the diversity of the generation results of the graph neural network. Specifically, in the traditional graph neural network, the calculation of the attention weights depends on the scoring function obtained by successive operations of the weight matrix and the vector. During the training process, this operation method may cause the attention weights to collapse into a linear layer, thus limiting the expressive ability of the model. To solve this problem, this embodiment first concatenates the query node feature (i.e., the node feature of node ) and the key node feature (i.e., the node feature of the neighbor node ), and then uses the learnable weight matrix to perform a linear transformation on the concatenated result. Then, after using the LeakyReLU activation function on it, it is multiplied by the attention parameter vector to calculate the relevance coefficient , and this process is formally expressed as the formula:
[0052] ;
[0053] Among them, is a learnable weight matrix that maps the original feature vector to a high-dimensional space; is the node and the neighbor node The correlation coefficient between them, that is, the influence coefficient of the node on the neighbor node ; Indicates the concatenation operation; is the attention parameter vector; is the activation function; is the node feature of the node ; is the node feature of the neighbor node ;
[0054] After obtaining the correlation coefficient , perform weighted summation with the node features of the neighbor nodes and pass through a non-linear activation function to obtain the final output feature vector of each node . To enhance the model's expressive ability, this paper adopts the multi-head attention mechanism, that is, uses multiple attention mechanisms to independently calculate the attention coefficients and concatenate or average their outputs. As shown in the formula:
[0055] ;
[0056] ;
[0057] Among them, represents the weight coefficient between the node and the neighbor node ; σ is the non-linear activation function, is the number of attention groups, is the layer of the attention layer The weight coefficient between the node and the neighbor node calculated by the group of attention mechanisms, layer of the attention layer is the input linear transformation matrix of the group of attention mechanisms, is the node The feature vector after being calculated by the layer of the attention layer, will be input into the next layer of the attention layer for iterative calculation; is the set of neighbor nodes of the node ; is the neighbor node of the node of the node The feature vectors calculated by the -layer attention layer, where .
[0058] In this embodiment, the feature vectors of each node are integrated into a feature matrix , where represents the feature vector input to the -layer attention layer (i.e., the feature vector output by the -layer). Therefore, the above graph dynamic attention encoder can be formally described by the formula:
[0059] ;
[0060] where is the function of the graph dynamic attention encoder; is the feature matrix of the -layer.
[0061] In this embodiment, the feature matrix output by each layer is aggregated into an aggregated vector .
[0062] The graph dynamic attention decoder also uses the above network model of the graph dynamic attention encoder as the backbone network, which is just a mirror image of the graph dynamic attention encoder. Specifically, using the boundary encoding vector of the sub-region and the attribute features of the sub-region as the input of the graph dynamic attention decoder, and sampling the aggregated vector to obtain the hidden layer vector of the graph dynamic attention decoder, and the hidden layer vector is also used as the input of the graph dynamic attention decoder. First, the graph dynamic attention decoder takes the hidden layer vector as the input and sends the result into a multi-layer perceptron to obtain the initial feature matrix , is the feature vector of node after being decoded by the 0th layer of the graph dynamic attention decoder; then message passing and feature learning are performed layer by layer, and after layers, the final feature matrix is obtained, where each feature component is decoded by different multi-layer perceptrons to obtain the node features corresponding to each building monomer, where indicates whether there is actually a building monomer at the decoded node ; indicates the type of the building monomer at the decoded node ; indicates the decoded node The abscissa of the center position of the building unit at Indicates the decoded node The ordinate of the center position of the building unit at Indicates the decoded node The height of the building unit at Indicates the decoded node The length of the minimum rotated bounding box of the floor plan outline of the building unit at Indicates the decoded node The width of the minimum rotated bounding box of the floor plan outline of the building unit at Indicates the decoded node The proportion of the floor plan outline of the building unit at in the minimum rotated bounding box.
[0063] In this embodiment, the graph dynamic attention encoder-decoder is trained by conditional variational learning. During the training process, conditional variational learning is carried out using a conditional variational autoencoder. The goal of the conditional variational autoencoder is to learn and model the complex distribution of the building floor plans in the dataset , which represents the distribution of the first graph structure data and attribute features under the condition of the given boundary encoding vector . In this paper, common methods in conditional autoencoders are used for training, that is, the model models the distribution by maximizing its variational lower bound (evidence lower bound, ELBO). The loss function of its training process is as follows:
[0064] ;
[0065] Among them, is the loss function; is a hyperparameter; is the expectation of the posterior distribution of the hidden layer vector under the condition of the boundary encoding vector (encoding vector) , attribute features and the first graph structure data ; represents the posterior distribution of the hidden layer vector under the condition of the boundary encoding vector , attribute features and the first graph structure data ; that is, the distribution of the encoder output result. is the boundary encoding vector , attribute features and the hidden layer vector Under the condition of distribution of is the hidden layer vector prior distribution of, which is set as the standard normal distribution in this paper; is and KL divergence distance between the two distributions. is a hyperparameter; represents the attribute att of the building in the graph structure data input to the encoder (i.e., the first graph structure data), (for example, the attribute att can be the building height); is the attribute att of the building in the graph structure data output by the decoder (i.e., the second graph structure data); is the L2 norm; in this paper, the L2 loss is used to measure the attributes of each building in the graph structure recovered by the decoder and the attributes of each building in the real graph structure similarity degree of.
[0066] During the training process, in this embodiment, the aggregated vector output by the graph dynamic attention encoder is sampled, and the sampling result is input into the graph dynamic attention decoder, and then the graph dynamic attention encoder-decoder is trained. After the training is completed, a trained graph dynamic attention encoder-decoder is obtained, and the graph dynamic attention decoder is used for the inference (generation) process of step S3.
[0067] S4: Convert the graph structure data into a two-dimensional vector layout of sub-regions, and then splice all the two-dimensional vector layouts to obtain a scene plane layout.
[0068] In step S4 of the present invention, the decoder output result (graph structure data) of step S3 is converted into a two-dimensional vector layout of the sub-regions, that is, the reverse process of step S3. Specifically: each node in the graph structure data obtained by decoding is restored to the outline of a single building, and the node features are restored to the features of the outline of a single building. For example, one of the node features obtained by decoding is restored to the corresponding type in the single building type, and then the two-dimensional vector layouts of all sub-regions are merged to form the scene plane layout of the entire scene. This result is the scene plane layout generated by the method provided in this embodiment.
[0069] S5: Construct a scene information table based on the scene plane layout, and obtain a 3D model library; traverse the floor plan outline of each building in the scene information table, and select multiple three-dimensional models that match its outline from the 3D model library to form a candidate set for the building.
[0070] Based on the scene plane layout generated in step S4, this embodiment maintains a scene information table and a 3D model library. The scene information table maintained in this embodiment consists of information such as the ID identifier of the building monomer, the ID of the sub-region of the plot it belongs to, the type of the building monomer, the outline of the building monomer floor plan, the height of the building monomer, and the area size. In addition, this embodiment constitutes a 3D model library by collecting three-dimensional building monomer models (i.e., 3D models) with different shapes and materials from the Internet. The length, width, and height information of these three-dimensional building monomer models can all be edited. In order to more comprehensively extract the information and features of the three-dimensional building monomer models, this embodiment takes snapshots of each three-dimensional building monomer model in the 3D model library from different angles to obtain its key information. Specifically, as shown in Figure 5 the present invention takes snapshots of the three-dimensional building monomer model from the front view at an inclination of degrees (where takes 45, 90, 135, 180, 225, 270, 325, 360) for a total of eight perspectives, and obtains the top view outline of each three-dimensional building monomer model. These snapshots and top view outlines are part of the 3D model library.
[0071] For the problem of matching the plane outline of the building monomer, this embodiment selects the building monomers in the 3D model library with a relatively high degree of matching with the floor plan outline of the building in the scene information table as candidates, and all candidates form a candidate set. During the matching process, this embodiment solves the editable parameters of the three-dimensional building monomer model in the 3D model library (i.e., the length q1 and width q2 of the minimum bounding box of the top view plane outline) to adjust and obtain the parameters corresponding to the maximum matching degree. Specifically, the outline matching problem is transformed into an optimization problem, which can be formally expressed as: given a polygon outline with fixed parameters , that is, the floor plan outline of the building in the scene information table; there is another variable polygon outline , that is, the top view outline of the three-dimensional building monomer model in the 3D model library, and the length and width of the minimum outer bounding box of this outline are two adjustable parameters, and the cost function is defined as the maximum intersection-over-union ratio between the two polygon outlines. The specific calculation formula is: its
[0072] ;
[0073] where is the degree of matching between the floor plan outline of the building and the top view outline of the three-dimensional model and width in the case of the length of the minimum outer bounding box of the top view outline of the three-dimensional model ; The floor plan outline of the building and the top - down outline of the 3D model of the intersection - over - union ratio; The floor plan outline of the building and the top - down outline of the 3D model of the area of the intersecting part; Denote the area of the floor plan outline of the building , Denote the length of the minimum outer bounding box of the top - down outline of the 3D model and the width in the case of the top - down outline of the 3D model of the area.
[0074] To solve this optimization problem, in this embodiment, the Powell algorithm is used. By iterative search, the parameters that minimize the cost function are found, that is, the length of the minimum outer bounding box of the 3D building single - body model under the best contour matching and the width . The intersection - over - union ratio between the top - down view outline of the 3D building single - body model after the final iteration and the floor plan outline of the building in the scene information table is , which is used as the matching degree score. In the present invention, the 3D building single - body models with a matching degree score greater than the set threshold (in a specific embodiment of the present invention, the threshold is set to 0.8) are all used as candidates, that is, as elements in the candidate set. For each 3D building single - body model in the candidate set, the present invention sets another adjustable parameter (height parameter) to the value of the "height" attribute corresponding in the scene information table. After this step, this embodiment obtains the candidate set of buildings in the scene plane layout, so the search space for the 3D building single - body models corresponding to the buildings in the scene plane layout is compressed. Therefore, step S6 performs the matching of the asset materials of the 3D building single - body models in the pruned search space.
[0075] S6: Extract the features of each 3D model in the candidate set, perform optimized matching of the 3D models based on the positional relationship and material similarity of the buildings in the scene plane layout, and then replace the buildings in the scene plane layout with the optimized - matched 3D models to finally obtain a 3D scene.
[0076] The matching of the asset materials of the 3D building single - body models is based on extracting the features of the asset materials to achieve efficient feature matching. In this embodiment, the zero - shot learning ability of the CLIP large model is utilized to extract the material features of the 3D building single - body models, where the features are the materials of the 3D building single - body models to achieve the matching between asset materials. Specifically, as shown in the appendix Figure 6As shown, after the encoder processes the images of each perspective of the same 3D building monomer model stored in the 3D model library, they are converted into feature vectors. In the present invention, the feature vectors of the images of each perspective are concatenated to obtain the encoding vector of the 3D building monomer model, that is, the material feature of the 3D building monomer model is obtained. During the process of asset material matching, for the encoding vector of the 3D building monomer model of another building input, in this embodiment, the cosine similarity between the two vectors is calculated.
[0077] Further, compared with the plane contour matching in the previous steps which is for a single building monomer, the asset material matching in this step is carried out in the entire scene layout. Specifically, for each building in the scene layout there is a candidate set , and the goal of this method is to select a 3D building monomer model from each candidate set to form the entire scene, so that the material features of the buildings in the same sub-region are as similar as possible, while the material features of the buildings between different sub-regions are as different as possible. This problem can be formally expressed as a binary linear programming problem: for a set of candidate sets (where , is the total number of buildings in the scene plane), and the corresponding set of material feature vectors encoded by the CLIP model (where , is the set of material features of the 3D building monomer models in the candidate set , is the material feature of the jth 3D building monomer model in ), find a subset of each candidate set (also a set) in the set of candidate sets such that under the constraint conditions, the cost function is minimized.
[0078] ;
[0079] ;
[0080] ;
[0081] ;
[0082] Among them, is the kth building in the scene plane; is the number of 3D models in the candidate set of the kth building; represents whether the ith 3D model in the candidate set of the kth building is selected; Indicates the material characteristics of the i-th 3D model in the candidate set of the k-th building; Used to indicate a building and the building Whether they are in the same sub-region. When the value is -1, it means they are in the same sub-region. When the value is 1, it means they are not in the same sub-region.
[0083] Among them, the meaning of the first line of the constraint condition is: If two buildings are located in the same plot sub-region, make the materials of their 3D building monomer models as similar as possible, that is, when their cosine similarity is the largest, the value of the cost function is the smallest; If they are located in different sub-regions, make the materials of their 3D building monomer models as dissimilar as possible according to the distance, that is, when their cosine similarity is the smallest, the value of the cost function is the largest. The meaning of the second line of the constraint condition is: For the 3D building monomer models in the candidate set There are two situations: selected or not selected; The meaning of the third line of the constraint condition is: Two 3D building monomer models cannot be assigned to the same building monomer at the same time.
[0084] In this embodiment, the binary linear programming problem is solved with reference to the method of Bouzas et al. [BOUZAS V, LEDOUX H, NAN L. Structure-aware Building Mesh Polygonization [J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2020, 167: 432-442.].
[0085] In the second aspect, corresponding to the above-provided 3D scene controllable generation method, the present invention also correspondingly provides a 3D scene controllable generation system for autonomous driving simulation. The structure of the 3D scene controllable generation system for autonomous driving simulation proposed by the present invention is as Figure 1 shown, including:
[0086] An encoding vector and attribute feature acquisition module M100, which is used to acquire the road network topology map to be used for 3D scene generation, and construct a geometric minimum loop set based on the road network topology map;
[0087] Divide the scene plane corresponding to the road network topology map with the geometric minimum loop as the boundary to obtain multiple sub-regions, generate binary images corresponding to each sub-region; Acquire a trained convolutional encoder and use it to extract the encoding vector of the binary image; Then assign attribute features to each sub-region;
[0088] A graph structure data acquisition module M200, which is used to acquire a trained graph dynamic attention decoder, and use it to decode the attribute features, encoding vectors, and hidden layer vectors sampled from the standard normal distribution to obtain the graph structure data of each sub-region;
[0089] A scene plane layout acquisition module M300, which is used to convert the graph structure data into a two-dimensional layout, and splice all the two-dimensional layouts to obtain the scene plane layout;
[0090] A candidate set acquisition module M400, which is used to construct a scene information table based on the scene plane layout and obtain a 3D model library; traverse the floor plan contours of each building in the scene information table, and screen out 3D models with a contour matching degree higher than the threshold from the 3D model library to construct a candidate set;
[0091] A three-dimensional scene generation module M500, which is used to perform optimized matching based on the positional relationship and feature similarity of the buildings in the scene plane layout, and finally execute a 3D model replacement operation to generate the final three-dimensional scene.
[0092] Furthermore, the three-dimensional scene controllable generation system further includes a scene plane layout data acquisition and processing module, a roadside element addition and scene export module, and a generated scene quality evaluation module.
[0093] The scene plane layout data acquisition and processing module is used to collect and process the scene plane layout data in the real world to form a data set for controllable generation training of the scene plane layout.
[0094] In addition to adding three-dimensional building monomer models to the scene floor plan, the roadside element addition and scene export module uses a procedural method in this embodiment to equidistantly add scene asset elements with regular distributions such as trees and fences on the roadside to increase the scene richness for the perception of autonomous driving vehicles, and package the generated entire three-dimensional scene and export it in the FBX format for use by autonomous driving simulation software such as CARLA.
[0095] In addition to qualitatively analyzing its effect, the generated scene quality evaluation module calculates its generation quality indicators (including visual quality, generation diversity, key attribute distribution) for the generated three-dimensional scene in this embodiment to verify the rationality of the generated three-dimensional scene and the effectiveness of the three-dimensional scene generation method, that is, to verify that the three-dimensional scene generation method proposed in this embodiment is capable of generating three-dimensional scenes with high visual quality and strong diversity.
[0096] The evaluation method of the generated scene quality evaluation module includes the following steps:
[0097] S701: In the CARLA autonomous driving simulation software, extract the roadside perspective image sequence of the generated three-dimensional scene, and calculate its visual quality using the FID (Frechet Inception Distance) metric. Specifically, the calculation method of this metric is as follows. The smaller the FID value, the more similar the distribution of the generated data is to the real data, and the higher the visual quality.
[0098] ;
[0099] wherein, is the similarity degree between the generated image sequence and the real image sequence; X and Y respectively represent the feature distributions of the generation result (derived from the extracted generated three-dimensional scene roadside perspective image) and the real result (derived from the roadside perspective image in the real world) extracted by the Inception V3 model; and represent the mean values of the corresponding distributions; and represent the mean covariance of the corresponding distribution.
[0100] S702: Based on the two-dimensional projection of the buildings in the scene, calculate the diversity of its generation results. Specifically, the present invention calculates the mean intersection over union (mIoU) between the generation result and the real result to evaluate its diversity (denoted as DIV). The formula is specifically:
[0101] ;
[0102] wherein, is all the building floor plans (real data) in the collected data set; is the two-dimensional planar projection of all the buildings in the generated scene (generation result). The larger this value is, the stronger the diversity of the generation result is.
[0103] S703: To further measure the difference in distribution between the generation result and the real data, the present invention observes two key attributes, namely building density and building height, and measures the earth mover's distance Bden(WD) and Bhgt(WD) between the buildings in the generated three-dimensional scene (generation result) and the buildings in the collected data set (real result) in terms of the distribution of these two attributes.
[0104] S704: The present invention uses the Validity index to evaluate whether the generated result meets the constraint limitations of the input conditions. Specifically, the calculation method is the number of buildings exceeding the layout boundary limit in each generated sample on average.
[0105] First, this embodiment extracts the roadside perspective image sequence of the generated three-dimensional scene and calculates its visual quality using the FID (Frechet Inception Distance) index. The scene quality evaluation results generated by this embodiment are shown in Table 1 (where ↓ indicates that the lower the value, the better, and ↑ indicates that the higher the value, the better):
[0106] Table 1
[0107]
[0108] Table 1 shows the quantitative comparison between this embodiment and the prior art in terms of visual quality, diversity of generation results, etc. It can be seen from Table 1 that the visual quality index of the three-dimensional scene generated in this embodiment has decreased by 64.61% numerically compared with the prior art, indicating that the visual quality of the generated scene has been greatly improved compared with the prior art. The diversity of the generated scene has increased by 45.39% compared with the prior art. Attached Figure 7 shows the qualitative comparison between the two-dimensional plane layout of the three-dimensional scene generated in this embodiment and the two-dimensional plane layout generated by the prior art. From the attached Figure 7 it can be seen that the buildings in the two-dimensional plane layout generated in this embodiment have clearer boundaries. Attached Figure 8 is the qualitative comparison between the image rendered by the three-dimensional scene generated in this embodiment in the CARLA simulation software and the image generated by the prior art. From the attached Figure 8 it can be seen that the results generated in this embodiment have high visual quality, strong diversity, and there is no problem of inconsistent perspectives.
[0109] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A three-dimensional scene controllable generation method for autonomous driving simulation, characterized in that Including: 1) Obtain a road network topology graph to be used for 3D scene generation, and construct a set of geometric minimum loops based on the road network topology graph; 2) Divide the scene plane corresponding to the road network topology graph with the geometric minimum loops as boundaries, obtain multiple sub-regions, assign attribute features to each sub-region, and then generate binary images corresponding to each sub-region; Obtain a trained convolutional encoder and use it to extract the encoded vectors of the binary images; 3) Obtain a trained graph dynamic attention decoder, and use it to decode the attribute features, encoded vectors, and hidden layer vectors sampled from the standard normal distribution to obtain the graph structure data of each sub-region; 4) Convert the graph structure data into a 2D layout, and splice all the 2D layouts to obtain the scene plane layout; 5) Construct a scene information table based on the scene plane layout, and obtain a 3D model library; Traverse the floor plan contours of each building in the scene information table, and screen out 3D models with a contour matching degree higher than the threshold from the 3D model library to construct a candidate set; 6) Perform optimization matching based on the positional relationship and feature similarity of the buildings in the scene plane layout, and finally perform a 3D model replacement operation to generate the final 3D scene.
2. The three-dimensional scene controllable generation method for autonomous driving simulation according to claim 1, wherein In step 2), the trained convolutional encoder is obtained by training a convolutional encoder-decoder using a self-constructed training set. The convolutional encoder-decoder includes a convolutional encoder and a convolutional decoder. The convolutional encoder consists of 4 convolutional layers, and the convolutional decoder is mirror-symmetrical to the convolutional encoder.
3. The three-dimensional scene controllable generation method for autonomous driving simulation according to claim 1, wherein In step 2), the construction of the training set includes: Obtain multiple road network topology graphs in the real scene and the corresponding scene plane layout vector graphs. Divide the scene plane corresponding to each scene plane layout vector graph into multiple sub-regions according to the method in step 1), and adjust the scaling scale of the scene plane layout vector graph of the sub-region so that each building in the sub-region is aligned one by one with each building in the public dataset containing building height information in this sub-region; If there is a building with a missing height value in the scene plane layout vector graph of the sub-region, find the building in the public dataset, and find the heights of all buildings within 300 meters of the radius of this building in the public dataset. Calculate the average height of all buildings within 300 meters of the radius and use this average height as the height of the building with the missing height value, update the scene plane layout vector graph of the sub-region. Finally, all the sub-regions of the scene plane and the updated scene plane layout vector graphs of the sub-regions form a dataset; Then generate binary images corresponding to each sub-region in the dataset, and unify the sizes of the obtained binary images. All the unified binary images form the training set for training the convolutional encoder-decoder; Among them, the average height and density of buildings in each sub-region of the dataset are also calculated based on the public dataset, and then attribute features are assigned to each sub-region based on the calculation results; the attribute features include the density of the building group in the sub-region and the height of the building group. The density of the building group is divided into high density and low density. When the density of the buildings in the sub-region is greater than or equal to 0.1, the density of the building group is high density, otherwise it is low density; the height of the building group is divided into high-rise and low-rise. When the average height of the building group is greater than or equal to 20m, it is high-rise, otherwise it is low-rise.
4. The three-dimensional scene controllable generation method for autonomous driving simulation according to claim 3, characterized in that In step 3), the obtaining of the trained graph dynamic attention decoder includes: Introduce a learnable dynamic attention mechanism into the graph neural network to form a graph dynamic attention encoder, and then construct a graph dynamic attention decoder that is mirror-image to the graph dynamic attention encoder to form a graph dynamic attention encoder-decoder; Normalize the scene plane layout vector graph of each sub-region in the dataset to obtain the graph structure data corresponding to each sub-region; Construct a virtual node connected to all nodes in the graph structure data to obtain the first graph structure data, input the first graph structure data into the graph dynamic attention encoder to obtain an aggregated vector; obtain the encoding vector and attribute features of the sub-region in the dataset, input the encoding vector and attribute features of the sub-region into the graph dynamic attention decoder, and at the same time sample the aggregated vector to obtain the hidden layer vector of the graph dynamic attention decoder, and input the hidden layer vector into the graph dynamic attention decoder. Finally, the graph dynamic attention decoder outputs the second graph structure data; Use the conditional variational learning method to train the graph dynamic attention encoder-decoder to obtain the trained graph dynamic attention decoder.
5. The three-dimensional scene controllable generation method for autonomous driving simulation according to claim 4, wherein The inputting the first graph structure data into the graph dynamic attention encoder to obtain an aggregated vector includes: Each node in the first figure structure data has node characteristics as , , where indicates whether there is actually a building unit at node ; indicates the type of the building unit at node ; indicates the abscissa of the center position of the building unit at node ; indicates the ordinate of the center position of the building unit at node ; indicates the height of the building unit at node ; indicates the length of the minimum rotated bounding box of the floor plan outline of the building unit at node ; indicates the width of the minimum rotated bounding box of the floor plan outline of the building unit at node ; indicates the proportion of the floor plan outline of the building unit at node in the minimum rotated bounding box; First, the node 's node features and the node features of the neighbor nodes are concatenated. Then, a learnable weight matrix is used to perform a linear transformation on the concatenated result. The LeakyReLU activation function is applied to the result of the linear transformation, and then it is multiplied by the attention parameter vector to obtain the correlation coefficient ; the specific expression is: ; Then, the relevance coefficient is weighted and summed with the node features of the neighbor nodes, and after passing through a non-linear activation function, the feature vector of each node is obtained; the specific expression is: ; ; Among them, represents the weight coefficient between the node and its neighbor nodes; is the set of neighbor nodes of the node; is the feature vector of the node calculated by the -th layer attention layer; is the non-linear activation function; is the number of attention groups; is the -th group attention mechanism of the -th layer attention layer, which calculates the weight coefficient between the node and its neighbor nodes ; is the input linear transformation matrix of the -th group attention mechanism of the -th layer attention layer; Integrate the feature vectors of each node after passing through the -layer attention layer into a feature matrix , and aggregate the feature matrices of all attention layers into an aggregated vector , .
6. The three-dimensional scene controllable generation method for autonomous driving simulation according to claim 5, wherein Loss function when training the graph dynamic attention decoder using conditional variational learning is as follows: ; where, is the loss function; is a hyperparameter; is the expectation of the posterior distribution of the hidden layer vector under the conditions of the encoding vector , the attribute feature and the first graph structure data ; represents the posterior distribution of the hidden layer vector under the conditions of the encoding vector , the attribute feature and the first graph structure data ; is the distribution of the second graph structure data output by the graph dynamic attention decoder under the conditions of the encoding vector , the attribute feature and the hidden layer vector ; is the prior distribution of the hidden layer vector ; represents the KL divergence distance between the two distributions of ; is a hyperparameter; represents the attribute att of the building in the first graph structure data input to the graph dynamic attention encoder; represents the attribute att of the building in the second graph structure data output by the graph dynamic attention decoder; is the L2 norm.
7. The three-dimensional scene controllable generation method for autonomous driving simulation according to claim 1, wherein In step 5), the scene information table includes the ID identification of the building monomer in the sub-region, the sub-region ID, the type of the building monomer, the floor plan outline of the building monomer, the height and area of the building monomer; The 3D model library includes 3D models with different shapes and materials, as well as eight-view pictures taken from the front view tilted degrees for each 3D model and the top view contour of each 3D model.
8. The three-dimensional scene controllable generation method for autonomous driving simulation according to claim 7, characterized in that In step 5), the screening of three-dimensional models with a contour matching degree higher than the threshold from the 3D model library to construct a candidate set includes: First, traverse the floor plan outlines of each building in the scene information table and the top-down outlines of the three-dimensional models in the 3D model library, and calculate the matching degree between the floor plan outline of the building and the top-down outline of the three-dimensional model. The calculation formula of the matching degree is: ; Among them, is the length of the minimum outer bounding box and width of the top-down contour of the 3D model, indicating the matching degree between the floor plan contour of the building and the top-down contour of the 3D model; is the intersection over union between the floor plan contour of the building and the top-down contour of the 3D model; is the area of the intersection part between the floor plan contour of the building and the top-down contour of the 3D model, where the length and width of the minimum outer bounding box of the top-down contour of the 3D model are given; Taking the maximum intersection over union as the cost function, using the Powell method and minimizing the cost function through iterative search to obtain the top-down contour of the 3D model at the optimal matching degree The length and width of the minimum outer bounding box; finally, calculate the intersection over union between the top-down contour of the 3D model at the optimal matching degree and the floor plan contour of the building, and take the 3D models with an intersection over union greater than the set threshold as candidates in the candidate set of the building.
9. The three-dimensional scene controllable generation method for autonomous driving simulation according to claim 8, wherein In step 6), the optimization matching based on the positional relationship and feature similarity of the buildings in the scene plane layout includes: Use the CLIP image encoder to encode the pictures of each perspective of the three-dimensional models in the 3D model library to obtain feature vectors, and then splice the feature vectors of each perspective to obtain the encoding vector of the three-dimensional model, that is, obtain the material feature of the three-dimensional model; Obtain the material features of all 3D models in the 3D model library, and then extract the material features of each 3D model in the candidate set to form a material feature set corresponding to the corresponding candidate set; aiming at the similarity of the material features of the buildings in the same sub-region of the scene layout and the difference of the material features of the buildings between different sub-regions, construct a cost function Perform optimization matching; the cost function The formula of is: ; ; ; ; Among them, is the k-th building in the scene plane; is the number of 3D models in the candidate set of the k-th building; indicates whether the i-th 3D model in the candidate set of the k-th building is selected; represents the material feature of the i-th 3D model in the candidate set of the k-th building; represents the building and whether they are in the same sub-region. being -1 indicates they are in the same sub-region; being 1 indicates they are not in the same sub-region.
10. A three-dimensional scene controllable generation system for autonomous driving simulation that implements the method described in claim 1, characterized in that, Includes: An encoding vector and attribute feature acquisition module, which is used to acquire the road network topology graph to be used for three-dimensional scene generation, and construct a geometric minimum cycle set based on the road network topology graph; Divide the scene plane corresponding to the road network topology map with the geometric minimum ring as the boundary, obtain multiple sub-regions, assign attribute features to each sub-region, and then generate binary images corresponding to each sub-region; obtain the trained convolutional encoder and use it to extract the encoded vector of the binary image; A graph structure data acquisition module, which is used to obtain the trained graph dynamic attention decoder, and use it to decode the attribute features, encoded vectors, and hidden layer vectors sampled from the standard normal distribution to obtain the graph structure data of each sub-region; A scene plane layout acquisition module, which is used to convert the graph structure data into a two-dimensional layout, and splice all the two-dimensional layouts to obtain the scene plane layout; A candidate set acquisition module, which is used to construct a scene information table based on the scene plane layout and obtain a 3D model library; traverse the floor plan outlines of each building in the scene information table, and screen out 3D models with a contour matching degree higher than the threshold from the 3D model library to construct a candidate set; A three-dimensional scene generation module, which is used to perform optimization matching based on the positional relationship and feature similarity of the buildings in the scene plane layout, and finally perform a 3D model replacement operation to generate the final three-dimensional scene.
Citation Information
Patent Citations
Image generation method of simulation scene, electronic equipment and storage medium
CN110998663A
Interactive simulation test system of automatic driving algorithm
CN112417756A