Street building group automatic layout and iterative optimization method based on TFM architecture
Through the automatic layout and iterative optimization method of neighborhood building complexes based on Transformer architecture, the problem of building complex layout design under multi-constraint conditions is solved, efficient and flexible adaptive learning and dynamic optimization are achieved, and the accuracy of building complex layout design and multi-objective optimization capabilities are improved.
Patent Information
- Application Number
- CN202510354098.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing building complex layout design methods are difficult to achieve efficient and flexible adaptive learning and dynamic optimization under multi-constraint conditions, especially when considering multi-dimensional factors such as spatial constraints, functional constraints and environmental performance, there is a lack of effective automated design methods.
The automatic layout and iterative optimization method of neighborhood building complexes based on Transformer architecture is adopted, and high-precision point cloud data is obtained through LiDAR scanning, and the spatial topological relationship matrix of building units is established. Combined with the self-attention mechanism and the multi-head graph attention module, the relationship between nodes is dynamically perceived, and layout schemes that meet multi-constraint conditions are generated through adaptive learning.
It realizes efficient automatic generation and optimization of building complex layout plans, shortens the design cycle, improves the accuracy and scientificity of the generation plan, and can meet the requirements of multi-dimensional planning indicators at the same time, and improves the feasibility and practicality of the design plan.
Smart Images

Figure CN120296840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of urban planning, and in particular to an automatic layout and iterative optimization method for neighborhood building groups based on the TFM (Transformer) architecture. Background Art
[0002] With the acceleration of the urbanization process, the planning and design of neighborhood building groups face increasingly high complexity and diverse requirements. Traditional building group layout design methods mainly rely on manual experience and rule-driven, and have problems such as low efficiency, poor flexibility, and difficulty in meeting the multi-objective optimization requirements. In recent years, with the rapid development of artificial intelligence technology, automated design methods based on deep learning have gradually become a research hotspot. However, existing automated design methods are mostly limited to single-objective optimization (such as maximizing the floor area ratio), and it is difficult to comprehensively consider multi-dimensional factors such as spatial constraints, functional constraints, and environmental performance, and lack the ability to dynamically model the spatial topological relationship of building groups.
[0003] As a powerful sequence modeling tool, the Transformer (TFM) architecture has achieved remarkable results in the fields of natural language processing and computer vision. Its self-attention mechanism can effectively capture long-range dependence relationships, providing a new idea for modeling complex spatial relationships in building group layout design. However, existing research has not fully explored how to apply the Transformer architecture to the automatic layout and iterative optimization of neighborhood building groups, especially there are still large gaps in adaptive learning and dynamic optimization under multi-constraint conditions. Therefore, there is an urgent need for an automatic layout and iterative optimization method for neighborhood building groups based on the Transformer architecture, which can efficiently process high-precision point cloud data, dynamically model the spatial topological relationship of building groups, and realize the automatic generation and optimization of layout schemes under multi-constraint conditions to meet the high efficiency, flexibility and sustainability requirements of modern urban planning and design. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to provide an automatic layout and iterative optimization method for neighborhood building groups based on the TFM architecture to realize the generation and optimization of building group spatial layout schemes under multi-constraint conditions.
[0005] Technical Solution Adopted by the Present Invention: The automatic layout and iterative optimization method for neighborhood building groups based on the TFM architecture includes the following steps:
[0006] Step 1: Prepare the data chassis and the case library
[0007] Perform three-dimensional scanning on the target neighborhood through a LiDAR (Light Detection and Ranging) to obtain high-precision point cloud data containing terrain elevation, road boundary lines, and existing building outlines, and synchronously import the planning condition constraint data. The constraint data includes plot ratio, building setback distance, and height limit area range; extract multi-city historical built-up area data from the public urban planning database, screen valid cases containing building coordinates, building scale, and functional type information, and establish a standardized neighborhood building complex database containing more than 100 cases; encode the spatial topological relationship matrix and economic and technical index statistical table of each case data containing building units; perform feature processing on the original data through a hierarchical encoder, and convert the basic data of the case library into a structured learnable graph attention matrix that can be input through a Transformer.
[0008] Step 2: Learn the association relationship of the neighborhood building complex
[0009] Through the self-attention mechanism of the Transformer, use the spatial position and its attributes of the neighborhood building complex as input embedding vectors, and use a multi-scale convolutional kernel group of a convolutional neural network as a pre-module to extract features from the input embedding vectors, and model the spatial association relationship between the building units in the case library.
[0010] Map the spatial attribute vectors of each case and convert them into a graph structure based on the spatial topological relationship between building units; based on the multi-head graph attention Transformer module, introduce edge weight adjustment of the graph structure to adjust the attention distribution, so that the model can dynamically perceive the relationship between nodes and perform adaptive attention learning; through stacking N layers of Transformer encoders, output an implicit association matrix A∈R{M×M} between building units to form a relationship graph network, where M is the total number of building units.
[0011] Step 3: Set design constraint conditions
[0012] Divide the design constraints into two categories: spatial constraints and functional constraints, and inject constraint conditions at different levels of the Transformer encoder; among them, the spatial constraints include: maximum building height, plot ratio threshold, building setback distance, minimum greening rate; the functional constraints include: land use compatibility table, radiation range of public service facilities, traffic noise isolation requirements; use a mixed method of numerical coding and symbolic coding to generate a constraint feature matrix, where the threshold-type constraints are converted into normalized scalar values; the logical-type constraints are converted into binary mask vectors.
[0013] Step 4: Initially generate the layout plan of the building complex
[0014] Input the chassis data and design constraint data of the plot to be designed into the neighborhood building complex spatial layout generation model. Convert the chassis data of the plot to be designed into a plot feature vector through the plot shape encoder, and convert the design constraint condition data into a global condition vector through the global condition encoder. Concatenate the plot feature vector and the global condition vector to obtain an initial condition vector. Input the initial condition vector into the autoregressive graph generation model based on the transformer to generate a preliminary building complex layout plan;
[0015] Step Five: Plan Iteration and Optimization
[0016] Encode the building complex layout graph obtained in Step Four into a high-dimensional node matrix. The rows in the matrix include the function, height, base area, and normalized coordinates of the buildings. Combine with the global condition vector to form a complete representation of the generated plan. Calculate the plot ratio, building density, and the highest building height of the current generated plan. Use the fitness function to quantitatively evaluate the plan. If the fitness is equal to or greater than the threshold, save the building complex layout plan graph network generated currently. If the fitness is less than the threshold, use the optimization algorithm to perform plan optimization iteration based on the high-dimensional node matrix of the building complex layout, and repeat the optimization loop of Step Four and Step Five until the plan indicators converge near the input design constraint data or reach the maximum number of iterations, and output the generated building complex layout plan graph network;
[0017] Step Six: Layout Plan Verification and Output
[0018] Save the building complex layout plan graph network output in Step Five. According to the attributes of the nodes and the weights of the edges in the graph network, convert the graph network into a three-dimensional digital model of the building complex layout plan. After outputting it to the simulation device for feasibility verification of the plan, perform refined adjustment and human-computer interaction display.
[0019] Furthermore, for the basic data conversion method of the structured learnable graph attention matrix in Step One, perform feature processing on the original data through a hierarchical encoder, perform voxelization downsampling on the point cloud data, and extract multi-scale terrain feature vectors through a 3D convolutional neural network. The formula for voxelization sampling is as follows:
[0020]
[0021] where V (x,y) is the coordinate of the voxel center point, p i is the point cloud data point falling within this voxel, and N is the number of point cloud data points within this voxel;
[0022] Convert the building layout data in the case library into a serialized representation: construct a feature tuple with the central coordinates (x, y), floor area (S), height (H), and functional type code (F) of each building unit, and generate a sequence after sorting by spatial proximity; convert the adjacency constraints into a learnable graph attention matrix, and use a graph neural network to pre-generate the relationship embedding vector between buildings.
[0023] Further, in step 2, the spatial position and its attributes of the neighborhood building group are used as input embedding vectors. The four-dimensional attribute vector includes the spatial coordinates (x, y), area S, height H, and functional type code F of the building unit, and is converted into a d-dimensional initial embedding vector E through a linear mapping layer. i ∈R d ; Extract spatial features from the initial embedding vector through a multi-scale convolution kernel group: the first convolutional layer uses a 3×3 kernel to capture the neighborhood building density feature; the second convolutional layer uses a 5×5 kernel to extract the medium-range functional group feature; the third convolutional layer uses a 7×7 kernel to capture the global spatial distribution pattern, and outputs the enhanced feature vector E'. i ∈R d .
[0024] Further, in step 2, when converting the spatial topological relationship between building units into a graph structure, the case data is converted into a graph structure G=(V, E), where the node V corresponds to the building unit, the node feature is E'i, and the weight of the edge E is calculated by the exponential decay function of the building spacing as follows:
[0025]
[0026] where σ is the Gaussian kernel radius parameter, and the Gaussian kernel radius parameter σ is dynamically adjusted according to the building function type: σ = 50 meters for commercial building groups; σ = 30 meters for residential building groups; σ = 100 meters for industrial building groups.
[0027] Further, in step 3, constraint conditions are injected at different levels of the Transformer encoder. Among them, in the feed-forward layer injection: the spatial constraint feature vector and the building unit feature vector are concatenated and then input into the fully connected layer; in the attention layer injection: the functional constraint is encoded as an attention mask matrix, and the calculation formula is:
[0028]
[0029] where M constraint is a 0-1 mask matrix generated by the functional compatibility rule.
[0030] Further, the plot shape encoder module in step 4 refers to using a convolutional neural network and a multi-layer perceptron to extract the geometric features of the vector data of the plot to be designed, and converting them into a shape feature vector h. shape; The global condition encoder refers to converting building density, plot ratio, and maximum building height into a condition vector through a multi-layer perceptron, encoding land use functions into a category vector through an embedding layer, and concatenating the condition vector and the category vector to obtain the global condition vector h cond .
[0031] Furthermore, the step four Transformer-based autoregressive graph generation model refers to an intelligent generation model that can generate a building complex layout graph network, consisting of a Transformer decoder, a node generation module, an edge generation module, a coordinate predictor module, and a dynamic termination controller. The initial condition vector and the embedding of the start symbol are input into the Transformer-based autoregressive model graph generator. The node generation module includes a function classifier for predicting building functions and a height and area regressor for predicting numerical values; the edge generation module includes an edge existence classifier for predicting whether there is an edge connection between nodes and an edge weight regressor for predicting the spacing numerical value.
[0032] Furthermore, the step four of generating a preliminary building complex layout plan refers to generating the first building node with attributes including building function, height, and base area based on the initial condition vector and the embedding of the start symbol, using the Transformer decoder through the multi-head attention mechanism and based on the learned building complex layout rules. The position of the building is determined within the boundary of the plot to be designed through the coordinate prediction module in combination with the plot mask, and the remaining constructible capacity is updated according to the difference between the total floor area constraint index and the floor area of all generated nodes, and it is judged whether to continue generating new nodes. The embedding of the generated nodes is concatenated with the initial condition vector and then input into the decoder again, repeating the above node generation process to generate the next node. At the same time, the spatial topological relationship between the new node and all generated nodes is calculated through the edge prediction module. If an edge is predicted to exist between nodes, its weight is recorded; the above process is looped until the dynamic termination condition is met and then the generation of new nodes stops, and the building complex layout graph network is output.
[0033] Furthermore, the formula for the fitness function in step five is:
[0034] Fitness = exp(-α·(L + β·Penalty))
[0035] L = λ1∣BD in -BD g ∣ + λ2∣FAR in -FAR g ∣ + λ3∣H max -H g ∣
[0036] Among them, L refers to the loss of the index proximity, Penalty is the penalty term for exceeding the plot boundary or insufficient spacing, α and β are weight coefficients, BD in is the building density of the generated building complex scheme, BD g is the input building density constraint value, FAR in is the plot ratio of the generated building complex scheme, FAR g is the input plot ratio constraint value, H max is the maximum building height of the generated building complex scheme, H g is the input maximum building height constraint value;
[0037] Among them, the calculation formula of Penalty is:
[0038] Penalty = Penalty boudary + Penalty spacing
[0039]
[0040] Among them, penalty boundary is the penalty term for exceeding the plot boundary, penalty spacing is the building spacing penalty term, x min 、x max 、y min 、y max are the plot boundary coordinates, x i 、y i are the ground center point coordinates of building i, d min is the minimum spacing requirement between buildings.
[0041] Furthermore, the optimization algorithm used in the fifth step refers to using the genetic algorithm to perform perturbations in the high-dimensional space. After flattening the node matrix into a vector, the parent solutions with high fitness are selected through tournament selection, and block crossover and Gaussian mutation are applied to them, respectively exchanging part of the building sub-matrix and randomly perturbing the height and coordinates to generate offspring solutions; subsequently, the mutated coordinates are forced to be constrained within the plot boundary using the projection operator, and the edge matrix is recalculated to ensure the minimum spacing. After each round of iteration, the population is updated based on the elitist retention strategy, and the top 10% of the solutions with high fitness are retained and the inefficient individuals are replaced.
[0042] Furthermore, the simulation and simulation device in the sixth step refers to a digital sand table integrated with algorithms for simulating the physical environment performance of the building complex, algorithms for simulating the flow of people, AR reality enhancement glasses devices, and voice and gesture recognition module devices, which enables urban designers to finely optimize and adjust the layout scheme of the building complex according to the results of the simulation algorithms, and display the three-dimensional model of the scheme, the building functions, height, area indicators, and the technical and economic indicators of the building complex through the AR reality enhancement glasses.
[0043] Beneficial effects: Compared with the prior art, the advantages of the present invention are as follows. Based on the Transformer architecture, the automatic layout and iterative optimization of neighborhood building complexes are realized, and the dynamic association between the spatial topological relationship and planning constraint conditions of the building complex is established. Through the self-attention mechanism and adaptive learning, the automatic generation and optimization of the layout scheme of the building complex that meets multiple constraint conditions are realized, specifically including the following advantages:
[0044] 1. By constructing a prediction model for the spatial topological relationship of the building complex, the present invention establishes a dynamic association relationship between the spatial positions, functional attributes of building units and planning constraint conditions, breaks through the deficiencies in the traditional layout design of building complexes that rely solely on manual experience and rule-driven, and improves the accuracy and scientificity of the generation of layout schemes.
[0045] 2. Through the self-attention mechanism of Transformer and the multi-head graph attention module, the present invention realizes the adaptive learning and iterative optimization of the layout scheme of the building complex, shortens the original design cycle that takes more than several weeks to within several hours, and improves the standardization and standardized output of the scheme.
[0046] 3. Through the comprehensive application of the fitness function and optimization algorithm, the present invention realizes multi-objective optimization, can simultaneously meet the requirements of multi-dimensional planning indicators such as plot ratio, building density, and greening rate, and significantly improves the feasibility and practicality of the design scheme. Description of the Drawings
[0047] Figure 1 is the flow chart of the automatic optimization method of the present invention;
[0048] Figure 2 is the schematic diagram of the modeling of the spatial topological relationship of the building complex and the Transformer architecture of the present invention;
[0049] Figure 3 is the schematic diagram of the generation and optimization process of the layout scheme of the building complex of the present invention. Detailed Embodiments
[0050] The technical solution of the present invention will be further described below with reference to the drawings.
[0051] As Figures 1 - 3 shown, the automatic layout and iterative optimization method of the neighborhood building complex based on the TFM architecture includes the following steps.
[0052] 1. Prepare the data chassis and case library. Use a LiDAR (Light Detection and Ranging) to perform 3D scanning on the target neighborhood to obtain high-precision point cloud data including terrain elevation, road boundary lines, and existing building outlines. Synchronously import the planning condition constraint data, where the constraint data includes floor area ratio, building setback distance, and height limit area range. Extract historical built-up area data from multiple cities in the public urban planning database, screen valid cases containing building coordinates, building scale, and function type information, and establish a standardized neighborhood building complex database containing more than 100 cases. Encode the spatial topological relationship matrix and economic and technical index statistical table of each case data containing building units. Use a hierarchical encoder to perform feature processing on the original data, and convert the basic data of the case library into a structured learnable graph attention matrix that can be input through a Transformer.
[0053] Among them, for the method of converting the basic data of the structured learnable graph attention matrix, use a hierarchical encoder to perform feature processing on the original data, perform voxelization downsampling on the point cloud data, and extract multi-scale terrain feature vectors through a 3D convolutional neural network. The formula for voxelization sampling is as follows:
[0054]
[0055] Among them, V (x,y) is the coordinate of the voxel center point, p i is the point cloud data point falling within this voxel, and N is the number of point cloud data points within this voxel;
[0056] Convert the building layout data in the case library into a serialized representation: form a feature tuple with the center coordinates (x, y), floor area (S), height (H), and function type encoding (F) of each building unit, and generate a sequence after sorting by spatial proximity. Convert the adjacency constraint into a learnable graph attention matrix, and use a graph neural network to pre-generate the relationship embedding vector between buildings.
[0057] 2. Learn the association relationship of the neighborhood building complex. Through the self-attention mechanism of the Transformer, use the spatial position and its attributes of the neighborhood building complex as input embedding vectors, use a multi-scale convolutional kernel group of a convolutional neural network as a pre-module to extract features from the input embedding vectors, and model the spatial association relationship between the building units in the case library;
[0058] Map the spatial attribute vectors of each case and convert them into a graph structure based on the spatial topological relationships between building units. Based on the multi-head graph attention Transformer module, introduce edge weights of the graph structure to adjust the attention distribution, enabling the model to dynamically perceive the relationships between nodes and perform adaptive attention learning. By stacking N layers of Transformer encoders, output an implicit association matrix A ∈ R{M×M} between building units to form a relational graph network, where M is the total number of building units.
[0059] Among them, the spatial positions and their attributes of the neighborhood building complex are used as input embedding vectors, and a four-dimensional attribute vector including the spatial coordinates (x, y), area (S), number of floors (H), and functional type encoding (F) of the building unit is converted into a d-dimensional initial embedding vector E i ∈R d ; perform spatial feature extraction on the initial embedding vector through a multi-scale convolutional kernel group: the first convolutional layer uses a 3×3 kernel to capture the neighborhood building density feature; the second convolutional layer uses a 5×5 kernel to extract the medium-range functional group feature; the third convolutional layer uses a 7×7 kernel to capture the global spatial distribution pattern, and output an enhanced feature vector E' i ∈R d .
[0060] Among them, in the conversion of the spatial topological relationships between building units into a graph structure, the case data is converted into a graph structure G = (V, E), where the node V corresponds to the building unit, the node feature is E'i, and the weight of the edge E is calculated by the exponential decay function of the building spacing as:
[0061]
[0062] Among them, σ is the Gaussian kernel radius parameter, and the Gaussian kernel radius parameter σ is dynamically adjusted according to the building functional type: σ = 50 meters is set for commercial building complexes; σ = 30 meters is set for residential building complexes; σ = 100 meters is set for industrial building complexes.
[0063] Third, set design constraint conditions, divide the design constraints into two categories: spatial constraints and functional constraints, and inject constraint conditions at different levels of the Transformer encoder; among them, the spatial constraints include: maximum building height, plot ratio threshold, building setback distance, minimum greening rate; the functional constraints include: land use compatibility table, radiation range of public service facilities, traffic noise isolation requirements; use a mixed method of numerical encoding and symbolic encoding to generate a constraint feature matrix, where the threshold-type constraints are converted into normalized scalar values; the logical-type constraints are converted into binary mask vectors.
[0064] Among them, constraint conditions are injected at different levels of the Transformer encoder. Among them, for the feed-forward layer injection: the spatial constraint feature vector and the building unit feature vector are concatenated and then input into the fully connected layer; for the attention layer injection: the functional constraint is encoded as an attention mask matrix, and the calculation formula is:
[0065]
[0066] Among them, M constraint is a 0-1 mask matrix generated by the functional compatibility rule.
[0067] IV. Initially generate the building complex layout plan. Input the chassis data and design constraint data of the plot to be designed into the neighborhood building complex spatial layout generation model. The chassis data of the plot to be designed is converted into a plot feature vector through the plot shape encoder, and the design constraint data is converted into a global condition vector through the global condition encoder. The plot feature vector and the global condition vector are concatenated to obtain the initial condition vector; the initial condition vector is input into the autoregressive graph generation model based on Transformer to generate the initial building complex layout plan.
[0068] Among them, the plot shape encoder module refers to using a convolutional neural network and a multi-layer perceptron to extract the geometric features of the vector data of the plot to be designed and convert them into a shape feature vector h shape ; the global condition encoder refers to converting the building density, plot ratio, and maximum building height into condition vectors through a multi-layer perceptron, encoding the land use function into a category vector through an embedding layer, and concatenating the condition vector and the category vector to obtain the global condition vector h cond .
[0069] Among them, the autoregressive graph generation model based on Transformer refers to an intelligent generation model that can generate a building complex layout graph network and is composed of a Transformer decoder, a node generation module, an edge generation module, a coordinate predictor module, and a dynamic termination controller. The initial condition vector and the embedding of the start symbol are input into the autoregressive model graph generator based on Transformer. Among them, the node generation module includes a function classifier for predicting the building function and a height and area regressor for predicting numerical values; the edge generation module includes an edge existence classifier for predicting whether there is an edge connection between nodes and an edge weight regressor for predicting the spacing numerical value.
[0070] Among them, the generation of the preliminary building complex layout plan refers to generating the first building node with attributes including building function, height, and base area based on the embedding of the initial condition vector and the starting symbol, using a Transformer decoder through the multi-head attention mechanism and based on the learned building complex layout rules. The position of the building is determined within the boundary of the plot to be designed through the coordinate prediction module in combination with the plot mask, and the remaining constructible capacity is updated according to the difference between the total floor area constraint index and the floor area of all generated nodes, and it is judged whether to continue generating new nodes. The embedding of the generated nodes is concatenated with the initial condition vector and then input into the decoder again, repeating the above node generation process to generate the next node. At the same time, the spatial topological relationship between the new node and all generated nodes is calculated through the edge prediction module. If there is an edge between the predicted nodes, its weight is recorded; the above process is cycled until new nodes are no longer generated after meeting the dynamic termination condition, and the building complex layout diagram network is output.
[0071] V. Scheme iteration and optimization: Encode the building complex layout diagram network obtained in step four into a high-dimensional node matrix. The rows in the matrix include the function, height, base area, and normalized coordinates of the building, and form a complete representation of the generation scheme in combination with the global condition vector; calculate the plot ratio, building density, and maximum building height of the current generation scheme, and use the fitness function to quantitatively evaluate the scheme. If the fitness is equal to or greater than the threshold, save the current generated building complex layout scheme diagram network; if the fitness is less than the threshold, use the optimization algorithm to perform scheme optimization iteration based on the high-dimensional node matrix of the building complex layout, and repeat the optimization cycle of step four and step five until the scheme index converges near the input design constraint data or reaches the maximum number of iterations, and output the generated building complex layout scheme diagram network.
[0072] Among them, the formula of the fitness function is:
[0073] Fitness = exp(-α·(L + β·Penalty))
[0074] L = λ1∣BD in -BD g ∣ + λ2∣FAR in -FAR g ∣ + λ3∣H max -H g ∣
[0075] Among them, L refers to the loss of index proximity, Penalty is the penalty term for exceeding the plot boundary or insufficient spacing, α and β are weight coefficients, BD in is the building density of the generated building complex scheme, BD g is the input building density constraint value, FAR in is the plot ratio of the generated building complex scheme, FARg is the input floor area ratio constraint value, H max is the maximum building height of the generated building complex plan, H g is the input maximum building height constraint value;
[0076] The calculation formula of Penalty is as follows:
[0077] Penalty = Penalty boundary + Penalty spacing
[0078]
[0079] where penalty boundary is the penalty term for exceeding the plot boundary, penalty spacing is the building spacing penalty term, x min 、x max 、y min 、y max are the plot boundary coordinates, x i 、y i are the ground center point coordinates of building i, d min is the minimum spacing requirement between buildings.
[0080] Among them, the use of the optimization algorithm means using the genetic algorithm to perform perturbations in the high-dimensional space. After flattening the node matrix into a vector, the parent solutions with high fitness are selected through tournament selection. Block crossover and Gaussian mutation are applied to them, and part of the building sub-matrix is exchanged and the height and coordinates are randomly perturbed respectively to generate offspring solutions; Subsequently, the mutated coordinates are forced to be constrained within the plot boundary by using the projection operator, and the edge matrix is recalculated to ensure the minimum spacing. After each iteration, the population is updated based on the elitist retention strategy, and the top 10% of the solutions with high fitness are retained and the inefficient individuals are replaced.
[0081] VI. Layout plan verification and output, save the building complex layout plan graph network output in step five. According to the attributes of the nodes and the weights of the edges in the graph network, convert the graph network into a three-dimensional digital model of the building complex layout plan, output it to the simulation device for feasibility verification of the plan, and then perform fine-tuning and human-computer interaction display.
[0082] Among them, the simulation device refers to a digital sand table integrated with building complex physical environment performance simulation algorithms, pedestrian flow simulation algorithms, AR reality augmented glasses devices and voice and gesture recognition module devices, which enables urban designers to perform fine optimization and adjustment of the building complex layout plan according to the results of the simulation algorithms, and display the three-dimensional model of the plan, building functions, height, area indicators and technical and economic indicators of the building complex through the AR reality augmented glasses.
[0083] Embodiment
[0084] The technical solution of the present invention will be described in detail below taking Nanjing Hexi CBD as an example.
[0085] (1) Prepare the data chassis and the case library. Perform 3D scanning on Nanjing Hexi CBD through a LiDAR (Light Detection and Ranging) to obtain high-precision point cloud data including terrain elevation, road boundary lines, and existing building outlines, and synchronously import the planning condition constraint data. The constraint data includes floor area ratio, building setback distance, and height limit area range; extract multi-city historical built-up area data from the public urban planning database, screen valid cases containing building coordinates, building scale, and function type information, and establish a standardized neighborhood building complex database containing more than 100 cases; encode the spatial topological relationship matrix and economic and technical index statistical table of each case data containing building units; perform feature processing on the original data through a hierarchical encoder, and convert the basic data of the case library into a structured learnable graph attention matrix that can be input through a Transformer.
[0086] (1.1) The method for converting the basic data of the structured learnable graph attention matrix performs feature processing on the original data through a hierarchical encoder, performs voxelization downsampling on the point cloud data, and extracts multi-scale terrain feature vectors through a 3D convolutional neural network; the formula for voxelization sampling is as follows:
[0087]
[0088] where V (x,y) is the coordinate of the voxel center point, p i is the point cloud data point falling within the voxel, and N is the number of point cloud data points within the voxel.
[0089] Convert the building layout data in the case library into a serialized representation: form a feature tuple with the center coordinates (x, y), floor area (S), height (H), and function type code (F) of each building unit, and generate a sequence after sorting by spatial proximity; convert the adjacency constraint into a learnable graph attention matrix, and use a graph neural network to pre-generate the relationship embedding vector between buildings.
[0090] (2) Learn the association relationship of the neighborhood building complex. Through the self-attention mechanism of the Transformer, use the spatial position and its attributes of the Nanjing Hexi CBD neighborhood building complex as the input embedding vector, and use a multi-scale convolutional kernel group of a convolutional neural network as a pre-module to extract features from the input embedding vector, and model the spatial association relationship between the building units in the case library.
[0091] Map the spatial attribute vectors of each case and convert them into a graph structure based on the spatial topological relationships between building units; based on the multi-head graph attention Transformer module, introduce edge weights of the graph structure to adjust the attention distribution, enabling the model to dynamically perceive the relationships between nodes and perform adaptive attention learning; by stacking N layers of Transformer encoders, output the implicit association matrix A ∈ R{M×M} between building units to form a relational graph network, where M is the total number of building units;
[0092] (2.1) Use the spatial location and its attributes of the Nanjing Hexi CBD neighborhood building complex as input embedding vectors. The four-dimensional attribute vector includes the spatial coordinates (x, y), area S, height H, and functional type code F of the building unit, and is converted into a d-dimensional initial embedding vector E through a linear mapping layer i ∈R d ; Extract spatial features from the initial embedding vector through a multi-scale convolution kernel group: The first convolutional layer uses a 3×3 kernel to capture the neighborhood building density feature; the second convolutional layer uses a 5×5 kernel to extract the medium-range functional group feature; the third convolutional layer uses a 7×7 kernel to capture the global spatial distribution pattern, and output the enhanced feature vector E' i ∈R d .
[0093] (2.2) In the conversion of the spatial topological relationships between the building units of the Nanjing Hexi CBD into a graph structure, convert the case data into a graph structure G = (V, E), where the node V corresponds to the building unit, the node feature is E'i, and the weight of the edge E is calculated by the exponential decay function of the building spacing as follows:
[0094]
[0095] where σ is the Gaussian kernel radius parameter, and the Gaussian kernel radius parameter σ is dynamically adjusted according to the building function type: σ = 50 meters for commercial building complexes; σ = 30 meters for residential building complexes; σ = 100 meters for industrial building complexes.
[0096] (3) Set design constraint conditions, divide the design constraints into two categories: spatial constraints and functional constraints, and inject constraint conditions at different levels of the Transformer encoder; among them, the spatial constraints include: maximum building height, plot ratio threshold, building setback distance, minimum greening rate; the functional constraints include: land use compatibility table, radiation range of public service facilities, traffic noise isolation requirements; use a mixed method of numerical coding and symbolic coding to generate a constraint feature matrix, where the threshold-type constraints are converted into normalized scalar values; the logical-type constraints are converted into binary mask vectors.
[0097] Inject constraint conditions at different levels of the Transformer encoder. Among them, for the feed-forward layer injection: concatenate the spatial constraint feature vector and the building unit feature vector and then input them into the fully connected layer; for the attention layer injection: encode the functional constraints as an attention mask matrix, and the calculation formula is:
[0098]
[0099] where M constraint is a 0-1 mask matrix generated by the functional compatibility rule.
[0100] (4) Initially generate the layout plan of the Nanjing Hexi CBD building complex. Input the chassis data and design constraint data of the plot to be designed into the Nanjing Hexi CBD neighborhood building complex spatial layout generation model. Convert the chassis data of the Nanjing Hexi CBD plot into a plot feature vector through the plot shape encoder, and convert the design constraint condition data into a global condition vector through the global condition encoder. Concatenate the plot feature vector and the global condition vector to obtain the initial condition vector; input the initial condition vector into the Transformer-based autoregressive graph generation model to generate the initial layout plan of the Nanjing Hexi CBD building complex.
[0101] (4.1) The Nanjing Hexi CBD plot shape encoder module refers to using a convolutional neural network and a multi-layer perceptron to extract the geometric features of the vector data of the plot to be designed and convert them into a shape feature vector h shape ; the global condition encoder refers to converting the building density, floor area ratio, and maximum building height into a condition vector through a multi-layer perceptron, encoding the land use function into a category vector through an embedding layer, and obtaining the global condition vector h cond .
[0102] (4.2) The Transformer-based autoregressive graph generation model refers to an intelligent generation model that can generate a building complex layout graph network and is composed of a Transformer decoder, a node generation module, an edge generation module, a coordinate predictor module, and a dynamic termination controller. Input the initial condition vector and the embedding of the start symbol into the Transformer-based autoregressive model graph generator. Among them, the node generation module includes a function classifier for predicting building functions and a height and area regressor for predicting numerical values; the edge generation module includes an edge existence classifier for predicting whether there is an edge connection between nodes and an edge weight regressor for predicting the spacing numerical value.
[0103] (4.3) The generation of the preliminary layout plan for the Nanjing Hexi CBD building complex refers to the embedding based on the initial condition vector and the starting symbol. Using the transformer decoder through the multi-head attention mechanism and based on the learned building complex layout rules, the first building node with attributes including building function, height, and floor area is generated. The position of the building is determined within the boundary of the plot to be designed through the coordinate prediction module in combination with the plot mask, and the remaining buildable capacity is updated according to the difference between the total floor area constraint index and the floor areas of all generated nodes, and it is judged whether to continue generating new nodes. The embedding of the generated nodes is concatenated with the initial condition vector and then input into the decoder again. The above node generation process is repeated to generate the next node. At the same time, the spatial topological relationship between the new node and all generated nodes is calculated through the edge prediction module. If there is an edge between the predicted nodes, its weight is recorded; this process is looped until new node generation stops after meeting the dynamic termination condition, and the layout diagram network of the Nanjing Hexi CBD building complex is output.
[0104] (5) Scheme iteration and optimization: Encode the layout diagram network of the Nanjing Hexi CBD building complex obtained in step four into a high-dimensional node matrix. The rows in the matrix include the function, number of floors, floor area, and normalized coordinates of the building, and form a complete representation of the generation scheme in combination with the global condition vector; calculate the plot ratio, building density, and maximum number of floors of the current Nanjing Hexi CBD generation scheme, and use the fitness function to quantitatively evaluate the scheme. If the fitness is equal to or greater than the threshold, save the layout diagram network of the currently generated building complex scheme; if the fitness is less than the threshold, then use the optimization algorithm to perform scheme optimization iteration based on the high-dimensional node matrix of the building complex layout, and repeat the optimization loop of step four and step five until the Nanjing Hexi CBD scheme indicators converge near the input design constraint data or reach the maximum number of iterations, and output the layout diagram network of the generated Nanjing Hexi CBD building complex scheme.
[0105] Among them, the formula of the fitness function is:
[0106] Fitness = exp(-α·(L + β·Penalty))
[0107] L = λ1∣BD in -BD g ∣ + λ2∣FAR in -FAR g ∣ + λ3∣H max -H g ∣
[0108] Among them, L refers to the loss of index proximity, Penalty is the penalty term for exceeding the plot boundary or insufficient spacing, α and β are weight coefficients, BD in is the building density of the generated building complex scheme, BD gis the input building density constraint value, FAR in is the floor area ratio of the generated building cluster scheme, FAR g is the input floor area ratio constraint value, H max is the maximum building height of the generated building cluster scheme, H g is the input maximum building height constraint value;
[0109] Among them, the calculation formula of Penalty is:
[0110] Penalty = Penalty boundary + Penalty spacing
[0111]
[0112] Among them, penalty boundary is the penalty term for exceeding the plot boundary, penalty spacing is the building spacing penalty term, x min 、x max 、y min 、y max are the plot boundary coordinates, x i 、y i are the ground center point coordinates of building i, d min is the minimum spacing requirement between buildings.
[0113] (5.1) The described use of the optimization algorithm means using the genetic algorithm to perform perturbations in the high-dimensional space. After flattening the node matrix into a vector, the parent solutions with high fitness are selected through tournament selection. Block crossover and Gaussian mutation are applied to them, and part of the building sub-matrix is exchanged and the height and coordinates are randomly perturbed respectively to generate offspring solutions; subsequently, the mutated coordinates are forced to be constrained within the plot boundary using the projection operator, and the edge matrix is recalculated to ensure the minimum spacing. After each iteration, the population is updated based on the elitist retention strategy, and the top 10% of the solutions with high fitness are retained and the inefficient individuals are replaced.
[0114] (6) Verification and output of the Nanjing Hexi CBD layout scheme. Save the network of the Nanjing Hexi CBD building cluster layout scheme diagram output in step five. According to the attributes of the nodes and the weights of the edges in the diagram network, convert the diagram network into a three-dimensional digital model of the building cluster layout scheme, output it to the simulation device for feasibility verification of the scheme, and then perform fine-tuning and human-computer interaction display.
[0115] (6.1) The simulation device refers to a digital sand table integrated with algorithms for simulating the physical environment performance of building complexes, algorithms for simulating the flow of people, AR reality enhancement glasses devices, and voice and gesture recognition module devices, which enables urban designers to finely optimize and adjust the layout plan of building complexes according to the results of the simulation algorithms, and display the 3D models of the plan, the building functions, height, area indicators, and the technical and economic indicators of the building complexes through the AR reality enhancement glasses.
[0116] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0117] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed.
Claims
1. An automatic layout and iterative optimization method for neighborhood building complexes based on the TFM architecture, characterized in that, The steps are as follows: Step 1: Prepare the data chassis and the case library Perform 3D scanning on the target neighborhood through a laser scanner to obtain high-precision point cloud data including terrain elevation, road boundary lines, and existing building outlines, and synchronously import the planning condition constraint data. The constraint data includes floor area ratio, building setback distance, and height limit area range; extract multi-city historical built-up area data from the public urban planning database, screen out valid cases containing building coordinates, building scale, and functional type information, and establish a standardized neighborhood building complex database containing more than 100 cases; encode the spatial topological relationship matrix and economic and technical index statistical table of each case data containing building units; perform feature processing on the original data through a hierarchical encoder, and convert the basic data of the case library into a structured learnable graph attention matrix that can be input through a Transformer; Step 2: Learn the association relationship of the neighborhood building complex Through the self-attention mechanism of the Transformer, use the spatial position and its attributes of the neighborhood building complex as input embedding vectors, and use the multi-scale convolutional kernel group of the convolutional neural network as a pre-module to extract features from the input embedding vectors, and model the spatial association relationship between the building units in the case library; Map the spatial attribute vectors of each case and convert them into a graph structure based on the spatial topological relationship between building units; based on the multi-head graph attention Transformer module, introduce the edge weights of the graph structure to adjust the attention distribution, so that the model can dynamically perceive the relationship between nodes and perform adaptive attention learning; Output the implicit association matrix A∈R{M×M} between building units by stacking N layers of Transformer encoders to form a relational graph network, where M is the total number of building units; Step 3: Set the design constraint conditions Divide the design constraints into two categories: spatial constraints and functional constraints, and inject constraint conditions at different levels of the Transformer encoder; among them, the spatial constraints include: maximum building height, floor area ratio threshold, building setback distance, minimum greening rate; the functional constraints include: land use compatibility table, radiation range of public service facilities, traffic noise isolation requirements; use a mixed method of numerical coding and symbolic coding to generate a constraint feature matrix, where the threshold-type constraints are converted into normalized scalar values; the logical-type constraints are converted into binary mask vectors; Step 4: Initially generate the layout plan of the building complex Input the chassis data and design constraint data of the plot to be designed into the neighborhood building complex spatial layout generation model. Convert the chassis data of the plot to be designed into a plot feature vector through the plot shape encoder, and convert the design constraint condition data into a global condition vector through the global condition encoder. Concatenate the plot feature vector and the global condition vector to obtain an initial condition vector; input the initial condition vector into the autoregressive graph generation model based on the transformer to generate the initial layout plan of the building complex; Step 5: Scheme iteration and optimization Encode the building complex layout graph network obtained in step four into a high-dimensional node matrix. The rows in the matrix include the functions, heights, base areas, and normalized coordinates of the buildings, and combine with the global conditional vector to form a complete representation of the generated scheme; calculate the plot ratio, building density, and the height of the tallest building of the current generated scheme, and use the fitness function to quantitatively evaluate the scheme. If the fitness is equal to or greater than the threshold, save the current generated building complex layout scheme graph network; if the fitness is less than the threshold, then use the optimization algorithm to perform scheme optimization iteration based on the high-dimensional node matrix of the building complex layout, and repeat the optimization loop of step four and step five until the scheme indicators converge near the input design constraint data or reach the maximum number of iterations, and output the generated building complex layout scheme graph network; Step six: Layout scheme verification and output Save the building complex layout scheme graph network output in step five. According to the attributes of the nodes and the weights of the edges in the graph network, convert the graph network into a three-dimensional digital model of the building complex layout scheme, output it to the simulation device for feasibility verification of the scheme, and then perform refined adjustment and human-computer interaction display.
2. The automatic layout and iterative optimization method for neighborhood building groups based on the Transformer architecture according to claim 1, wherein, The basic data conversion method of the structured learnable graph attention matrix in step one performs feature processing on the original data through a hierarchical encoder, performs voxelization downsampling on the point cloud data, and extracts multi-scale terrain feature vectors through a 3D convolutional neural network; the formula for voxelization sampling is as follows: Among them, V (x,y) is the coordinate of the voxel center point, p i is the point cloud data point falling within the voxel, and N is the number of point cloud data points within the voxel; Convert the building layout data in the case library into a serialized representation: form a feature tuple with the center coordinates (x, y), floor area S, height H, and function type encoding F of each building unit, and generate a sequence after sorting by spatial proximity; convert the adjacency constraint into a learnable graph attention matrix, and use the graph neural network to pre-generate the relationship embedding vector between buildings.
3. The automatic layout and iterative optimization method for neighborhood building groups based on the Transformer architecture according to claim 2, wherein, In the second step, the spatial location and its attributes of the neighborhood building complex are used as input embedding vectors. A four-dimensional attribute vector composed of the spatial coordinates (x, y), area S, height H, and functional type code F of the building unit is converted into a d-dimensional initial embedding vector E through a linear mapping layer. i ∈R d ; The spatial features of the initial embedding vector are extracted through a multi-scale convolution kernel group: the first convolutional layer uses a 3×3 kernel to capture the neighborhood building density features; The second convolutional layer extracts mid-range functional group features using a 5×5 kernel; the third convolutional layer captures the global spatial distribution pattern using a 7×7 kernel and outputs the enhanced feature vector E'. i ∈R d .
4. The automatic layout and iterative optimization method for neighborhood building groups based on the Transformer architecture according to claim 3, characterized in that In step two, when converting the spatial topological relationship between building units into a graph structure, convert the case data into a graph structure G=(V, E), where the node V corresponds to the building unit, the node feature is E'i, and the weight of the edge E is calculated by the exponential decay function of the building spacing as: Among them, σ is the Gaussian kernel radius parameter, and the Gaussian kernel radius parameter σ is dynamically adjusted according to the building function type: σ = 50 meters for commercial building complexes; σ = 30 meters for residential building complexes; σ = 100 meters for industrial building complexes.
5. The automatic layout and iterative optimization method for neighborhood building complexes based on the Transformer architecture according to claim 4, characterized in that In step three, constraint conditions are injected at different levels of the Transformer encoder. Among them, for the feed-forward layer injection: concatenate the spatial constraint feature vector and the building unit feature vector and then input them into the fully connected layer; for the attention layer injection: encode the function constraint as an attention mask matrix, and the calculation formula is: Among them, M constraint is a 0-1 mask matrix generated by the functional compatibility rule.
6. The automatic layout and iterative optimization method for neighborhood building groups based on the Transformer architecture according to claim 5, characterized in that The plot shape encoder module in Step 4 refers to using a convolutional neural network and a multi-layer perceptron to extract the geometric features of the vector data of the plot to be designed and convert them into a shape feature vector h shape The global condition encoder refers to converting building density, plot ratio, and maximum building height into a condition vector through a multi-layer perceptron, encoding the land use function into a category vector through an embedding layer, and obtaining a global condition vector h after concatenating the condition vector and the category vector cond ; The autoregressive graph generation model based on Transformer in Step 4 refers to an intelligent generation model that can generate a building complex layout graph network, consisting of a Transformer decoder, a node generation module, an edge generation module, a coordinate predictor module, and a dynamic termination controller. The initial condition vector and the embedding of the start symbol are input into the autoregressive model graph generator based on Transformer. The node generation module includes a function classifier for predicting building functions and height and area regressors for predicting numerical values. The edge generation module includes an edge existence classifier for predicting whether there is an edge connection between nodes and an edge weight regressor for predicting the spacing numerical value. The preliminary building complex layout scheme generated in Step 4 means that, based on the initial condition vector and the embedding of the start symbol, the Transformer decoder uses the multi-head attention mechanism and based on the learned building complex layout rules to generate the attributes of the first building node, including building function, height, and floor area. The position of the building is determined within the boundary of the plot to be designed through the coordinate prediction module in combination with the plot mask. The remaining constructible capacity is updated according to the difference between the total floor area constraint index and the floor area of all generated nodes, and it is judged whether to continue generating new nodes. The embedding of the generated nodes is concatenated with the initial condition vector and then input into the decoder again. The above node generation process is repeated to generate the next node. At the same time, the spatial topological relationship between the new node and all generated nodes is calculated through the edge prediction module. If an edge is predicted between nodes, its weight is recorded. The above process is looped until the generation of new nodes stops after meeting the dynamic termination condition, and the building complex layout graph network is output.
7. The automatic layout and iterative optimization method for neighborhood building groups based on the Transformer architecture according to claim 6, characterized in that The formula for the fitness function in Step 5 is: Fitness = exp(-α·(L + β·Penalty)) L = λ1|BD in - BD g | + λ2|FAR in - FAR g | + λ3|H max - H g | Among them, L refers to the loss of the index proximity, Penalty is the penalty term for exceeding the plot boundary or insufficient spacing, α and β are weight coefficients, BD in is the building density of the generated building complex plan, BD g is the input building density constraint value, FAR in is the floor area ratio of the generated building complex plan, FAR g is the input floor area ratio constraint value, H max is the maximum building height of the generated building complex plan, H g is the input maximum building height constraint value; where the calculation formula for Penalty is: Penalty=Pendlty boundary +Penalty spacing where penalty boundary is the penalty term for exceeding the plot boundary, and penalty spacing is the penalty term for building spacing, x min and x max and y min and y max are the plot boundary coordinates, x i and y i are the ground center point coordinates of building i, and d min is the minimum spacing requirement between buildings.
8. A method for automatic layout and iterative optimization of neighborhood building complexes based on the Transformer architecture according to claim 7, characterized in that, The optimization algorithm used in Step 5 means using the genetic algorithm to perform perturbations in the high-dimensional space. After flattening the node matrix into a vector, the parent solutions with high fitness are selected through tournament selection. Block crossover and Gaussian mutation are applied to them, respectively, to exchange part of the building sub-matrix and randomly perturb the height and coordinates to generate offspring solutions. Subsequently, the projection operator is used to force the mutated coordinates within the plot boundary and recalculate the edge matrix to ensure the minimum spacing. After each iteration, the population is updated based on the elitist retention strategy, and the top 10% of the solutions with high fitness are retained and the inefficient individuals are replaced.
9. The automatic layout and iterative optimization method for neighborhood building complexes based on the Transformer architecture according to claim 8, characterized in that, The simulation device in Step 6 refers to a digital sand table integrated with building complex physical environment performance simulation algorithms, pedestrian flow simulation algorithms, AR reality augmentation glasses devices, and voice and gesture recognition module devices, which enables urban designers to finely optimize and adjust the building complex layout scheme according to the results of the simulation algorithms, and display the 3D model of the scheme, building functions, height, area indicators, and technical and economic indicators of the building complex through the AR reality augmentation glasses.
Citation Information
Cited By
Production line layout optimization method based on constraint perception Transform and soft permutation matrix
CN122047152A