TFM architecture-based automatic layout and iterative optimization method for urban block building clusters

WO2026199952A1PCT designated stage Publication Date: 2026-10-01SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/134547
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2025-11-13
Publication Date
2026-10-01

Smart Images

  • Figure CN2025134547_01102026_PF_FP_ABST
    Figure CN2025134547_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention is a TFM architecture-based automatic layout and iterative optimization method for urban block building clusters, the method comprising: constructing a structured learnable graph attention matrix by means of high-precision point cloud data, and, on the basis of a historical case library, encoding spatial topological relationships and economic indicators of building units; by means of a multi-head graph attention Transformer module, dynamically modelling spatial and functional association relationships of building clusters; integrating spatial and functional constraint conditions by means of a hierarchical constraint injection mechanism; by means of an autoregressive graph generation model, generating a preliminary layout scheme on the basis of plot features and global constraints; performing multi-objective iterative optimization by means of a genetic algorithm and a fitness function; and finally outputting a three-dimensional layout scheme satisfying multi-dimensional indicators such as the floor area ratio, the building density and the greening rate. The present invention overcomes the limitations in conventional human experience-driven approaches, supports adaptive learning and optimization under multi-constraint conditions, shortens the design cycle from weeks to hours, and significantly improves the scientific soundness and feasibility of schemes.
Need to check novelty before this filing date? Find Prior Art

Description

Automatic Layout and Iterative Optimization Method for Neighborhood Building Clusters Based on TFM Architecture Technical Field

[0001] This invention relates to the field of urban planning technology, and in particular to an automatic layout and iterative optimization method for neighborhood building clusters based on the TFM (Transformer) architecture. Background Technology

[0002] With the acceleration of urbanization, the planning and design of neighborhood building complexes face increasingly complex and diverse demands. Traditional building complex layout design methods mainly rely on human experience and rule-driven approaches, resulting in low efficiency, poor flexibility, and difficulty in meeting multi-objective optimization needs. In recent years, with the rapid development of artificial intelligence technology, automated design methods based on deep learning have gradually become a research hotspot. However, existing automated design methods are mostly limited to single-objective optimization (such as maximizing floor area ratio), making it difficult to comprehensively consider multi-dimensional factors such as spatial constraints, functional constraints, and environmental performance, and lacking the ability to dynamically model the spatial topological relationships of building complexes.

[0003] The Transformer (TFM) architecture, as a powerful sequence modeling tool, has achieved remarkable results in natural language processing and computer vision. Its self-attention mechanism can effectively capture long-distance dependencies, providing a new approach to modeling complex spatial relationships in building cluster layout design. However, existing research has not fully explored how to apply the Transformer architecture to the automatic layout and iterative optimization of neighborhood building clusters, especially in adaptive learning and dynamic optimization under multiple constraints. Therefore, there is an urgent need for an automatic layout and iterative optimization method for neighborhood building clusters based on the Transformer architecture, capable of efficiently processing high-precision point cloud data, dynamically modeling the spatial topology of building clusters, and automatically generating and optimizing layout schemes under multiple constraints to meet the efficiency, flexibility, and sustainability requirements of modern urban planning and design. Summary of the Invention

[0004] Purpose of the invention: The purpose of this invention is to provide an automatic layout and iterative optimization method for neighborhood building clusters based on the TFM architecture, so as to realize the generation and optimization of building cluster spatial layout schemes under multiple constraints.

[0005] The technical solution adopted in this invention is an automatic layout and iterative optimization method for neighborhood building clusters based on TFM architecture, which includes the following steps:

[0006] Step 1: Prepare the data platform and case library

[0007] The target neighborhood is 3D scanned using a LiDAR scanner to obtain high-precision point cloud data including terrain elevation, road boundaries, and existing building outlines. Simultaneously, planning constraint data, including floor area ratio, building setback distance, and height restriction area, is imported. Data on historical built-up areas from multiple cities is extracted from publicly available urban planning databases, and valid cases containing building coordinates, scale, and functional type information are selected to establish a standardized neighborhood building cluster database containing over 100 cases. Each case data point includes a spatial topology matrix of building units and a statistical table of economic and technical indicators, which are then encoded. A hierarchical encoder is used to perform feature processing on the raw data, transforming the basic data of the case database into a structured, learnable graph attention matrix that can be input via a Transformer.

[0008] Step Two: Learn the relationships between neighborhood building clusters

[0009] By using the self-attention mechanism of Transformer, the spatial location and attributes of the neighborhood building complex are used as the input embedding vector. Multi-scale convolutional kernel groups of convolutional neural networks are used as the front-end module to extract features from the input embedding vector and model the spatial relationship between building units in the case library.

[0010] The spatial attribute vectors of each case are mapped and transformed into a graph structure based on the spatial topological relationships between building units. Using a multi-head graph attention Transformer module, edge weights from the graph structure are introduced to adjust the attention distribution, enabling the model to dynamically perceive the relationships between nodes and perform adaptive attention learning. By stacking N layers of Transformer encoders, the implicit association matrix A∈R between building units is output. {M×M} This forms a relational graph network, where M is the total number of building units;

[0011] Step 3: Set design constraints

[0012] Design constraints are divided into two categories: spatial constraints and functional constraints. Constraint conditions are injected at different levels of the Transformer encoder. Spatial constraints include: maximum building height, floor area ratio threshold, building setback distance, and minimum greening rate. Functional constraints include: land use compatibility table, public service facility coverage area, and traffic noise isolation requirements. A constraint feature matrix is ​​generated using a hybrid method of numerical and symbolic encoding. Threshold constraints are converted into normalized scalar values, and logical constraints are converted into binary mask vectors.

[0013] Step 4: Generate a preliminary layout plan for the building complex.

[0014] The chassis data and design constraint data of the plot to be designed are input into the spatial layout generation model of the block building complex. The chassis data of the plot to be designed is converted into a plot feature vector by the plot shape encoder. The design constraint data is converted into a global condition vector by the global condition encoder. The plot feature vector and the global condition vector are concatenated to obtain the initial condition vector. The initial condition vector is input into the autoregressive graph generation model based on transformer to generate a preliminary building complex layout scheme.

[0015] Step 5: Solution Iteration and Optimization

[0016] The building layout diagram network obtained in step four is encoded into a high-dimensional node matrix. The rows of the matrix include the building's function, height, base area, and normalized coordinates. Combined with the global condition vector, a complete representation of the generated scheme is formed. The plot ratio, building density, and maximum building height of the current generated scheme are calculated. The scheme is quantitatively evaluated using a fitness function. If the fitness is equal to or greater than the threshold, the generated building layout scheme diagram network is saved. If the fitness is less than the threshold, an optimization algorithm is used to iterate the scheme optimization based on the high-dimensional node matrix of the building layout. Steps four and five are repeated in the optimization loop until the scheme indicators converge to near the input design constraint data or the maximum number of iterations is reached. The generated building layout scheme diagram network is then output.

[0017] Step Six: Layout Scheme Verification and Output

[0018] Save the architectural layout scheme graph network output in step five. Based on the attributes of the nodes and the weights of the edges in the graph network, convert the graph network into a three-dimensional digital model of the architectural layout scheme. Output the model to a simulation device to verify the feasibility of the scheme, then make fine adjustments and display it in a human-computer interaction.

[0019] Furthermore, the basic data transformation method for the structured learnable graph attention matrix in step one involves feature processing of the original data through a hierarchical encoder, voxelization downsampling of the point cloud data, and extraction of multi-scale terrain feature vectors through a 3D convolutional neural network; the formula for voxelization sampling is as follows:

[0020] Among them, V (x,y) These are the coordinates of the voxel center point, p i N represents the point cloud data points that fall within the voxel, and N is the number of point cloud data points within the voxel.

[0021] The building layout data in the case library is converted into a serialized representation: feature tuples are constructed using the center coordinates (x,y), floor area (S), height (H), and function type code (F) of each building unit, and then the sequence is generated after being sorted by spatial proximity; the adjacency constraints are converted into a learnable graph attention matrix, and graph neural networks are used to pre-generate the relationship embedding vectors between buildings.

[0022] Furthermore, in step two, the spatial location and attributes of the neighborhood building complex are used as the input embedding vector. A four-dimensional attribute vector is constructed, comprising the spatial coordinates (x, y) of the building unit, area S, height H, and functional type code F. This vector is then converted into a d-dimensional initial embedding vector E through a linear mapping layer. i ∈R d Spatial features are extracted from the initial embedding vector using a multi-scale convolutional kernel group: the first convolutional layer uses a 3×3 kernel to capture the density features of neighboring buildings; the second convolutional layer uses a 5×5 kernel to extract the features of mid-range functional clusters; and the third convolutional layer uses a 7×7 kernel to capture the global spatial distribution pattern, outputting the enhanced feature vector E′. i ∈R d .

[0023] Furthermore, in step two, based on the spatial topological relationships between building units, the case data is converted into a graph structure G = (V, E), where node V corresponds to a building unit, and the node feature is E′. i The weight of edge E is calculated using the exponential decay function of the building spacing as follows:

[0024] Wherein, σ is the Gaussian kernel radius parameter, which is dynamically adjusted according to the building function type: σ = 50 meters for commercial building complexes; σ = 30 meters for residential building complexes; and σ = 100 meters for industrial building complexes.

[0025] Furthermore, in step three, constraints are injected at different levels of the Transformer encoder. Specifically, in the feedforward layer injection: the spatial constraint feature vector is concatenated with the building unit feature vector and then input into the fully connected layer; in the attention layer injection: the functional constraints are encoded into an attention mask matrix, calculated using the following formula:

[0026] Among them, M constraint This is a 0-1 mask matrix generated by the functional compatibility rules.

[0027] Furthermore, the fourth step, the land parcel shape encoder module, refers to using a convolutional neural network and a multilayer perceptron to extract the geometric features of the vector data of the land parcel to be designed, and converting it into a shape feature vector h. shapeThe global condition encoder refers to converting building density, floor area ratio, and maximum building height into condition vectors using a multilayer perceptron, encoding land use functions into category vectors using an embedding layer, and concatenating the condition vectors and category vectors to obtain the global condition vector h. cond .

[0028] Furthermore, the fourth step, the transformer-based autoregressive graph generation model, refers to an intelligent generation model capable of generating a building cluster layout network, consisting of a transformer decoder, a node generation module, an edge generation module, a coordinate predictor module, and a dynamic termination controller. The initial condition vector and the embedded start symbol are input into the transformer-based autoregressive graph generator. The node generation module includes a function classifier for predicting building functions and height and area regressors for predicting numerical values. The edge generation module includes an edge existence classifier for predicting whether there are edge connections between nodes and an edge weight regressor for predicting spacing values.

[0029] Furthermore, step four, generating a preliminary building cluster layout scheme, refers to using a transformer decoder with a multi-head attention mechanism and learned building cluster layout rules to generate the first building node, whose attributes include building function, height, and base area, based on the embedding of the initial condition vector and the starting symbol. The coordinate prediction module, combined with the plot mask, determines that the building's location is within the boundary of the plot to be designed. The remaining buildable capacity is updated based on the difference between the total building area constraint index and the building area of ​​all generated nodes, and it is determined whether to continue generating new nodes. The embedding of the generated node is concatenated with the initial condition vector and input into the decoder again. The above node generation process is repeated to generate the next node. At the same time, the edge prediction module calculates the spatial topological relationship between the new node and all generated nodes. If there is an edge between the predicted nodes, its weight is recorded. The above process is repeated until the dynamic termination condition is met, at which point the generation of new nodes stops, and the building cluster layout graph network is output.

[0030] Furthermore, the formula for the fitness function in step five is: Fitness=exp(-α·(L+β·Penalty)) L=λ1∣BD in -BD g ∣+λ2∣FAR in -FAR g |+λ3|H max -H g |

[0031] Where L refers to the loss due to proximity to the indicator, Penalty is the penalty for exceeding the plot boundary or insufficient spacing, α and β are weighting coefficients, and BD in BD is used to calculate the building density of the generated building complex scheme.g For the input building density constraint value, FAR in For the floor area ratio of the generated building complex scheme, FAR g H is the input floor area ratio constraint value. max H represents the maximum building height of the generated building complex design. g This is the maximum building height constraint value that you input.

[0032] The formula for calculating Penalty is: Penalty = Penalty boundary +Penalty spacing

[0033] Among them penalty boundary Penalty for exceeding land parcel boundaries spacing x is the penalty for building spacing. min x max y min y max Let x be the boundary coordinates of the land parcel. i y i Let d be the coordinates of the ground center point of building i. min It refers to the minimum spacing requirement between buildings.

[0034] Furthermore, step five uses an optimization algorithm, which refers to using a genetic algorithm to perturb the nodes in a high-dimensional space. After flattening the node matrix into a vector, a tournament selection is used to select parent schemes with high fitness. Block crossover and Gaussian mutation are applied to these parent schemes, and partial building submatrices are exchanged and the height and coordinates are randomly perturbed to generate offspring schemes. Subsequently, the projection operator is used to force the mutated coordinates to be constrained to the plot boundary, and the edge matrix is ​​recalculated to ensure the minimum spacing. After each iteration, the population is updated based on the elite retention strategy, retaining the top 10% of the schemes with the highest fitness and replacing inefficient individuals.

[0035] Furthermore, the simulation equipment mentioned in step six refers to a digital sand table that integrates building complex physical environment performance simulation algorithms, pedestrian flow simulation algorithms, AR augmented reality glasses, and voice and gesture recognition modules. This enables urban designers to finely optimize and adjust the building complex layout plan based on the results of the simulation algorithms, and to display the three-dimensional model of the plan and the building functions, height, area indicators, and technical and economic indicators of the building complex through AR augmented reality glasses.

[0036] Beneficial effects: Compared with the prior art, the advantages of this invention are that it realizes automatic layout and iterative optimization of neighborhood building clusters based on the Transformer architecture, establishes a dynamic association between the spatial topology of building clusters and planning constraints, and realizes automatic generation and optimization of building cluster layout schemes that meet multiple constraints through self-attention mechanism and adaptive learning. Specifically, the advantages include the following:

[0037] 1. This invention establishes a dynamic correlation between the spatial location, functional attributes and planning constraints of building units by constructing a spatial topology prediction model for building clusters. This overcomes the shortcomings of traditional building cluster layout design that relies solely on human experience and rule-driven approaches, and improves the accuracy and scientific nature of layout scheme generation.

[0038] 2. This invention achieves adaptive learning and iterative optimization of building cluster layout schemes through the self-attention mechanism of Transformer and the multi-head graph attention module, shortening the original design cycle of several weeks to within a few hours, and improving the standardization and normalization of the scheme output.

[0039] 3. This invention achieves multi-objective optimization through the comprehensive application of fitness functions and optimization algorithms, which can simultaneously meet the planning index requirements of multiple dimensions such as plot ratio, building density, and greening rate, and significantly improve the feasibility and practicality of the design scheme. Attached Figure Description

[0040] Figure 1 is a flowchart of the automatic optimization method of the present invention;

[0041] Figure 2 is a schematic diagram of the building group spatial topology modeling and Transformer architecture of the present invention;

[0042] Figure 3 is a schematic diagram of the process of generating and optimizing the building group layout scheme of the present invention. Detailed Implementation

[0043] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0044] As shown in Figure 1-3, the automatic layout and iterative optimization method for neighborhood building clusters based on the TFM architecture includes the following steps.

[0045] I. Prepare the data platform and case library. Use a LiDAR scanner to perform 3D scanning of the target blocks to obtain high-precision point cloud data including terrain elevation, road boundaries, and existing building outlines. Simultaneously import planning constraint data, including floor area ratio, building setback distance, and height restriction area. Extract historical built-up area data from public urban planning databases of multiple cities, and filter valid cases containing building coordinates, building scale, and functional type information to establish a standardized block building cluster database containing 100+ cases. Encode the spatial topology matrix and economic and technical indicator statistics table of each case data. Perform feature processing on the raw data using a hierarchical encoder, converting the basic data of the case library into a structured, learnable graph attention matrix that can be input via a Transformer.

[0046] The structured, learnable graph attention matrix-based data transformation method involves feature processing of the original data using a hierarchical encoder, voxelization downsampling of the point cloud data, and extraction of multi-scale terrain feature vectors using a 3D convolutional neural network. The voxelization sampling formula is as follows:

[0047] Among them, V (x,y) These are the coordinates of the voxel center point, p i N represents the point cloud data points that fall within the voxel, and N is the number of point cloud data points within the voxel.

[0048] The building layout data in the case library is converted into a serialized representation: feature tuples are constructed using the center coordinates (x, y), floor area (S), height (H), and function type code (F) of each building unit, and then sorted by spatial proximity to generate a sequence. Adjacency constraints are transformed into learnable graph attention matrices, and graph neural networks are used to pre-generate relationship embedding vectors between buildings.

[0049] Second, learn the relationship between neighborhood building groups. Through the self-attention mechanism of Transformer, the spatial location and attributes of neighborhood building groups are used as input embedding vectors. Multi-scale convolutional kernel groups of convolutional neural networks are used as front-end modules to extract features from the input embedding vectors and model the spatial relationship between building units in the case library.

[0050] The spatial attribute vectors of each case are mapped and transformed into a graph structure based on the spatial topological relationships between building units. Using a multi-head graph attention Transformer module, edge weights from the graph structure are introduced to adjust the attention distribution, enabling the model to dynamically perceive the relationships between nodes and perform adaptive attention learning. By stacking N layers of Transformer encoders, the implicit association matrix A∈R between building units is output. {M×M}This forms a relational graph network, where M is the total number of building units;

[0051] Specifically, the spatial location and attributes of the neighborhood building complex are used as input embedding vectors. A four-dimensional attribute vector is constructed, comprising the spatial coordinates (x, y), area (S), number of floors (H), and functional type code (F) of the building units. This vector is then converted into a d-dimensional initial embedding vector E through a linear mapping layer. i ∈R d Spatial features are extracted from the initial embedding vector using a multi-scale convolutional kernel group: the first convolutional layer uses a 3×3 kernel to capture the density features of neighboring buildings; the second convolutional layer uses a 5×5 kernel to extract the features of mid-range functional clusters; and the third convolutional layer uses a 7×7 kernel to capture the global spatial distribution pattern, outputting the enhanced feature vector E′. i ∈R d .

[0052] In the process of converting the spatial topological relationships between building units into a graph structure, the case data is converted into a graph structure G = (V, E), where node V corresponds to a building unit and the node feature is E′. i The weight of edge E is calculated using the exponential decay function of the building spacing as follows:

[0053] Wherein, σ is the Gaussian kernel radius parameter, which is dynamically adjusted according to the building function type: σ = 50 meters for commercial building complexes; σ = 30 meters for residential building complexes; and σ = 100 meters for industrial building complexes.

[0054] 3. Set design constraints, which are divided into two categories: spatial constraints and functional constraints. These constraints are injected into different levels of the Transformer encoder. Spatial constraints include: maximum building height, floor area ratio threshold, building setback distance, and minimum greening rate. Functional constraints include: land use compatibility table, public service facility coverage area, and traffic noise isolation requirements. A constraint feature matrix is ​​generated using a hybrid numerical and symbolic encoding method. Threshold constraints are converted into normalized scalar values, and logical constraints are converted into binary mask vectors.

[0055] The Transformer encoder injects constraints at different levels. Specifically, in the feedforward layer, the spatial constraint feature vector is concatenated with the building unit feature vector and then input into the fully connected layer. In the attention layer, functional constraints are encoded into an attention mask matrix, calculated using the following formula:

[0056] Among them, M constraint This is a 0-1 mask matrix generated by the functional compatibility rules.

[0057] IV. Initially generate the building cluster layout scheme. Input the chassis data and design constraint data of the plot to be designed into the block building cluster spatial layout generation model. Convert the chassis data of the plot to be designed into a plot feature vector through the plot shape encoder. Convert the design constraint data into a global condition vector through the global condition encoder. Concatenate the plot feature vector and the global condition vector to obtain the initial condition vector. Input the initial condition vector into the transformer-based autoregressive graph generation model to generate the initial building cluster layout scheme.

[0058] The plot shape encoder module refers to using convolutional neural networks and multilayer perceptrons to extract the geometric features of the vector data of the plot to be designed, and converting it into a shape feature vector h. shape The global condition encoder refers to converting building density, floor area ratio, and maximum building height into condition vectors using a multilayer perceptron, encoding land use functions into category vectors using an embedding layer, and concatenating the condition vectors and category vectors to obtain the global condition vector h. cond .

[0059] The transformer-based autoregressive graph generation model refers to an intelligent generation model capable of generating a network of building cluster layout diagrams, consisting of a transformer decoder, a node generation module, an edge generation module, a coordinate predictor module, and a dynamic termination controller. The initial condition vector and the embedded start symbol are input into the transformer-based autoregressive graph generator. The node generation module includes a function classifier for predicting building functions and height and area regressors for predicting numerical values. The edge generation module includes an edge existence classifier for predicting whether there are edge connections between nodes and an edge weight regressor for predicting spacing values.

[0060] The process of generating a preliminary building cluster layout scheme involves using an embedding of an initial condition vector and a starting symbol, employing a transformer decoder with a multi-head attention mechanism and learned building cluster layout patterns to generate the first building node, whose attributes include building function, height, and floor area. A coordinate prediction module, combined with a plot mask, determines that the building's location is within the boundary of the plot to be designed. The remaining buildable capacity is updated based on the difference between the total building area constraint and the building areas of all generated nodes, and a decision is made on whether to continue generating new nodes. The embeddings of the generated nodes are concatenated with the initial condition vector and input back into the decoder. This node generation process is repeated to generate the next node. Simultaneously, an edge prediction module calculates the spatial topological relationship between the new node and all generated nodes; if an edge exists between predicted nodes, its weight is recorded. This process is repeated until a dynamic termination condition is met, at which point the generation of new nodes stops, and the building cluster layout graph network is output.

[0061] V. Scheme Iteration and Optimization: The building cluster layout network obtained in Step IV is encoded into a high-dimensional node matrix. The rows of the matrix include the building's function, height, base area, and normalized coordinates. Combined with the global condition vector, a complete representation of the generated scheme is formed. The plot ratio, building density, and maximum building height of the current generated scheme are calculated. The fitness function is used to quantitatively evaluate the scheme. If the fitness is equal to or greater than the threshold, the current generated building cluster layout scheme network is saved. If the fitness is less than the threshold, the optimization algorithm is used to iterate the scheme optimization based on the high-dimensional node matrix of the building cluster layout. Steps IV and V are repeated in an optimization loop until the scheme indicators converge to near the input design constraint data or the maximum number of iterations is reached. The generated building cluster layout scheme network is then output.

[0062] The fitness function is defined as follows: Fitness=exp(-α·(L+β·Penalty)) L=λ1∣BD in -BD g ∣+λ2∣FAR in -FAR g |+λ3|H max -H g |

[0063] Where L refers to the loss due to proximity to the indicator, Penalty is the penalty for exceeding the plot boundary or insufficient spacing, α and β are weighting coefficients, and BD in BD is used to calculate the building density of the generated building complex scheme. g For the input building density constraint value, FAR in For the floor area ratio of the generated building complex scheme, FAR g H is the input floor area ratio constraint value. max H represents the maximum building height of the generated building complex design. g This is the maximum building height constraint value that you input.

[0064] The formula for calculating Penalty is: Penalty = Penalty boundary +Penalty spacing

[0065] Among them penalty boundary Penalty for exceeding land parcel boundaries spacing x is the penalty for building spacing. min x max y min y max Let x be the boundary coordinates of the land parcel. i y i Let d be the coordinates of the ground center point of building i. minIt refers to the minimum spacing requirement between buildings.

[0066] The optimization algorithm mentioned refers to using a genetic algorithm to perturb in a high-dimensional space, flattening the node matrix into a vector, selecting parent schemes with high fitness through tournament selection, applying block crossover and Gaussian mutation to them, exchanging part of the building submatrices and randomly perturbing the height and coordinates to generate offspring schemes; then using the projection operator to force the mutated coordinates to be constrained to within the plot boundary, and recalculating the edge matrix to ensure minimum spacing. After each iteration, the population is updated based on the elite retention strategy, retaining the top 10% of the schemes with the highest fitness and replacing inefficient individuals.

[0067] VI. Layout Scheme Verification and Output: Save the architectural layout scheme graph network output in step five. Based on the attributes of the nodes and the weights of the edges in the graph network, convert the graph network into a three-dimensional digital model of the architectural layout scheme. Output the model to the simulation equipment to verify the feasibility of the scheme, make fine adjustments, and then perform human-computer interaction display.

[0068] The simulation equipment refers to a digital sand table that integrates algorithms for simulating the physical environment performance of building complexes, algorithms for simulating pedestrian flow, AR augmented reality glasses, and voice and gesture recognition modules. It enables urban designers to finely optimize and adjust the layout of building complexes based on the results of the simulation algorithms, and to display the three-dimensional model of the plan and the building functions, height, area indicators, and technical and economic indicators of the building complex through AR augmented reality glasses.

[0069] Example

[0070] The technical solution of this invention will be described in detail below using Nanjing Hexi CBD as an example.

[0071] (1) Prepare the data platform and case library. Use a laser scanner (LiDAR) to perform a 3D scan of Nanjing Hexi CBD to obtain high-precision point cloud data containing topographic elevation, road boundary lines, and existing building outlines. Simultaneously import planning condition constraint data, which includes plot ratio, building setback distance, and height restriction area range. Extract historical built-up area data from public urban planning databases of multiple cities, screen effective cases containing building coordinates, building scale, and functional type information, and establish a standardized neighborhood building group database containing 100+ cases. Encode the spatial topology relationship matrix of building units and the economic and technical indicator statistics table of each case data. Use a hierarchical encoder to perform feature processing on the original data and convert the basic data of the case library into a structured learnable graph attention matrix that can be input through Transformer.

[0072] (1.1) The basic data transformation method for the structured learnable graph attention matrix described above performs feature processing on the original data through a hierarchical encoder, performs voxelization downsampling on the point cloud data, and extracts multi-scale terrain feature vectors through a 3D convolutional neural network; the formula for voxelization sampling is as follows:

[0073] Among them, V (x,y) These are the coordinates of the voxel center point, p i N represents the point cloud data points that fall within the voxel, and N is the number of point cloud data points within the voxel.

[0074] The building layout data in the case library is converted into a serialized representation: feature tuples are constructed using the center coordinates (x,y), floor area (S), height (H), and function type code (F) of each building unit, and then the sequence is generated after being sorted by spatial proximity; the adjacency constraints are converted into a learnable graph attention matrix, and graph neural networks are used to pre-generate the relationship embedding vectors between buildings.

[0075] (2) Learn the relationship between the building clusters in the neighborhood. Using the self-attention mechanism of Transformer, the spatial location and attributes of the building clusters in Nanjing Hexi CBD are used as the input embedding vector. The multi-scale convolution kernel group of the convolutional neural network is used as the front module to extract features from the input embedding vector and model the spatial relationship between the building units in the case library.

[0076] The spatial attribute vectors of each case are mapped and transformed into a graph structure based on the spatial topological relationships between building units. Using a multi-head graph attention Transformer module, edge weights from the graph structure are introduced to adjust the attention distribution, enabling the model to dynamically perceive the relationships between nodes and perform adaptive attention learning. By stacking N layers of Transformer encoders, the implicit association matrix A∈R between building units is output. {M×M} This forms a relational graph network, where M is the total number of building units;

[0077] (2.1) The spatial location and attributes of the Nanjing Hexi CBD block building complex are used as the input embedding vector. The spatial coordinates (x, y), area S, height H, and functional type code F of the building unit are used to form a four-dimensional attribute vector, which is then converted into a d-dimensional initial embedding vector E through a linear mapping layer. i ∈R d Spatial features are extracted from the initial embedding vector using a multi-scale convolutional kernel group: the first convolutional layer uses a 3×3 kernel to capture the density features of neighboring buildings; the second convolutional layer uses a 5×5 kernel to extract the features of mid-range functional clusters; and the third convolutional layer uses a 7×7 kernel to capture the global spatial distribution pattern, outputting the enhanced feature vector E′. i ∈R d .

[0078] (2.2) In the conversion of the spatial topological relationship between building units in Nanjing Hexi CBD into a graph structure, the case data is converted into a graph structure G=(V,E), where node V corresponds to a building unit and the node feature is E′. i The weight of edge E is calculated using the exponential decay function of the building spacing as follows:

[0079] Wherein, σ is the Gaussian kernel radius parameter, which is dynamically adjusted according to the building function type: σ = 50 meters for commercial building complexes; σ = 30 meters for residential building complexes; and σ = 100 meters for industrial building complexes.

[0080] (3) Set design constraints and divide them into two categories: spatial constraints and functional constraints. Inject the constraints at different levels of the Transformer encoder. Among them, spatial constraints include: maximum building height, floor area ratio threshold, building setback distance, and minimum greening rate; functional constraints include: land use compatibility table, public service facility radiation range, and traffic noise isolation requirements. Generate the constraint feature matrix using a hybrid method of numerical encoding and symbolic encoding. Threshold constraints are converted into normalized scalar values, and logical constraints are converted into binary mask vectors.

[0081] Constraints are injected at different levels of the Transformer encoder. Specifically, in the feedforward layer, the spatial constraint feature vector is concatenated with the building unit feature vector and then input into the fully connected layer. In the attention layer, functional constraints are encoded as an attention mask matrix, calculated using the following formula:

[0082] Among them, M constraint This is a 0-1 mask matrix generated by the functional compatibility rules.

[0083] (4) A preliminary layout scheme for the Nanjing Hexi CBD building complex is generated. The base data and design constraint data of the plot to be designed are input into the spatial layout generation model of the Nanjing Hexi CBD block building complex. The base data of the Nanjing Hexi CBD plot is converted into a plot feature vector through the plot shape encoder. The design constraint data is converted into a global condition vector through the global condition encoder. The plot feature vector and the global condition vector are concatenated to obtain the initial condition vector. The initial condition vector is input into the autoregressive graph generation model based on transformer to generate a preliminary layout scheme for the Nanjing Hexi CBD building complex.

[0084] (4.1) The shape encoder module for the Nanjing Hexi CBD plot refers to the module that uses convolutional neural networks and multilayer perceptrons to extract the geometric features of the vector data of the plot to be designed, and converts it into a shape feature vector h. shapeThe global condition encoder refers to converting building density, floor area ratio, and maximum building height into condition vectors using a multilayer perceptron, encoding land use functions into category vectors using an embedding layer, and concatenating the condition vectors and category vectors to obtain the global condition vector h. cond .

[0085] (4.2) The transformer-based autoregressive graph generation model refers to an intelligent generation model that can generate a network of building layout diagrams, consisting of a transformer decoder, a node generation module, an edge generation module, a coordinate predictor module, and a dynamic termination controller. The initial condition vector and the embedded start symbol are input into the transformer-based autoregressive graph generator. The node generation module includes a function classifier for predicting building functions and height and area regressors for predicting numerical values. The edge generation module includes an edge existence classifier for predicting whether there are edge connections between nodes and an edge weight regressor for predicting spacing values.

[0086] (4.3) The generation of the preliminary Nanjing Hexi CBD building cluster layout scheme refers to the process of generating the first building node with attributes including building function, height, and base area based on the embedding of the initial condition vector and the starting symbol, using a transformer decoder through a multi-head attention mechanism and based on the learned building cluster layout rules. The coordinate prediction module is used to determine the location of the building within the boundary of the plot to be designed, and the remaining buildable capacity is updated according to the difference between the total building area constraint index and the building area of ​​all generated nodes. It is then determined whether to continue generating new nodes. The embedding of the generated node is concatenated with the initial condition vector and input into the decoder again. The above node generation process is repeated to generate the next node. At the same time, the spatial topological relationship between the new node and all generated nodes is calculated through the edge prediction module. If there is an edge between the predicted nodes, its weight is recorded. The above process is repeated until the dynamic termination condition is met, at which point the generation of new nodes stops and the Nanjing Hexi CBD building cluster layout map network is output.

[0087] (5) Scheme iteration and optimization: The Nanjing Hexi CBD building cluster layout map network obtained in step four is encoded into a high-dimensional node matrix. The rows in the matrix include the building function, number of floors, base area and normalized coordinates. Combined with the global condition vector, a complete representation of the generated scheme is formed. The plot ratio, building density and the highest number of building floors of the current Nanjing Hexi CBD generated scheme are calculated. The fitness function is used to quantitatively evaluate the scheme. If the fitness is equal to or greater than the threshold, the current generated building cluster layout scheme map network is saved. If the fitness is less than the threshold, the optimization algorithm is used to optimize the scheme iteratively based on the high-dimensional node matrix of the building cluster layout. The optimization loop of step four and step five is repeated until the Nanjing Hexi CBD scheme indicators converge to the vicinity of the input design constraint data or reach the maximum number of iterations. The generated Nanjing Hexi CBD building cluster layout scheme map network is then output.

[0088] The fitness function is defined as follows: Fitness=exp(-α·(L+β·Penalty)) L=λ1∣BD in -BD g ∣+λ2∣FAR in -FAR g |+λ3|H max -H g |

[0089] Where L refers to the loss due to proximity to the indicator, Penalty is the penalty for exceeding the plot boundary or insufficient spacing, α and β are weighting coefficients, and BD in BD is used to calculate the building density of the generated building complex scheme. g For the input building density constraint value, FAR in For the floor area ratio of the generated building complex scheme, FAR g H is the input floor area ratio constraint value. max H represents the maximum building height of the generated building complex design. g This is the maximum building height constraint value that you input.

[0090] The formula for calculating Penalty is: Penalty = Penalty boundary +Penalty spacing

[0091] Among them penalty boundary Penalty for exceeding land parcel boundaries spacing x is the penalty for building spacing. min x max y min y max Let x be the boundary coordinates of the land parcel. i yi Let d be the coordinates of the ground center point of building i. min It refers to the minimum spacing requirement between buildings.

[0092] (5.1) The optimization algorithm mentioned refers to using a genetic algorithm to perturb in a high-dimensional space, flattening the node matrix into a vector, selecting parent schemes with high fitness through tournament selection, applying block crossover and Gaussian mutation to them, exchanging part of the building submatrix and randomly perturbing the height and coordinates to generate offspring schemes; then using the projection operator to force the mutated coordinates to be constrained to the plot boundary, and recalculating the edge matrix to ensure the minimum spacing. After each iteration, the population is updated based on the elite retention strategy, retaining the top 10% of the schemes with fitness and replacing inefficient individuals.

[0093] (6) Verification and output of Nanjing Hexi CBD layout scheme: Save the Nanjing Hexi CBD building group layout scheme graph network output in step five. Based on the attributes of the nodes and the weights of the edges in the graph network, convert the graph network into a three-dimensional digital model of the building group layout scheme. Output the model to the simulation equipment to verify the feasibility of the scheme, make fine adjustments, and perform human-computer interaction display.

[0094] (6.1) The simulation equipment refers to a digital sand table that integrates building complex physical environment performance simulation algorithm, pedestrian flow simulation algorithm, AR augmented reality glasses equipment and voice and gesture recognition module equipment. It enables urban designers to finely optimize and adjust the building complex layout scheme based on the results of the simulation algorithm, and display the three-dimensional model of the scheme and the building functions, height, area indicators and technical and economic indicators of the building complex through AR augmented reality glasses.

[0095] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0096] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. An automatic layout and iterative optimization method for neighborhood building clusters based on TFM architecture, characterized in that, Includes the following steps: Step 1: Prepare the data platform and case library A laser scanner is used to perform a 3D scan of the target neighborhood, acquiring high-precision point cloud data including terrain elevation, road boundaries, and existing building outlines. Simultaneously, planning constraint data, including floor area ratio, building setback distance, and height restriction area, is imported. Data on historical built-up areas from multiple cities is extracted from publicly available urban planning databases, and valid cases containing building coordinates, building scale, and functional type information are selected to establish a standardized neighborhood building cluster database containing over 100 cases. Each case data point includes a spatial topology matrix of building units and a statistical table of economic and technical indicators, which are then encoded. A hierarchical encoder is used to perform feature processing on the raw data, transforming the basic data of the case database into a structured, learnable graph attention matrix that can be input via a Transformer. Step Two: Learn the relationships between neighborhood building clusters By using the self-attention mechanism of Transformer, the spatial location and attributes of the neighborhood building complex are used as input embedding vectors. Multi-scale convolutional kernel groups of convolutional neural networks are used as front-end modules to extract features from the input embedding vectors and model the spatial relationship between building units in the case library. The spatial attribute vectors of each case are mapped and converted into a graph structure based on the spatial topological relationship between building units. Based on the multi-head graph attention Transformer module, the edge weights of the graph structure are introduced to adjust the attention distribution, enabling the model to dynamically perceive the relationship between nodes and perform adaptive attention learning. By stacking N layers of Transformer encoders, the implicit correlation matrix A∈R between building units is output. {M×M} This forms a relational graph network, where M is the total number of building units; Step 3: Set design constraints Design constraints are divided into two categories: spatial constraints and functional constraints. Constraint conditions are injected at different levels of the Transformer encoder. Spatial constraints include: maximum building height, floor area ratio threshold, building setback distance, and minimum greening rate. Functional constraints include: land use compatibility table, public service facility coverage area, and traffic noise isolation requirements. A constraint feature matrix is ​​generated using a hybrid method of numerical and symbolic encoding. Threshold constraints are converted into normalized scalar values, and logical constraints are converted into binary mask vectors. Step 4: Generate a preliminary layout plan for the building complex. The chassis data and design constraint data of the plot to be designed are input into the spatial layout generation model of the block building complex. The chassis data of the plot to be designed is converted into a plot feature vector by the plot shape encoder. The design constraint data is converted into a global condition vector by the global condition encoder. The plot feature vector and the global condition vector are concatenated to obtain the initial condition vector. The initial condition vector is input into the autoregressive graph generation model based on transformer to generate a preliminary building complex layout scheme. Step 5: Solution Iteration and Optimization The building layout diagram network obtained in step four is encoded into a high-dimensional node matrix. The rows of the matrix include the building's function, height, base area, and normalized coordinates. Combined with the global condition vector, a complete representation of the generated scheme is formed. The plot ratio, building density, and maximum building height of the current generated scheme are calculated. The scheme is quantitatively evaluated using a fitness function. If the fitness is equal to or greater than the threshold, the generated building layout scheme diagram network is saved. If the fitness is less than the threshold, an optimization algorithm is used to iterate the scheme optimization based on the high-dimensional node matrix of the building layout. Steps four and five are repeated in the optimization loop until the scheme indicators converge to near the input design constraint data or the maximum number of iterations is reached. The generated building layout scheme diagram network is then output. Step Six: Layout Scheme Verification and Output Save the architectural layout scheme graph network output in step five. Based on the attributes of the nodes and the weights of the edges in the graph network, convert the graph network into a three-dimensional digital model of the architectural layout scheme. Output the model to a simulation device to verify the feasibility of the scheme, then make fine adjustments and display it in a human-computer interaction.

2. The method for automatic layout and iterative optimization of neighborhood building clusters based on Transformer architecture according to claim 1, characterized in that, The basic data transformation method for the structured learnable graph attention matrix in step one involves feature processing of the original data through a hierarchical encoder, voxelization downsampling of the point cloud data, and extraction of multi-scale terrain feature vectors through a 3D convolutional neural network. The voxelization sampling formula is as follows: Among them, V (x,y) These are the coordinates of the voxel center point, p i N represents the point cloud data points that fall within the voxel, and N is the number of point cloud data points within the voxel. The building layout data in the case library is converted into a serialized representation: feature tuples are constructed using the center coordinates (x,y), floor area S, height H, and function type code F of each building unit, and then generated into a sequence after being sorted by spatial proximity; the adjacency constraints are converted into a learnable graph attention matrix, and graph neural networks are used to pre-generate relationship embedding vectors between buildings.

3. The method for automatic layout and iterative optimization of neighborhood building clusters based on Transformer architecture according to claim 2, characterized in that, In step two, the spatial location and attributes of the neighborhood building complex are used as the input embedding vector. A four-dimensional attribute vector is constructed, which includes the spatial coordinates (x, y) of the building unit, area S, height H, and function type code F. This vector is then converted into a d-dimensional initial embedding vector E through a linear mapping layer. i ∈R d Spatial features are extracted from the initial embedding vector using a multi-scale convolutional kernel group: the first convolutional layer uses a 3×3 kernel to capture the density features of neighboring buildings; The second convolutional layer uses a 5×5 kernel to extract mid-range functional cluster features; the third convolutional layer uses a 7×7 kernel to capture the global spatial distribution pattern and outputs the enhanced feature vector E′. i ∈R d .

4. The method for automatic layout and iterative optimization of neighborhood building clusters based on Transformer architecture according to claim 3, characterized in that, In step two, based on the spatial topological relationships between building units, the case data is converted into a graph structure G = (V, E), where node V corresponds to a building unit, and the node feature is E′. i The weight of edge E is calculated using the exponential decay function of the building spacing as follows: Wherein, σ is the Gaussian kernel radius parameter, which is dynamically adjusted according to the building function type: σ = 50 meters for commercial building complexes; σ = 30 meters for residential building complexes; and σ = 100 meters for industrial building complexes.

5. The method for automatic layout and iterative optimization of neighborhood building complexes based on Transformer architecture according to claim 4, characterized in that, Step three involves injecting constraints at different levels of the Transformer encoder. Specifically, the feedforward layer injection involves concatenating the spatial constraint feature vector with the building unit feature vector and then inputting the concatenated vector into the fully connected layer. The attention layer injection involves encoding the functional constraints into an attention mask matrix, calculated using the following formula: Among them, M constraint This is a 0-1 mask matrix generated by the functional compatibility rules.

6. The method for automatic layout and iterative optimization of neighborhood building clusters based on Transformer architecture according to claim 5, characterized in that, The fourth step, the land parcel shape encoder module, refers to using a convolutional neural network and a multilayer perceptron to extract the geometric features of the vector data of the land parcel to be designed, and converting it into a shape feature vector h. shape The global condition encoder refers to converting building density, floor area ratio, and maximum building height into condition vectors using a multilayer perceptron, encoding land use functions into category vectors using an embedding layer, and concatenating the condition vectors and category vectors to obtain the global condition vector h. cond ; The fourth step, the transformer-based autoregressive graph generation model, refers to an intelligent generation model capable of generating a network of building cluster layout diagrams, consisting of a transformer decoder, a node generation module, an edge generation module, a coordinate predictor module, and a dynamic termination controller. The initial condition vector and the embedded start symbol are input into the transformer-based autoregressive graph generator. The node generation module includes a function classifier for predicting building functions and height and area regressors for predicting numerical values. The edge generation module includes an edge existence classifier for predicting whether there are edge connections between nodes and an edge weight regressor for predicting spacing values. Step four, generating a preliminary building cluster layout scheme, involves using a transformer decoder with a multi-head attention mechanism and learned building cluster layout patterns to generate the first building node, whose attributes include building function, height, and floor area. This is achieved by embedding the initial condition vector and starting symbols, and then generating the first building node. The coordinate prediction module, combined with a plot mask, determines that the building's location is within the boundary of the plot to be designed. The remaining buildable capacity is updated based on the difference between the total building area constraint and the building areas of all generated nodes, and a decision is made on whether to continue generating new nodes. The embeddings of the generated nodes are concatenated with the initial condition vector and input back into the decoder. This node generation process is repeated to generate the next node. Simultaneously, the edge prediction module calculates the spatial topological relationship between the new node and all generated nodes. If an edge exists between predicted nodes, its weight is recorded. This process is repeated until a dynamic termination condition is met, at which point the generation of new nodes stops, and the building cluster layout graph network is output.

7. The method for automatic layout and iterative optimization of neighborhood building clusters based on Transformer architecture according to claim 6, characterized in that, The formula for the fitness function in step five is: Fitness=exp(-α·(L+β·Penalty)) L=λ1∣BD in -BD g ∣+λ2∣FAR in -FAR g ∣+λ3∣H max -H g ∣ Where L refers to the loss due to proximity to the indicator, Penalty is the penalty for exceeding the plot boundary or insufficient spacing, α and β are weighting coefficients, and BD in BD is used to calculate the building density of the generated building complex scheme. g For the input building density constraint value, FAR in For the floor area ratio of the generated building complex scheme, FAR g H is the input floor area ratio constraint value. max H represents the maximum building height of the generated building complex design. g This is the maximum building height constraint value that you input. The formula for calculating Penalty is as follows: Penalty=Penalty boundary +Penalty spacing Among them penalty boundary Penalty for exceeding land parcel boundaries spacing x is the penalty for building spacing. min x max y min y max Let x be the boundary coordinates of the land parcel. i y i Let d be the coordinates of the ground center point of building i. min It refers to the minimum spacing requirement between buildings.

8. The method for automatic layout and iterative optimization of neighborhood building clusters based on Transformer architecture according to claim 7, characterized in that, Step five uses an optimization algorithm, which involves using a genetic algorithm to perturb the nodes in a high-dimensional space. After flattening the node matrix into a vector, a tournament selection process is used to select parent schemes with high fitness. Block crossover and Gaussian mutation are applied to these parent schemes, and partial building submatrices are swapped and the height and coordinates are randomly perturbed to generate offspring schemes. Subsequently, the projection operator is used to force the mutated coordinates to be constrained to the plot boundary, and the edge matrix is ​​recalculated to ensure the minimum spacing. After each iteration, the population is updated based on an elite retention strategy, retaining the top 10% of the schemes with the highest fitness and replacing inefficient individuals.

9. The method for automatic layout and iterative optimization of neighborhood building clusters based on Transformer architecture according to claim 8, characterized in that, The simulation equipment mentioned in step six refers to a digital sand table that integrates building complex physical environment performance simulation algorithms, pedestrian flow simulation algorithms, AR augmented reality glasses, and voice and gesture recognition modules. It enables urban designers to finely optimize and adjust the layout plan of the building complex based on the results of the simulation algorithms, and to display the three-dimensional model of the plan and the building functions, height, area indicators, and technical and economic indicators of the building complex through AR augmented reality glasses.