High-strength district thermal environment performance space design decision-making method based on reinforcement learning
Through training the building and tree agents based on reinforcement learning, the buildings and tree layout of high-intensity areas of the city are jointly optimized, and the problems of cumbersome thermal environment optimization process and insufficient automation in the existing technology are solved, and rapid and effective thermal comfort improvement is achieved.
Patent Information
- Application Number
- CN202510476725.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing technology has problems such as cumbersome optimization process, inability to iterate quickly, lack of dynamic response capabilities, and inability to automatically optimize building and tree layout in urban high-intensity areas, resulting in significant impacts on thermal comfort and energy consumption.
Using a reinforcement learning-based method, by defining the state space, action space and reward functions of the building and tree agents, a proximal strategy optimization algorithm is used to train the building and tree agents to jointly generate the optimal layout to meet thermal comfort and floor area ratio constraints.
It realizes efficient and automated optimization of building and tree layout, improves thermal comfort, adapts to different plot shapes and floor area ratios, and improves optimization efficiency and universality.
Smart Images

Figure CN120297150A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of urban thermal comfort evaluation and optimization, and particularly to a spatial design decision-making method for the thermal environment performance of high-intensity areas based on reinforcement learning. Background Art
[0002] With the acceleration of the urbanization process, the heat island effect in urban high-intensity areas (i.e., areas with high building density and large floor area ratio) is becoming increasingly severe. The high-density building layout and limited ventilation conditions have a significant impact on the thermal comfort of residents and building energy consumption. Therefore, reasonable building layout and tree layout are important means to alleviate the heat island effect and play an important role in the optimization of the thermal environment. Traditional urban thermal environment optimization methods mainly rely on manual design or rule-based optimization algorithms to improve the thermal environment through the subjective experience of architects or simple geometric adjustments. However, these methods have the following problems: First, the optimization process is cumbersome and difficult to meet the requirements of rapid iteration in urban planning, especially under the complex spatial and functional constraints of high-intensity areas, and the efficiency problem is particularly prominent; Second, traditional methods lack the dynamic response ability to complex thermal environment indicators (such as UTCI, Universal Thermal Climate Index), and cannot comprehensively consider the comprehensive impact of building layout and tree layout on the microclimate; In addition, existing methods usually fail to achieve efficient automated optimization under random plot and floor area ratio constraints, resulting in insufficient universality of the optimization results.
[0003] In the prior art, some studies have attempted to use numerical simulation tools (such as ENVI-met or Ladybug) to evaluate the thermal environment, but these tools have high computational costs and are difficult to provide real-time feedback on the results of design changes. At the same time, although some machine learning methods have been used for thermal environment prediction, they mainly focus on static prediction rather than dynamic optimization and lack the ability to actively adjust building parameters (such as location, size, angle) and tree parameters (such as location, height, crown diameter). Therefore, how to quickly optimize the building layout and tree layout to improve thermal comfort through an efficient automated method under given plot shape and floor area ratio conditions has become a technical problem to be solved urgently. Summary of the Invention
[0004] In view of the above deficiencies of the prior art, the purpose of the present invention is to provide a spatial design decision-making method for the thermal environment performance of high-intensity areas based on reinforcement learning, aiming to dynamically adjust the building layout and tree layout through the reinforcement learning algorithm, generate building arrangements and tree arrangement plans that meet the thermal comfort requirements under random plot shape and floor area ratio constraints, quickly optimize the thermal comfort indicators of the urban thermal environment, and provide scientific support for urban planning and the optimization of the built environment.
[0005] The technical solution of the present invention is as follows:
[0006] A decision-making method for the spatial design of the thermal environment performance in a high-intensity area based on reinforcement learning, the method comprising the following steps:
[0007] Define the state space, action space, and reward function of the building Agent and the state space, action space, and reward function of the tree Agent respectively;
[0008] Train the building Agent using the proximal policy optimization algorithm, and train the tree Agent using the proximal policy optimization algorithm based on the training results of the building Agent;
[0009] Apply the trained building Agent and the trained tree Agent to jointly generate the optimal layout of buildings and trees in the area to be optimized.
[0010] Further, according to the decision-making method for the spatial design of the thermal environment performance in a high-intensity area based on reinforcement learning, the state space of the building Agent represents the input information of the building Agent, including the geometric feature 1 of the plot, the target floor area ratio FAR {target} , the geometric parameters of the building, and the thermal environment state 1; the geometric feature 1 of the plot includes: the coordinates of each vertex of the plot boundary, the plot area; the geometric parameters of the building include: the center point coordinates (x i , y i ) of the bottom surface of each building, the length l i , the width w i , the building height h i , and the orientation angle θ i . If n buildings have been placed in the plot, the geometric parameters of the building are represented as [x1, y1, l1, w1, h1, θ1,..., x n , y n , l n , w n , h n , θ n ; the thermal environment state 1 includes the following parameters: comfort area ratio Comfort {ratio} , building overlap rate ArchOverlap {ratio} , current floor area ratio FAR {current} , maximum UTCI value UTCI max , building boundary violation index Boundary {out} 1, building size violation index Size {out} 1;
[0011] The state space of the tree Agent represents the input information of the tree Agent, including the geometric feature 2 of the plot, the target tree coverage rate Cover {current}, the geometric parameters of the trees and the thermal environment state 2; the geometric features 2 of the plot include: the coordinates of each vertex of the plot boundary, the floor area of the buildings already placed in the plot, and the geometric parameters of the buildings; the geometric parameters of the trees include the coordinates (x j , y j ) of each tree, the height h j and the crown diameter d j . If there are m trees already placed in the plot, the geometric parameters of the trees are expressed as [x1, y1, h1, d1,..., x m , y m , h m , d m ; the thermal environment state 2 includes the following parameters: comfort area ratio Comfort {ratio} , tree overlap rate TreeOverlap {ratio} , current tree coverage rate Cover {current} , maximum UTCI value UTCI max , tree boundary violation index Boundary {out} 2, tree size violation index Size {out} 2.
[0012] Furthermore, according to the above-mentioned method for making decisions on the spatial design of the thermal environment performance of high-intensity areas based on reinforcement learning, the parameters of the thermal environment state 1 and the thermal environment state 2 are all obtained through UTCI simulation calculations:
[0013] (1) Comfort area ratio Comfort {ratio} , which represents the ratio of the area within the grid range corresponding to UTCI from 9 to 26 to the area of the plot:
[0014]
[0015] In the formula, K is the total number of grids; UTCI k is the UTCI value of the kth grid; I(9 ≤ UTCI k < 26) is an indicator function. If UTCI k is between 9 and 26, then I(9 ≤ UTCI k < 26) is 1, otherwise I(9 ≤ UTCI k < 26) is 0; A grid is the area of a single grid, and A plot is the total area of the plot;
[0016] (2) Building overlap rate ArchOverlap {ratio} :
[0017]
[0018] In the formula, AOverlap,i,j Denote the overlapping bottom area between the \(i\)-th building and the \(j\)-th building; \(A\) i , \(A\) j are the bottom area of the \(i\)-th building and the bottom area of the \(j\)-th building respectively; \(I\) Overlap,i is an indicator function. If Ratio exceeds 0.6, then \(I\) Overlap,i = 1, otherwise \(I\) Overlap,i = 0; \(N\) is the total number of buildings within the current plot;
[0019] (3) Current floor area ratio FAR {current} :
[0020]
[0021] In the formula, \(F\) i is the storey height; is the building area of the \(i\)-th building; \(A\) plot is the site area;
[0022] (4) Maximum UTCI value UTCI max , the highest UTCI value in the grid:
[0023]
[0024] In the formula, UTCI k is the UTCI value of the \(k\)-th grid obtained through UTCI simulation;
[0025] (5) Building boundary violation index Boundary {out} 1. Statistic the ratio of the number of buildings exceeding the plot boundary line to the total number:
[0026]
[0027] In the formula, \(I\) boundary,i is an exponential function;
[0028] (6) Building size violation index Size {out} 1. Statistic the ratio of the number of buildings exceeding the size limit to the total number:
[0029]
[0030] In the formula, \(I\) size,i is an exponential function;
[0031] (7) Tree overlap rate TreeOverlap {ratio} :
[0032]
[0033] In the formula, for the newly placed \(i\)-th tree, check whether its crown overlaps with the existing crowns; \(A\) Overlap,i,j represents the overlapping bottom area between the \(i\)-th tree and the \(j\)-th tree; \(A\) i , \(A\) j are respectively the bottom areas of the \(i\)-th tree and the \(j\)-th tree; \(I\) Overlap,i is an indicator function. If Ratio exceeds 0.6, then \(I\) Overlap,i = 1, otherwise \(I\) Overlap,i = 0; \(A\) Overlap,i,j is the overlapping area between the \(i\)-th tree and the \(j\)-th tree; \(M\) is the total number of trees in the current plot.
[0034] (8) Current tree coverage rate Cover {current} :
[0035]
[0036] In the formula, \(d\) i is the crown diameter of the \(j\)-th tree; \(l\) i · \(w\) i is the building floor area of the \(i\)-th building, \(M\) is the total number of trees in the current plot, and \(N\) is the total number of buildings in the current plot;
[0037] (9) Tree boundary violation index Boundary {out} 2. Statistically, the ratio of the number of trees exceeding the plot boundary line to the total number:
[0038]
[0039] In the formula, \(I\) boundary,i is an exponential function. If the \(i\)-th tree exceeds the plot boundary, then \(I\) boundary,i = 1, otherwise \(I\) boundary,i = 0;
[0040] (10) Tree size violation index Size {out} 2. Statistically, the ratio of the number of trees exceeding the size limit to the total number:
[0041]
[0042] In the formula, \(I\) size,i is an exponential function.
[0043] Furthermore, according to the above-mentioned method for spatial design decision of the thermal environment performance in high-intensity areas based on reinforcement learning, the action space of the building Agent represents the output actions of the building Agent, including placing new buildings and adjusting existing buildings; placing new buildings refers to placing new buildings within the plot; adjusting existing buildings refers to specifying the adjustment amount for the \(i\)-th existing building: [Δx i , Δyi , Δl i , Δw i , Δh i , Δθ i , where Δx i , Δy i is the change in the coordinates of the center point of the building's base; Δl i , Δw i are the changes in the length and width of the building's base; Δθ i is the change in the building's orientation angle;
[0044] The action space of the tree Agent represents the output actions of the tree Agent, including placing new trees and adjusting existing trees; placing new trees refers to placing new trees within the plot; adjusting existing trees specifies the adjustment amount for the j-th existing tree: [Δx j , Δy j , Δh j , Δd j , where Δx j , Δy j are the coordinate changes of the tree; Δh j is the height change of the tree; Δd j is the change in the crown diameter of the tree.
[0045] Further, according to the above-mentioned decision-making method for the thermal environment performance space design of high-intensity areas based on reinforcement learning, the reward function r of the building Agent arch is:
[0046]
[0047]
[0048] In the formula, ω1·Comfort {ratio} encourages maximizing the comfort area ratio; ω2·ArchOverlap {ratio} punishes the situation where the overlap between the new building and the existing building exceeds 60%; punishes the floor area ratio deviation; ω4· max (0, UTCI max - 26) punishes the maximum value of UTCI, UTCI max exceeding 26; ω5·Boundary {out}1 punishes the building for exceeding the boundary; ω6·Size {out} 1 punishes the size for exceeding the limit; ω are the weights of different rewards;
[0049] The reward function r of the tree Agent tree is:
[0050]
[0051] where τ1·Comfort {ratio} encourages maximizing the comfort area ratio; τ2·TreeOverlap {ratio} penalizes the case where the overlap of new trees with existing trees exceeds 60%; penalizes the deviation of the tree coverage rate;
[0052] τ4· max (0, UTCI max - 26) penalizes the maximum value of UTCI, UTCI max exceeding 26; τ5·Boundary {out} 2 penalizes trees exceeding the boundary; τ6·Size {out} 2 penalizes trees with over - sized dimensions; τ is the weight of different rewards.
[0053] Furthermore, according to the above - mentioned decision - making method for the thermal environment performance space design of high - intensity areas based on reinforcement learning, the training of the building Agent using the Proximal Policy Optimization algorithm includes the following steps:
[0054] S2.1.1: Conduct a thermal environment simulation on the plot: Randomly generate a batch of empty plots, extract the geometric feature 1 of each plot and assign the target plot ratio to each plot, then generate a site status layout diagram representing the layout positions of buildings that can be arranged within the plot based on the geometric feature 1 of the plot, and conduct a UTCI simulation on each plot based on historical meteorological data to obtain the initial thermal environment state 1;
[0055] S2.1.2: Based on the state space s of the current building Agent arch and the site status layout diagram, generate the output action distribution parameters of the building Agent through the Actor network to obtain the current policy π θ (a arch |s arch );
[0056] S2.1.3: Based on the reward function of the building Agent, predict the state value V φ (s arch ) and calculate the advantage A of the action arch ;
[0057] S2.1.4: Based on the advantage A arch , the state value V φ (s arch ) and the policy π θ (a arch |s arch ), calculate the total loss L through the loss function;
[0058] S2.1.5: Update the Actor network parameters θ according to the total loss L using the Adam optimizer arch and the Critic network parameters φ arch ;
[0059] S2.1.6: Repeat steps S2.1.2 to S2.1.5 to iteratively update θ arch and φ arch until the preset maximum number of training rounds is reached, thereby obtaining the final parameters of the Actor network and the final parameters of the Critic network of the building Agent.
[0060] Furthermore, according to the above - mentioned method for high - intensity area thermal environment performance space design decision - making based on reinforcement learning, training the tree Agent using the Proximal Policy Optimization algorithm includes the following steps:
[0061] S2.2.1: Conduct a thermal environment simulation on the plot: Extract the geometric features 2 of the plot and the geometric parameters of the buildings after the training of the building Agent is completed, and specify the target greening rate. Then, generate a site current layout map representing the positions where trees can be arranged within the plot based on the geometric features 2 of the plot and the geometric parameters of the buildings, and conduct UTCI simulation on each plot based on historical meteorological data to obtain the initial thermal environment state 2;
[0062] S2.2.2: Based on the current state space s of the tree Agent tree and the site current layout map, generate the output action distribution parameters of the tree Agent through the Actor network to obtain the current policy π θ (a tree |s tree );
[0063] S2.2.3: Based on the reward function of the tree Agent, predict the state value V φ (s tree ) through the Critic network and calculate the advantage A of the action tree ;
[0064] S2.2.4: Based on the advantage A tree , the state value V φ (s tree ) and the policy π θ (a tree |s tree ), calculate the total loss L' through the loss function;
[0065] S2.2.5: Update the Actor network parameters θ tree and the Critic network parameters φ tree using the Adam optimizer according to the total loss L';
[0066] S2.2.6: Repeat S2.2.2 to S2.2.5 to iteratively update θ tree and φ tree until the preset maximum number of training epochs is reached, thereby obtaining the final parameters of the Actor network and the final parameters of the Critic network of the building Agent.
[0067] Furthermore, according to the above-mentioned decision-making method for the spatial design of the thermal environment performance of high-intensity areas based on reinforcement learning, the application of the trained building Agent and the trained tree Agent to collaboratively generate the optimal layout of buildings and trees in the area to be designed and decided includes the following steps:
[0068] S3.1: Obtain the state space s of the area to be optimized, i.e., the plot, according to the method of S2.1.1 arch , including the geometric feature 1 of the plot, the thermal environment state 1, and the current site layout map; set the target floor area ratio FAR {target} of this area; load the Actor network parameters θ of the trained building Agent arch ;
[0069] S3.2: Obtain the state space s of the area to be optimized, i.e., the plot, according to the method of S2.2.1 tree , including the geometric feature 2 of the plot, the thermal environment state 2, and the current site layout map; set the target tree coverage rate Cover {target} of this area; load the Actor network parameters θ of the trained tree Agent tree ;
[0070] S3.3: Take the current geometric feature 1 of the plot, the thermal environment state 1, and the current site layout map as inputs, and apply the trained building Agent to optimize the building layout of the area with the constraints of meeting the target floor area ratio and the preliminary thermal environment requirements, and output the current optimal building layout of this area, as well as the updated geometric feature 1 and thermal environment state 1;
[0071] S3.4: Take the current geometric feature 2 of the plot, the thermal environment state 2, and the current site layout map as inputs, and apply the trained tree Agent to optimize the tree layout of the area with the constraints of meeting the target tree coverage rate and the preliminary thermal environment requirements on the basis of the building layout generated in S3.3, and output the current optimal tree layout of this area, as well as the updated geometric feature 2 and thermal environment state 2;
[0072] S3.5: Repeat S3.3 to S3.4 until the preset termination condition is met;
[0073] S3.6: Output the final optimized layout of the area to be optimized, i.e., the plot, and its thermal environment status: According to the target floor area ratio and the indicators in the reward function of the building Agent, output the geometric feature 1 and the thermal environment status 1 corresponding to the best building layout scheme among the indicators, and according to the target tree coverage rate and the indicators in the reward function of the tree Agent, output the geometric feature 2 and the thermal environment status 2 corresponding to the best tree layout scheme among the indicators.
[0074] Furthermore, according to the above-mentioned method for spatial design decision-making of the thermal environment performance of high-intensity areas based on reinforcement learning, the current site layout map is a grayscale image of 256×256 pixels, and the height of buildings and / or trees is mapped by the depth of color.
[0075] Compared with the prior art, the present invention has the following beneficial effects:
[0076] (1) Efficient and automated optimization: Based on the fast decision-making mechanism of reinforcement learning, intelligent adaptive adjustment, and full utilization of modern computing power for parallel computing to dynamically adjust the building layout and tree layout, the efficiency is increased by dozens of times compared with traditional manual design or numerical simulation methods, meeting the requirements of rapid iteration of urban planning.
[0077] (2) Comprehensive optimization: The method of the present invention simultaneously satisfies multiple constraints such as plot boundaries, floor area ratio, overlap rate, and thermal comfort, and combines the optimization of tree layout, fully considering the regulating effect of greening on the thermal environment, ensuring the practicability and feasibility of the optimization results.
[0078] (3) Strong universality: Using random plots for algorithm training enables the method of the present invention to adapt to plots of different shapes and floor area ratio conditions, with strong generalization ability. Description of the Drawings
[0079] Figure 1 It is the flow chart of the method for spatial design decision-making of the thermal environment performance of high-intensity areas based on reinforcement learning in this embodiment;
[0080] Figure 2 It is the structural schematic diagram of the proximal policy optimization algorithm;
[0081] Figure 3 It is the structural diagram of the Actor network in the building Agent in this embodiment;
[0082] Figure 4 It is the structural diagram of the Critic network in the building Agent in this embodiment. Specific Embodiments
[0083] To facilitate the understanding of this application, the following will describe this application more comprehensively with reference to the relevant drawings.
[0084] Figure 1 It is a flowchart of the decision-making method for the thermal environment performance space design of high-intensity areas based on reinforcement learning in this embodiment.
[0085] As Figure 1 shown, the decision-making method for the thermal environment performance space design of high-intensity areas based on reinforcement learning includes the following steps:
[0086] S1: Define the state space, action space, and reward function of the building Agent and the state space, action space, and reward function of the tree Agent respectively;
[0087] S1.1: Define the state space of the building Agent and the state space of the tree Agent respectively;
[0088] The state space is a set of features used to describe the current state of the environment in reinforcement learning, representing the input information of the agent.
[0089] In this embodiment, the state space s of the building Agent arch represents the input information of the building Agent, including the geometric feature 1 of the plot, the target floor area ratio FAR {target} of the plot, the geometric parameters of the building, and the thermal environment state 1; the geometric feature 1 of the plot includes: the vertex coordinates of each boundary of the plot, the plot area; the target floor area ratio FAR {target} is randomly obtained within the design specification range and is used to constrain the total floor area of the building; the geometric parameters of the building include: the center point coordinates (x i , y i ) of the bottom surface of each building, the length l i , the width w i , the building height h i , and the orientation angle θ i . If there are n buildings placed in the plot, the geometric parameters of the building are expressed as [x1, y1, l1, w1, h1, θ1,..., x n , y n , l n , w n , h n , θ n ; the thermal environment state 1 includes the following parameters: comfort area ratio Comfort {ratio} , building overlap rate ArchOverlap {ratio} , current floor area ratio FAR {current} , maximum UTCI value UTCI max , building boundary violation index Boundary {out} 1, building size violation index Size {out}1. In this embodiment, the parameters of the thermal environment state 1 are all obtained through UTCI simulation calculations, specifically as follows:
[0090] (1) Comfort area ratio Comfort {ratio} , representing the ratio of the area within the grid range corresponding to UTCI from 9 to 26 to the plot area:
[0091]
[0092] In the formula, K is the total number of grids; UTCI k is the UTCI value of the kth grid; I(9 ≤ UTCI k < 26) is an indicator function. If UTCI k is between 9 and 26, then I(9 ≤ UTCI k < 26) is 1; otherwise, I(9 ≤ UTCI k < 26) is 0; A grid is the area of a single grid, and A plot is the total area of the plot.
[0093] (2) Building overlap ratio ArchOverlap {ratio} :
[0094]
[0095] In the formula, for the newly placed ith building in the plot, check whether its base area overlaps with the existing buildings (the 1st to i - 1th) in the plot; in this embodiment, Ratio is used to calculate the ratio of the overlapping area to the base area of the overlapping existing buildings; A Overlap,i,j represents the overlapping base area between the ith building and the jth building; A i , A j are the base areas of the ith building and the jth building respectively; I Overlap,i is an indicator function. If Ratio exceeds 0.6, then I Overlap,i = 1; otherwise, I Overlap,i = 0; N is the total number of buildings in the current plot.
[0096] (3) Current floor area ratio FAR {current} :
[0097]
[0098] In the formula, F i is the floor height, which is taken as F i = 3 in this embodiment; is the floor area of the ith building; A plot is the site area.
[0099] (4) Maximum UTCI value UTCImax , the highest UTCI value in the grid:
[0100]
[0101] where UTCI k is the UTCI value of the k-th grid obtained by performing UTCI simulation using Ladybug.
[0102] (5) Building boundary violation index Boundary {out} 1. Calculate the ratio of the number of buildings exceeding the plot boundary line to the total number:
[0103]
[0104] where I boundary,i is an exponential function. If the i-th building exceeds the plot boundary, then I boundary,i = 1; otherwise, I boundary,i = 0; N is the number of buildings within the current plot.
[0105] (6) Building size violation index Size {out} 1. Calculate the ratio of the number of buildings exceeding the size limit to the total number:
[0106]
[0107] where I size,i is an exponential function. If the length l of the i-th building i > 20m or the width w i > 20m, then I size,i = 1; otherwise, I size,i = 0.
[0108] In this embodiment, the state space s of the tree Agent tree represents the input information of the tree Agent, including the geometric features of the plot 2, the target tree coverage rate Cover {current} , the geometric parameters of the trees, and the thermal environment state 2; the geometric features of the plot 2 include: the coordinates of each vertex of the plot boundary, the floor area of the buildings already placed within the plot, and the geometric parameters of the buildings; the geometric parameters of the trees include the coordinates (x j , y j ) of each tree, the height h j and the crown diameter d j . If there are m trees already placed within the plot, the geometric parameters of the trees are represented as [x1, y1, h1, d1,..., x m , y m , h m , d m; The thermal environment state 2 includes the following parameters: Comfortable area ratio Comfort {ratio} , Tree overlap rate TreeOverlap {ratio} , Current tree coverage rate Cover {current} , Maximum value of UTCI UTCI max , Tree boundary violation index Boundary {out} 2. Tree size violation index Size {out} 2. In this embodiment, the parameters of the thermal environment state 2 are also obtained through UTCI simulation calculation. The formulas for the comfortable area ratio Comfort {ratio} and the maximum value of UTCI UTCI max have been given above. The calculation formulas for the other 4 parameters are as follows:
[0109] I. Tree overlap rate TreeOverlap {ratio} :
[0110]
[0111] In the formula, for the newly placed i-th tree, check whether its tree crown overlaps with the existing tree crowns (the 1st to i-1th); in this embodiment, the ratio of the overlapping bottom area to the bottom area of the existing overlapping trees is calculated through Ratio; A Overlap,i,j represents the overlapping bottom area between the i-th tree and the j-th tree; A i , A j are the bottom areas of the i-th tree and the j-th tree respectively; I Overlap,i is an indicator function. If Ratio exceeds 0.6, then I Overlap,i =1, otherwise I Overlap,i =0; A Overlap,i,j is the overlapping area between the i-th tree and the j-th tree; M is the total number of trees in the current plot.
[0112] II. Current tree coverage rate Cover {current} :
[0113]
[0114] In the formula, d i is the crown diameter of the j-th tree; l i ·w i is the building floor area of the i-th building, M is the total number of trees in the current plot, and N is the total number of buildings in the current plot.
[0115] III. Tree boundary violation index Boundary {out} 2. Statistically calculate the ratio of the number of trees exceeding the plot boundary line to the total number:
[0116]
[0117] In the formula, I boundary,i is an exponential function. If the i-th tree exceeds the plot boundary, then I boundary,i = 1; otherwise I boundary,i = 0; M is the number of trees within the current plot.
[0118] IV. Tree Size Violation Index Size {out} 2. Calculate the ratio of the number of trees exceeding the size limit to the total number:
[0119]
[0120] In the formula, I size,i is an exponential function. If the crown diameter d i of the i-th tree is < 2m or d i > 10m, or the tree height h i < 3m or h i > 30m, then I size,i = 1; otherwise I size,i = 0.
[0121] S1.2: Define the action space a arch of the building Agent and the action space a tree of the tree Agent respectively;
[0122] The action space is the set of operations that an Agent can perform in reinforcement learning, representing the output actions of the Agent. In this embodiment, the action space a arch of the building Agent represents the output actions of the building Agent, including two types of actions: placing a new building and adjusting an existing building. Placing a new building specifically refers to placing the n-th new building within the plot, and the geometric parameters of the building are [x n , y n , l n , w n , h n , θ n . The center point coordinates (x n , y n ) of the building's bottom surface are restricted within the plot boundary; the length l n and width w n of the building's bottom surface are set according to the actual situation. In this embodiment, they are in the range of 5m to 20m; the orientation angle of the building is in the range of 0° to 180°. Adjusting an existing building refers to specifying the adjustment amount for the i-th existing building: [Δx i , Δy i , Δl i , Δw i, Δh i , Δθ i , where Δx i , Δy i is the change in the coordinates of the center point of the building's base, with a step size of L1, and it remains within the plot after adjustment; Δl i , Δw i is the change in the length and width of the building's base, with a step size of L2. In this embodiment, after adjustment, l i , w i is in the range of 5 m to 20 m; Δθ i is the change in the building's orientation angle, with a step size of L3. After adjustment, θ i is in the range of 0° to 180°;
[0123] In this embodiment, the action space of the tree Agent represents the output actions of the tree Agent, including placing new trees and adjusting existing trees. Placing a new tree specifically means placing the m-th new tree within the plot, and the geometric parameters of the tree are [x m , y m , h m , d m . The coordinates (x m , y m ) of the tree are restricted within the plot boundary and do not coincide with the building coverage area; the height h m of the tree is set according to the actual situation. In this embodiment, it is in the range of 3 m to 30 m; the crown diameter d m of the tree is set according to the actual situation. In this embodiment, it is in the range of 2 m to 10 m. Adjusting an existing tree means specifying an adjustment amount for the j-th existing tree: [Δx j , Δy j , Δh j , Δd j . Among them, Δx j , Δy j is the change in the coordinates of the tree, with a step size of L4, and it remains within the plot and does not coincide with the building coverage area after adjustment; Δh j is the change in the height of the tree, with a step size of L5. After adjustment, h j is in the range of 3 m to 30 m; Δd j is the change in the crown diameter of the tree, with a step size of L6. After adjustment, d j is in the range of 2 m to 10 m.
[0124] S1.3: Define the reward function r arch of the building Agent and the reward function r tree of the tree Agent respectively;
[0125] The Reward Function is a metric used in reinforcement learning to evaluate the effects of an Agent's actions, guiding the Agent to optimize its goals.
[0126] The reward function r of the building Agent designed in this embodiment arch is:
[0127]
[0128] In the formula, ω1·Comfort {ratio} encourages maximizing the comfort area ratio; ω2·ArchOverlap {ratio} penalizes the case where the overlap between the new building and the existing building exceeds 60%; penalizes the floor area ratio deviation; ω4·max(0,UTCI max -26) penalizes the maximum value of UTCI, UTCI max exceeding 26; ω5·Boundary {out} 1 penalizes the building exceeding the boundary; ω6·Size {out} 1 penalizes the size exceeding the limit; ω are the weights of different rewards.
[0129] The reward function r of the tree Agent designed in this embodiment tree is:
[0130]
[0131] In the formula, τ1·Comfort {ratio} encourages maximizing the comfort area ratio; τ2·TreeOverlap {ratio} penalizes the case where the overlap between the new tree and the existing tree exceeds 60%; penalizes the deviation of the tree coverage rate;
[0132] τ4·max(0,UTCI max -26) penalizes the maximum value of UTCI, UTCI max exceeding 26; τ5·Boundary {out} 2 penalizes the tree exceeding the boundary; τ6·Size {out} 2 penalizes the tree size exceeding the limit; τ are the weights of different rewards.
[0133] S2: Use the Proximal Policy Optimization algorithm to train the building Agent and the tree Agent;
[0134] S2.1: Use the Proximal Policy Optimization algorithm to train the building Agent;
[0135] Figure 2It is a schematic diagram of the Proximal Policy Optimization algorithm. The Proximal Policy Optimization (PPO) algorithm is based on the Actor-Critic framework. It uses the policy network (Actor) to select actions and the value network (Critic) to evaluate the action values, and realizes the stability control of policy update through the dual-network collaborative optimization mechanism. The PPO algorithm ensures the stability of the training process by controlling the amplitude of policy update, and finally generates the optimal building layout plan. The following are the specific steps for training the building Agent using the PPO algorithm, including all the parameters and operation processes that need to be set.
[0136] S2.1.1: Conduct a thermal environment simulation for the plot: Randomly generate a batch of empty plots, extract the geometric feature 1 of each plot and assign the target floor area ratio for each plot, and then generate the current site layout map based on the geometric feature 1 of the plot and conduct a UTCI simulation for each plot to obtain the initial thermal environment state 1.
[0137] In this embodiment, first, Rhino software is used to randomly generate a batch of quadrilateral empty plots, and the geometric feature 1 of each plot is extracted. Then, based on the geometric feature 1 of the plot, a layout map of the current site is generated, indicating the positions where buildings can be laid out within the plot. Next, a 256×256 pixel grayscale image is generated using Rhino, and the height of the building is mapped through the shade of color.
[0138] Then, a UTCI simulation is conducted on the generated empty plots based on historical meteorological data to obtain the initial thermal environment state 1;
[0139] In this embodiment, historical meteorological data is mainly provided by the local online meteorological database in the research area, and the required meteorological data is exported in the ".epw" format to ensure that it can reflect the true meteorological conditions in the research area. The Ladybug plugin in the Grasshopper parametric platform of Rhino is used for UTCI simulation. The grid size for simulation calculation is set to 5 meters × 5 meters, and the monitoring height is 1.8 meters to carefully capture the urban microclimate changes. All input parameters are set according to the annual conventional climate data standard in the research area to ensure the accuracy of the simulation results.
[0140] Finally, the state space s of the plot arch consists of the geometric feature 1 of the plot, the target floor area ratio FAR of the plot {target} , the geometric parameters of the building, and the thermal environment state 1, which is used as the input for subsequent Agent training.
[0141] S2.1.2: Based on the state space s of the current building Agent archand the current layout diagram of the site. Generate the output action distribution parameters of the building agent through the Actor network to obtain the current policy π θ (a arch |s arch );
[0142] Figure 3 This is the structure diagram of the Actor network in the building agent of this embodiment. Initialize the Actor network parameters θ of the building agent arch and input the state space s of the building agent generated in S2.1.1 arch , including the geometric feature 1 of the plot, the target floor area ratio FAR {target} of the plot, the geometric parameters of the building, and the thermal environment state 1. Initialize the weights of θ arch using the Xavier method (the Xavier method is a common initialization method to ensure that the initial output of the network is neither too large nor too small), and initialize the bias to 0. The input layer of the Actor network receives the plot state s arch . The current layout diagram of the site is processed using a convolutional neural network (CNN, Convolutional Neural Network) with a dimension of 256×256×1. Then, spatial features are extracted through three convolutional layers. The size of each convolutional kernel is 3×3, the stride is 2, and the number of output channels is 16, 32, and 64 respectively. Each layer is followed by a ReLU activation function. Finally, through a fully connected layer, the plot image feature 1×256 is output; the geometric parameters of the building are processed using a long short-term memory network (LSTM, Long Short-Term Memory), with 128 hidden units set, and the 6-dimensional parameters of each building are processed step by step in time, and the hidden state T of the last time step is output n as 1×128; the thermal environment state 1 (dimension 1×6, including [Comfort {ratio} ,ArchOverlap {ratio} ,FAR {current} ,UTCI max ,Boundary {out} 1,Size {out} 1]), through a fully connected layer, outputs a dimension of 1×64. Then, the current layout diagram of the site (1×256), the LSTM hidden state T n (1×128) and the thermal environment state 1 (1×64) are concatenated, and the output dimension is 1×448. Then, the feature vector enters two fully connected layers, with an input dimension of 1×448 and an output dimension of 1×128, and is activated through ReLU. Finally, the output layer of the network generates the action distribution parameter π θ (a arch |s arch) where the mean μ represents the expected value of the action ("the most likely action"), maps the input 1×128 to the output 1×6 through a fully connected layer, and scales it to the action range through the tanh activation function; the standard deviation σ represents the variance of the action distribution ("the exploration range of the action"), also maps the input 1×128 to the output 1×6 through a fully connected layer, and ensures a positive value through the softplus activation function. The Actor network outputs the action distribution parameter π θ (a arch |s arch ) where π θ (a arch |s arch ) is the current policy, defined as the probability distribution of the action a at the state s at time t, samples the action a from the normal distribution N(μ, σ). arch When, the action a arch is passed to the environment, and the environment updates the plot layout, that is, adds new buildings or moves existing buildings on the basis of the current site layout map, and re - conducts the UTCI simulation (grid size is 5m×5m, monitoring height is 1.8m), calculates the reward r arch . And generates the next state s arch (including the updated geometric feature 1, thermal environment state 1 and the current site layout map), and saves (s arch , a arch+1 , r arch , s arch ) to the experience storage module, and repeats this process until the maximum decision - making step number N is collected or the preset termination condition is met. arch , s arch+1 )
[0143] S2.1.3: Based on the reward function of the building Agent, the Critic network predicts the state value V φ (s arch ) and calculates the advantage A of the action arch ;
[0144] Figure 4 is the structure diagram of the Critic network in the building Agent of this embodiment. Initialize the Critic network parameters φ arch of the building Agent, and use the stored experience data. The Critic network predicts the state value and calculates the advantage of the action. Traverse each piece of data (s arch , a arch , r arch , s arch+1 ) in the experience storage module, and take the states s arch and s arch+1Input the Critic network. The structure of the Critic network is the same as that of the Actor network, including a multi-branch structure. After being processed by the fusion layer, a feature vector with a dimension of 1×448 is obtained. Then, the feature vector enters two fully connected layers with an input dimension of 1×448 and an output dimension of 1×128, which is activated by ReLU. Finally, through a fully connected layer, the input 1×128 is mapped to the output 1×1 to generate the state value V φ (s arch )), where V φ (s arch ) is the value predicted by the Critic network, representing the expected value of the future discounted rewards that can be obtained starting from the state s arch according to the current policy. Similarly, calculate V φ (s arch+1 ).
[0145] Calculate the cumulative value R arch , which is the training objective of the Critic network, representing the current reward plus the discounted value of the next state. The formula is as follows:
[0146] R arch = r arch + γV φ (r arch+1 )
[0147] Among them, γ is the discount factor (indicating the importance of future rewards. The closer the value is to 1, the more the Agent focuses on long-term benefits). In this embodiment, γ = 0.99.
[0148] Calculate the temporal difference error δ t , which represents the error predicted by the Critic network and is used to calculate the advantage. The formula is as follows:
[0149] δ t = r arch + γV φ (s arch+1 ) - V φ (s arch )
[0150] Among them, δ t is the temporal difference error (TD error), representing the difference between the predicted value and the actual return at time step t.
[0151] Calculate the advantage A arch , which represents the relative goodness or badness of an action. The larger the value, the better the action, and it is used to guide the Actor network to improve the policy. The formula is as follows:
[0152]
[0153] Among them, A archis the advantage function, indicating the level above or below the average when choosing action a arch under state s arch ; λ is the GAE parameter used to calculate the advantage of the action. In this embodiment, λ = 0.95; q is a non - negative integer (q = 0, 1, 2, …, ∞), representing the offset from the current time step t to the q - th step in the future.
[0154] S2.1.4: Based on the advantage A arch , the state value V φ (s arch ) and the policy π θ (a arch |s arch ), calculate the total loss L through the loss function;
[0155] Using the advantage A arch , the state value V φ (s arch ) and the policy π θ (a arch |s arch ), calculate the loss function to prepare for network update. Calculate the policy gradient target L CLIP , which is used to measure the improvement direction of the policy and ensure stable update. The formula is as follows:
[0156]
[0157] where, π θ (a arch |s arch ) is the current policy; is the old policy; r arch (θ) is the probability ratio of the new and old policies, representing the change amplitude of the current policy; ∈ is the clipping parameter, and in this embodiment, ∈ = 0.2.
[0158] Calculate the value loss L VF , which is used to optimize the Critic network to make it more accurately predict the state value. The formula is as follows:
[0159]
[0160] where, represents the expectation, R arch is the cumulative value, V φ (s arch ) is the value predicted by the Critic.
[0161] Calculate the entropy loss S[π θ (s arch), representing the randomness of the policy (the larger the entropy value, the more random the policy, and the more inclined the Agent is to explore. The combined total loss L incorporates policy improvement, value prediction, and exploration, and is used to guide the update of network parameters. The formula is as follows:
[0162] L = L CLIP - c1L VF + c2S[πθ](s arch )
[0163] Among them, c1 is the value loss weight (the larger this value, the greater the training intensity of the Critic network); c2 is the entropy loss weight (the larger this value, the stronger the exploration of the Agent).
[0164] S2.1.5: According to the total loss L, use the Adam optimizer to update the Actor network parameter θ arch and the Critic network parameter θ arch ;
[0165] In this embodiment, the Adam optimizer is used to update the Actor network parameter θ arch . Because θ arch determines the policy of the Actor network, after the update, the construction Agent can select better actions. The formula is as follows:
[0166]
[0167] Among them, η is the learning rate, and in this embodiment, η = 0.0003; is the gradient of the loss with respect to θ arch , reflecting the direction and magnitude of the adjustment of the parameter θ arch .
[0168] Update the Critic network parameter φ arch . Because φ arch determines the value prediction of the Critic network, after the update, the Critic can more accurately evaluate the actions. The formula is as follows:
[0169]
[0170] Among them, is the gradient of the loss with respect to φ arch , reflecting the direction and magnitude of the adjustment of the parameter φ arch .
[0171] After updating the network parameters, clear the experience storage module and prepare for the next round of experience collection.
[0172] S2.1.6: Repeat the training until the maximum number of rounds is reached
[0173] Repeat steps S2.1.2 to S2.1.5 to iteratively update θ arch and θ arch until the preset maximum number of training rounds (set to 1000 rounds in this embodiment) is reached. The finally obtained θ arch and φ arch are the final parameters of the Actor and Critic networks of the building Agent. The θ of the Actor network arch corresponds to the optimal policy π θ which is the optimal building layout strategy that meets the requirements of the target plot ratio and thermal environment. After training is completed, save the Actor network parameter θ arch for subsequent generation of building layouts.
[0174] S2.2: Train the tree Agent using the Proximal Policy Optimization algorithm
[0175] The Proximal Policy Optimization (PPO) algorithm is based on the Actor-Critic framework. It uses the policy network (Actor) to select actions and the value network (Critic) to evaluate the action values, and realizes the stability control of policy update through a dual-network collaborative optimization mechanism. The PPO algorithm ensures the stability of the training process by controlling the amplitude of policy update, and finally generates the optimal tree layout plan. The following are the specific steps for training the tree Agent using the PPO algorithm, including all the parameters to be set and the operation process.
[0176] S2.2.1: Conduct a thermal environment simulation on the plot: Extract the geometric features 2 of the plot and the geometric parameters of the buildings after optimization in S2.1, and specify the target greening rate. Then, based on the geometric features 2 of the plot and the geometric parameters of the buildings, generate a layout map of the current site situation and conduct a UTCI simulation for each plot.
[0177] In this embodiment, first, use Rhino software to model according to the geometric features 2 of the plot and the geometric parameters of the buildings in S2.2. Then, based on the geometric features 2 of the plot and the geometric parameters of the buildings, generate a layout map of the current site situation, indicating the positions where trees can be laid out within the plot. Then use Rhino to generate a grayscale image of 256×256 pixels, and map the height of the trees through the depth of the color.
[0178] Then conduct a UTCI simulation on the generated empty plots based on historical meteorological data to obtain the initial thermal environment state 2;
[0179] In this embodiment, historical meteorological data is mainly provided by the local online meteorological database in the study area, and the required meteorological data is exported in the ".epw" format to ensure that it can reflect the true meteorological conditions in the study area. The Ladybug plug-in in the Grasshopper parametric platform of Rhino is used for UTCI simulation. The grid size for simulation calculation is set to 5 meters × 5 meters, and the monitoring height is 1.8 meters to carefully capture urban microclimate changes. All input parameters are set according to the annual conventional climate data standard of the study area to ensure the accuracy of the simulation results.
[0180] Finally, the plot state space consists of the geometric features of the plot 2, the target floor area ratio FAR of the plot {target} , the geometric parameters of the building, and the thermal environment state 2, which serves as the input for subsequent Agent training.
[0181] S2.2.2: Based on the state space s of the current tree Agent tree and the current site layout plan, generate the output action distribution parameters of the tree Agent through the Actor network to obtain the current policy π θ (a tree |s tree );
[0182] Initialize the Actor network parameters θ of the tree Agent tree , and input the state space s generated in S2.2.1 tree , which includes geometric features 2, the target floor area ratio FAR of the plot {target} , the geometric parameters of the tree, and the thermal environment state 2. Initialize the weights of θ tree using the Xavier method (the Xavier method is a common initialization method to ensure that the initial output of the network is not too large or too small), and the bias is initialized to 0. The input layer of the Actor network receives the plot state s tree , and the current site layout plan is processed using a convolutional neural network (CNN, Convolutional Neural Network) with a dimension of 256×256×1. Then, spatial features are extracted through three convolutional layers. The size of each convolutional kernel is 3×3, the stride is 2, and the number of output channels is 16, 32, and 64 respectively. Each layer is followed by a ReLU activation function, and finally, through a fully connected layer, the plot image feature 1×256 is output; the geometric parameters of the tree are processed using a long short-term memory network (LSTM, Long Short-Term Memory) with 128 hidden units, and the 4D parameters of each tree are processed step by step in time, and the hidden state T n of the last time step is output as 1×128; the thermal environment state 2 (dimension 1×6, including
[0183] [Comfort {ratio} , ArchOverLap {ratio} , FAR {current} , UTCI max , Boundary {out} , Size {out} ),through the fully connected layer, the output dimension is 1×64. Then, the current site layout map (1×256), the LSTM hidden state T n (1×128) and the thermal environment state (1×64) are concatenated, and the output dimension is 1×448. Then, the feature vector enters two fully connected layers, the input dimension is 1×448, the output dimension is 1×128, and is activated by ReLU. Finally, the output layer of the network generates the action distribution parameter π θ (a tree |s tree ), where the mean μ represents the expected value of the action (“the most likely action”), maps the input 1×128 to the output 1×6 through the fully connected layer and scales it to the action range through the tanh activation function; the standard deviation σ represents the variance of the action distribution (“the exploration range of the action”), also maps the input 1×128 to the output 1×6 through the fully connected layer and ensures a positive value through the softplus activation function. The Actor network outputs the action distribution parameter π θ (a tree |s tree ), where π θ (a tree |s tree ) is the current policy, defined as the probability distribution of the action a tree given the state s tree , sample the action a tree from the normal distribution N(μ,σ). Pass the action a tree to the environment, the environment updates the plot layout (add new trees or move existing trees based on the current site layout map), re - conducts the UTCI simulation (the simulation parameters are the same as in S2.2.1, the grid size is 5m×5m, and the monitoring height is 1.8m), calculates the reward r tree and generates the next state s arch+1 (including the updated geometric features, thermal environment state and current site layout map), and saves (s tree , a tree , r tree , s tree+1 ) to the experience storage module, and repeat this process until the maximum decision - making step N is collected or the preset termination condition is met.
[0184] S2.2.3: Reward function based on tree Agent, predict the state value V through the Critic networkφ (s tree ) and calculate the advantage A of the action tree ;
[0185] Initialize the parameters φ of the Critic network of the tree Agent tree , and using the stored experience data, the Critic network predicts the state value and calculates the advantage of the action. Traverse each piece of data (s tree , a tree , r tree , s tree+1 ) in the experience storage module, and input the states s tree and s tree+1 into the Critic network. The structure of the Critic network is the same as that of the Actor network, including a multi-branch structure. After being processed by the fusion layer, a feature vector (with a dimension of 1×448) is obtained. Then, the feature vector enters two fully connected layers, with an input dimension of 1×448 and an output dimension of 1×128, activated by ReLU. Finally, through the fully connected layer, the input 1×128 is mapped to the output 1×1 to generate the state value V φ (s tree ), where V φ (s tree ) is the value predicted by the Critic network (representing the expected value of the future discounted reward that can be obtained starting from the state s tree according to the current policy). Similarly, calculate V φ (s tree+1 ).
[0186] Calculate the cumulative value R arch , which is the training target of the Critic network, representing the current reward plus the discounted value of the next state. The formula is as follows:
[0187] R tree = r tree + γV φ (r tree+1 )
[0188] Among them, γ is the discount factor (representing the importance of future rewards. The closer the value is to 1, the more the Agent focuses on long-term benefits). In this embodiment, γ = 0.99.
[0189] Calculate the temporal difference error δ t , representing the error predicted by the Critic network, used to calculate the advantage. The formula is as follows:
[0190] δ t = r tree + γV φ (s tree+1 ) - V φ (s tree)
[0191] Among them, δ t is the temporal difference error (TD error), representing the difference between the predicted value and the actual return at time step t.
[0192] Calculate the advantage A tree , representing the relative goodness or badness of an action. The larger the value, the better the action, which is used to guide the Actor network to improve the policy. The formula is as follows:
[0193]
[0194] Among them, A tree is the advantage function, representing how much higher or lower the action a tree selected under the state s tree is compared to the average level. λ is the GAE parameter used to calculate the advantage of the action. In this embodiment, λ = 0.95; q is a non - negative integer (k = 0, 1, 2,..., ∞), representing the offset from the current time step t to the q - th step in the future.
[0195] S2.2.4: Based on the advantage A tree , the state value V φ (s tree ) and the policy π θ (a tree |s tree ), calculate the total loss L' through the loss function;
[0196] As Figure 2 shown, use the advantage A tree , the state value V φ (s tree ) and the policy π θ (a tree |s tree ) to calculate the loss function for preparing the network update. Calculate the policy gradient objective L CLIP , which is used to measure the improvement direction of the policy to ensure stable update. The formula is as follows:
[0197]
[0198] Among them, π θ (a tree |s tree ) is the current policy, is the old policy, r arch (θ) is the probability ratio of the new and old policies, representing the change amplitude of the current policy. ∈ is the clipping parameter. In this embodiment, ∈ = 0.2.
[0199] Calculate the value loss L VF, which is used to optimize the Critic network to make it more accurately predict the state value. The formula is as follows:
[0200]
[0201] Among them, represents the expectation, R tree is the cumulative value, V φ (s tree ) is the value predicted by the Critic.
[0202] Calculate the entropy loss S[π θ (s tree ), which represents the randomness of the policy. The larger the entropy value, the more random the policy, and the more inclined the Agent is to explore. The combined total loss l' combines policy improvement, value prediction, and exploration, and is used to guide the update of network parameters. The formula is as follows:
[0203] l' = L CLIP -c1L VF +c2S[π θ (s tree )
[0204] Among them, c1 is the value loss weight (the larger this value, the greater the training intensity of the Critic network), and c2 is the entropy loss weight (the larger this value, the stronger the exploration of the Agent).
[0205] S2.2.5: According to the total loss L', use the Adam optimizer to update the Actor network parameters θ tree and the Critic network parameters φ tree ;
[0206] In this embodiment, the Adam optimizer is used to update the Actor network parameters θ tree . Because θ tree determines the policy of the Actor network, after the update, the Agent can select better tree actions. The formula is as follows:
[0207]
[0208] Among them, η is the learning rate, and in this embodiment, η = 0.0003; is the gradient of the loss with respect to θ tree , which reflects the direction and magnitude of the parameter θ tree adjustment.
[0209] Update the Critic network parameters φ tree . Because φ tree determines the value prediction of the Critic network, after the update, the Critic can more accurately evaluate the action. The formula is as follows:
[0210]
[0211] Among them, is the gradient of the loss pair φ tree , which reflects the direction and magnitude of the adjustment of the parameter φ tree .
[0212] After updating the network parameters, clear the experience storage module and prepare for the next round of experience collection.
[0213] S2.2.6: Repeat training until the maximum number of rounds is reached
[0214] Repeat steps S2.2.2 to S2.2.6 to iteratively update θ tree and φ tree , until the preset maximum number of training rounds (set to 1000 rounds in this embodiment) is reached. The finally obtained θ tree and φ tree are the final parameters of the Actor and Critic networks of the tree Agent. The θ tree of the Actor network corresponds to the optimal policy π θ , which is the optimal tree layout strategy that meets the requirements of the target floor area ratio and the thermal environment. After the training is completed, save the Actor network parameter θ tree for subsequent generation of the tree layout.
[0215] S3: Apply the building Agent and the tree Agent to jointly generate the optimal layout
[0216] In this step, use the building Agent and the tree Agent trained in S2.1 and S2.2 to jointly generate the optimal building layout and tree layout on the same plot, dynamically balancing the requirements of the target floor area ratio and the thermal environment performance. The building Agent and the tree Agent alternately execute actions to gradually optimize the layout until the termination condition is met. The following are the specific steps, including the operation process and the result evaluation.
[0217] S3.1: Input the status and target parameters of the plot to be optimized.
[0218] In this step, input the state space of the plot to be optimized, the target floor area ratio FAR {target} , the target tree coverage rate Cover {target} , to provide the initial conditions for the layout optimization. The plot state space is generated based on S2.1.1 and includes the geometric feature 1 of the plot, the thermal environment state 1, and the current site layout map. At the same time, load the Actor network parameter of the building Agent trained in S2.1.
[0219] S3.2: Obtain the state space s of the plot to be optimized, i.e., the plot, according to the method in S2.2.1tree , including the geometric features 2 of the plot, the thermal environment status 2, and the current site layout plan; set the target tree coverage rate Cover for this area {target} ; load the Actor network parameters θ of the trained tree Agent tree ;
[0220] Next, the building Agent and the tree Agent alternately execute actions to collaboratively optimize the building layout and tree layout of the plot. In each iteration, first the building Agent generates or adjusts the building layout, and then the tree Agent generates or adjusts the tree layout.
[0221] S3.3: Take the geometric features 1, thermal environment status 1, and the current site layout plan of the current plot as inputs, and use the trained building Agent to optimize the building layout of the area with the constraints of meeting the target plot ratio and the preliminary thermal environment requirements, and output the current optimal building layout of the area, as well as the updated geometric features 1 and thermal environment status 1;
[0222] In this step, use the trained building Agent in S2.1 to optimize the building layout of the plot to meet the target plot ratio and the preliminary thermal environment requirements. Take s arch (including geometric features 1, thermal environment status 1, and the current site layout plan) as the input to generate an optimized building layout until the preset termination condition is met. Output the updated geometric features 1 and thermal environment status 1 for evaluating the thermal comfort and the effect of the building layout.
[0223] S3.4: Take the geometric features 2, thermal environment status 2, and the current site layout plan of the current plot as inputs, and use the trained tree Agent to optimize the tree layout of the area on the basis of the building layout generated in S3.3 with the constraints of meeting the target tree coverage rate and the preliminary thermal environment requirements, and output the current optimal tree layout of the area, as well as the updated geometric features 2 and thermal environment status 2;
[0224] In this step, use the trained tree Agent in S2.2 to optimize the tree layout on the basis of the building layout generated in S3.3 to further improve the thermal environment performance. Take s tree (including geometric features 2, thermal environment features 2, and the current site layout plan) as the input to generate an optimized tree layout until the preset termination condition is met. Output the updated geometric features 2 and thermal environment status 2 for evaluating the thermal comfort and the effect of the tree layout.
[0225] S3.5: Repeat S3.3 to S3.4 until the preset termination condition is met;
[0226] S3.6: Output the final optimized layout of the plot to be optimized and its thermal environment status.
[0227] In this step, according to the target plot ratio and the indicators in the building Agent reward function, the building geometric feature 1 and the thermal environment state 1 of the best building layout scheme among the various indicators are output. Then, according to the target tree coverage rate and the indicators in the tree Agent reward function, the tree geometric feature 2 and the thermal environment state 2 of the best tree layout scheme among the various indicators are output for the user to use for scheme screening or further optimizing the layout scheme.
[0228] It should be understood that those skilled in the art, inspired by the technical concept of the present invention and without departing from the content of the present invention, can also make various improvements or transformations based on the above content, which still fall within the protection scope of the present invention.
Claims
1. A spatial design decision-making method for the thermal environment performance of high-intensity areas based on reinforcement learning, characterized in that The method includes the following steps: Define the state space, action space, and reward function of the building Agent and the state space, action space, and reward function of the tree Agent respectively; Train the building Agent using the Proximal Policy Optimization algorithm, and train the tree Agent using the Proximal Policy Optimization algorithm based on the training results of the building Agent; Apply the trained building Agent and the trained tree Agent to jointly generate the optimal layout of buildings and trees in the area to be optimized.
2. The decision-making method for the spatial design of the thermal environment performance in high-intensity areas based on reinforcement learning according to claim 1, wherein, The state space of the building Agent represents the input information of the building Agent, including the geometric features of the plot 1, the target floor area ratio (FAR) of the plot {target} , the geometric parameters of the building, and the thermal environment state 1; The geometric features 1 of the said plot include: the vertex coordinates of each plot boundary, the plot area; the geometric parameters of the building include: the center point coordinates (x i , y i ) of the bottom surface of each building, the length l i , the width w i , the building height h i and the orientation angle θ i . If n buildings have been placed within the plot, the geometric parameters of the buildings are expressed as [x1, y1, l1, w1, h1, θ1,..., x n , y n , l n , w n , h n , θ n ; The thermal environment state 1 includes the following parameters: comfort area ratio Comfort {ratio} , building overlap rate ArchOverlap {ratio} , current floor area ratio FAR {current} , maximum value of UTCI UTCI max , building boundary violation index Boundary {out} 1, building size violation index Size {out} 1; The state space of the tree agent represents the input information of the tree agent, including the geometric features of the plot 2, the target tree coverage rate Cover {current} , the geometric parameters of the trees, and the thermal environment state 2; the geometric features of the plot 2 include: the coordinates of each vertex of the plot boundary, the floor area of the buildings already placed in the plot, and the geometric parameters of the buildings; the geometric parameters of the trees include the coordinates (x j , y j ) of each tree, the height h j and the crown diameter d j . If there are m trees already placed in the plot, the geometric parameters of the trees are represented as [x1, y1, h1, d1,..., x m , y m , h m , d m ; the thermal environment state 2 includes the following parameters: the comfortable area ratio Comfort {ratio} , the tree overlap rate TreeOverlap {ratio} , the current tree coverage rate Cover {current} , the maximum value of UTCI UTCI max , the tree boundary violation index Boundary {out} 2, the tree size violation index Size {out} 2.
3. The method for making a decision on the spatial design of the thermal environment performance of a high-intensity area based on reinforcement learning according to claim 2, wherein, The parameters of the thermal environment state 1 and the thermal environment state 2 are both obtained through UTCI simulation calculations: (1) Comfort area ratio {ratio} , representing the ratio of the area within the grid corresponding to a UTCI of 9 to 26 to the plot area: where K is the total number of grids; UTCI k is the UTCI value of the k-th grid; I(9 ≤ UTCI k <26) is an indicator function. If UTCI k is between 9 and 26, then I(9 ≤ UTCI k <26) is 1; otherwise, I(9 ≤ UTCI k <26) is 0; A grid is the area of a single grid, and A plot is the total area of the plot. (2) ArchOverlap (Building Overlap Rate) {ratio} : Wherein, A Overlap,i,j represents the overlapping bottom area between the i-th building and the j-th building; A i , A j are respectively the bottom area of the i-th building and the bottom area of the j-th building; I Overlap,i is an indicator function. If Ratio exceeds 0.6, then I Overlap,i = 1, otherwise I Overlap,i = 0; N is the total number of buildings within the current plot; (3) Current Floor Area Ratio FAR {current} : where F i is the storey height; is the floor area of the i-th building; A plot is the site area; (4) Maximum UTCI value, UTCI max , the highest UTCI value in the grid: where UTCI k is the UTCI value of the k-th grid obtained by UTCI simulation; (5) Building boundary violation index BOundary {out} 1. Calculate the ratio of the number of buildings exceeding the plot boundary line to the total number: Where, i boundary,i is an exponential function. If the i-th building exceeds the plot boundary, then I boundary,i = 1; otherwise, I boundary,i = 0; N is the number of buildings within the current plot; (6) Building Size Violation Index Size {out} 1. Calculate the ratio of the number of buildings exceeding the size limit to the total number of buildings: where I size,i is an exponential function. If the length l i or the width w i of the i-th building is not within the preset numerical range, then I size,i = 1; otherwise, I size,i = 0; (7) TreeOverlap {ratio} : In the formula, for the newly placed $i$-th tree, check whether its tree crown overlaps with the existing tree crowns; $A$ OVerlap,i,j represents the overlapping bottom area between the $i$-th tree and the $j$-th tree; $A$ i , $A$ j are the bottom areas of the $i$-th tree and the $j$-th tree respectively; $I$ Overlap,i is an indicator function. If Ratio exceeds 0.6, then $I$ Overlap,i = 1, otherwise $I$ Overlap,i = 0; $A$ Overlap,i,j is the overlapping area between the $i$-th tree and the $j$-th tree; $M$ is the total number of trees in the current plot; (8) Current tree coverage COver {current} : where d i is the crown diameter of the j-th tree; l i ·w i is the floor area of the i-th building, M is the total number of trees in the current plot, and N is the total number of buildings in the current plot; (9) Tree boundary violation index Boundary {out} 2. Calculate the ratio of the number of trees exceeding the plot boundary line to the total number of trees: where, I boundary,i is an exponential function, if the i-th tree exceeds the plot boundary, then I boundary,i = 1, otherwise I boundary,i = 0; (10) Tree size violation index Size {out} 2. Calculate the ratio of the number of trees exceeding the size limit to the total number: Wherein, I size,i is an exponential function. If the crown diameter d i or the tree height h i of the i-th tree is not within the preset numerical range, then I size,i = 1; otherwise, I size,i = 0.
4. The method for spatial design decision of the thermal environment performance of a high-intensity area based on reinforcement learning according to claim 1, characterized in that The action space of the building Agent represents the output actions of the building Agent, including new building placement and existing building adjustment; the new building placement refers to placing a new building within a plot; the existing building adjustment refers to specifying the adjustment amount for the i-th existing building: [Δx i , Δy i , Δl i , Δw i , Δh i , Δθ i , where Δx i , Δy i are the changes in the coordinates of the center point of the building's bottom surface; Δl i , Δw i are the changes in the length and width of the building's bottom surface; Δθ i is the change in the orientation angle of the building; The action space of the tree agent represents the output actions of the tree agent, including new tree placement and adjustment of existing trees; the new tree placement refers to placing new trees within the plot; the adjustment of existing trees specifies the adjustment amount for the j-th existing tree: [Δx j , Δy j , Δh j , Δd j , where Δx j , Δy j are the coordinate changes of the tree; Δh j is the height change of the tree; Δd j is the change in the crown diameter of the tree.
5. The decision-making method for the spatial design of the thermal environment performance of high-intensity areas based on reinforcement learning according to claim 3, wherein Reward function r of the building Agent arch is as follows: where ω1·Comfort {ratio} encourages maximizing the comfort area ratio; ω2·ArchOverlap {ratio} penalizes cases where the overlap between the new building and the existing building exceeds 60%; penalizes the floor area ratio deviation; ω4· max (0, UTCI max - 26) penalizes the maximum value of UTCI, UTCI max exceeding 26; ω5·Boundary {out}1 penalizes the building for exceeding the boundary; ω6·Size {out} 1 penalizes for exceeding the size limit; ω is the weight of different rewards; Reward function r of the Tree Agent tree is as follows: where τ1·Comfort {ratio} encourages maximizing the comfort area ratio; τ2·TreeOverlap {ratio} penalizes cases where the overlap of new trees with existing trees exceeds 60%; penalizes the deviation of the tree coverage rate; τ4· max (0, UTCI max - 26) penalizes the maximum value of UTCI, UTCI max exceeding 26; τ5·Boundary {out} 2 penalizes trees exceeding the boundary; τ6·Size {out} 2 penalizes trees with over - sized dimensions; τ is the weight of different rewards.
6. The method for making a decision on the spatial design of the thermal environment performance of a high-intensity area based on reinforcement learning according to claim 5, wherein The training of the building Agent using the Proximal Policy Optimization algorithm includes the following steps: S2.1.1: Conduct thermal environment simulation on the plot: Randomly generate a batch of empty plots, extract the geometric feature 1 of each plot and assign the target plot ratio to each plot, then generate a site current layout diagram representing the positions where buildings can be laid out within the plot based on the geometric feature 1 of the plot, and conduct UTCI simulation on each plot based on historical meteorological data to obtain the initial thermal environment state 1; S2.1.2: Based on the state space s of the current building Agent arch and the current site layout diagram, generate the output action distribution parameters of the building Agent through the Actor network to obtain the current policy π θ (a arch |s arch ); S2.1.3: Based on the reward function of the building Agent, predict the state value V through the Critic network φ (s arch ) and calculate the advantage A of the action arch ; S2.1.4: Based on Advantage A arch , State Value V φ (s arch ) and Policy π θ (a arch |s arch ), calculate the total loss L through the loss function; S2.1.5: Update the Actor network parameters θ according to the total loss L using the Adam optimizer arch and the Critic network parameters φ arch ; S2.1.6: Repeat steps S2.1.2 to S2.1.5 to iteratively update θ arch and φ arch , until the preset maximum number of training rounds is reached, so as to obtain the final parameters of the Actor network and the final parameters of the Critic network of the building Agent.
7. The decision-making method for the spatial design of the thermal environment performance of high-intensity areas based on reinforcement learning according to claim 6, characterized in that The training of the tree Agent using the Proximal Policy Optimization algorithm includes the following steps: S2.2.1: Conduct thermal environment simulation on the plot: Extract the geometric feature 2 of the plot and the geometric parameters of the building after the training of the building Agent is completed, and assign the target greening rate, then generate a site current layout diagram representing the positions where trees can be laid out within the plot based on the geometric feature 2 of the plot and the geometric parameters of the building, and conduct UTCI simulation on each plot based on historical meteorological data to obtain the initial thermal environment state 2; S2.2.2: Based on the state space s of the current tree Agent tree and the current site layout diagram, generate the output action distribution parameters of the tree Agent through the Actor network to obtain the current policy π θ (a tree |s tree ); S2.2.3: Based on the reward function of the tree Agent, predict the state value V through the Critic network φ (s tree ) and calculate the advantage A of the action tree ; S2.2.4: Based on Advantage A tree , State Value V φ (s tree ) and Policy π θ (a tree |s tree ), calculate the total loss L' through the loss function; S2.2.5: Update the Actor network parameters θ and the Critic network parameters φ according to the total loss L' using the Adam optimizer tree and the Critic network parameters φ tree ; S2.2.6: Repeat S2.2.2 to S2.2.5 to iteratively update θ tree and φ tree , until the preset maximum number of training rounds is reached, thereby obtaining the final parameters of the Actor network and the final parameters of the Critic network of the building Agent.
8. The decision-making method for the spatial design of the thermal environment performance in high-intensity areas based on reinforcement learning according to claim 7, wherein The application of the trained building Agent and the trained tree Agent to jointly generate the optimal layout of buildings and trees in the area to be designed and decided includes the following steps: S3.1: Obtain the state space s of the area to be optimized, i.e., the plot, according to the method in S2.1.1 arch , including the geometric feature 1 of the plot, the thermal environment state 1, and the current site layout plan; set the target floor area ratio FAR of this area {target} ; load the Actor network parameters θ of the trained building Agent arch ; S3.2: Obtain the state space s of the area to be optimized, i.e., the plot, according to the method in S2.2.1 tree , including the geometric features 2 of the plot, the thermal environment state 2, and the current site layout plan; set the target tree coverage rate Vover of this area {target} ; load the Actor network parameters θ of the trained tree Agent tree ; S3.3: Take the geometric feature 1, thermal environment state 1, and site current layout diagram of the current plot as inputs, and apply the trained building Agent to optimize the building layout of the area with the constraints of meeting the target plot ratio and the preliminary thermal environment requirements, and output the current optimal building layout of the area and the updated geometric feature 1 and thermal environment state 1; S3.4: Take the geometric feature 2, thermal environment state 2, and site current layout diagram of the current plot as inputs, and apply the trained tree Agent to optimize the tree layout of the area with the constraints of meeting the target tree coverage rate and the preliminary thermal environment requirements on the basis of the building layout generated in S3.3, and output the current optimal tree layout of the area and the updated geometric feature 2 and thermal environment state 2; S3.5: Repeat S3.3 to S3.4 until the preset termination condition is met; S3.6: Output the final optimized layout and its thermal environment status of the area to be optimized, i.e., the plot: According to the target plot ratio and the indicators in the reward function of the building Agent, output the geometric feature 1 and thermal environment status 1 corresponding to the best building layout scheme among each indicator, and according to the target tree coverage rate and the indicators in the reward function of the tree Agent, output the geometric feature 2 and thermal environment status 2 corresponding to the best tree layout scheme among each indicator.
9. The decision-making method for the spatial design of the thermal environment performance of high-intensity areas based on reinforcement learning according to any one of claims 6-8, characterized in that The current site layout diagram is a grayscale image of 256×256 pixels, and the height of buildings and / or trees is mapped by the shade of color.
Citation Information
Patent Citations
Spatial layout design method for urban buildings
CN111090899A
Automatic building space combination method based on multi-agent deep reinforcement learning
CN118296702A
Automatic driving behavior decision-making method and system based on reinforcement learning
CN118770284A
Two-stage urban thermal comfort rapid prediction method based on building and greening layout
CN119476043A