High-intensity district thermal environment performance spatial design decision method based on reinforcement learning

By training building and tree agents based on reinforcement learning, the optimal layout is generated collaboratively, solving the problem of inefficient urban thermal environment optimization in traditional methods and achieving rapid iteration and thermal comfort optimization under multiple constraints.

CN120297150BActive Publication Date: 2025-10-17TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510476725.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-10-17
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Traditional urban thermal environment optimization methods are inefficient in high-intensity areas, difficult to iterate quickly, unable to dynamically respond to complex thermal environment indicators, and lack the ability to automatically optimize building and tree layouts, resulting in insufficient universality of optimization results.

Method used

A reinforcement learning-based approach is adopted. By defining the state space, action space, and reward function of building and tree agents, the agents are trained using a proximal policy optimization algorithm to collaboratively generate the optimal layout of buildings and trees to meet thermal comfort requirements.

Benefits of technology

It achieves efficient and automated optimization, improves the iteration efficiency of urban planning, adapts to different plot shapes and floor area ratios, comprehensively considers multiple constraints, and ensures the practicality and feasibility of the optimization results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297150B_ABST
    Figure CN120297150B_ABST
Patent Text Reader

Abstract

The application discloses a high-strength piece area thermal environment performance space design decision method based on reinforcement learning, and relates to the technical field of urban thermal comfort evaluation and optimization. The method comprises the following steps: defining the state space, action space and reward function of the building Agent and the tree Agent respectively; training the building Agent by using a proximal policy optimization algorithm, and training the tree Agent by using the proximal policy optimization algorithm based on the training result of the building Agent; and applying the trained building Agent and the trained tree Agent to cooperatively generate the optimal layout of the building and the tree in the to-be-optimized piece area. The method dynamically adjusts the building layout and the tree layout through the reinforcement learning algorithm, generates the building arrangement and the tree arrangement scheme satisfying the thermal comfort requirement under the constraint of random plot shape and volume rate, can quickly optimize the thermal comfort index of the urban thermal environment, and provides scientific support for urban planning and built environment optimization.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of urban thermal comfort assessment and optimization, and particularly relates to a high-intensity area thermal environment performance space design decision method based on reinforcement learning. BACKGROUND

[0002] With the acceleration of urbanization, the heat island effect in high-intensity areas of cities (i.e., areas with high building density and large plot ratio) is increasingly severe. High-density building layout and limited ventilation conditions have a significant impact on the thermal comfort of residents and building energy consumption. Therefore, reasonable building layout and tree layout are important means to alleviate the heat island effect and play an important role in thermal environment optimization. Traditional urban thermal environment optimization methods mainly rely on manual design or rule-based optimization algorithms to improve the thermal environment through the subjective experience of architects or simple geometric adjustments. However, these methods have the following problems: first, the optimization process is tedious and difficult to adapt to the rapid iteration requirements of urban planning, especially under the complex spatial and functional constraints of high-intensity areas, the efficiency problem is particularly prominent; second, traditional methods lack dynamic response capability to complex thermal environment indicators (such as UTCI, Universal Thermal Climate Index), and cannot comprehensively consider the comprehensive influence of building layout and tree layout on microclimate; in addition, existing methods usually fail to achieve efficient automated optimization under random plot and plot ratio constraints, resulting in insufficient universality of the optimization results.

[0003] In the prior art, some studies attempt to use numerical simulation tools (such as ENVI-met or Ladybug) to evaluate the thermal environment, but these tools have high computational cost and are difficult to provide real-time feedback on the results of design changes. At the same time, although some machine learning methods are used for thermal environment prediction, they mainly focus on static prediction rather than dynamic optimization, and lack the ability to actively adjust building parameters (such as location, size, angle) and tree parameters (such as location, height, crown diameter). Therefore, how to quickly optimize building layout and tree layout to improve thermal comfort under the given plot shape and plot ratio conditions through an efficient automated method has become a technical problem to be solved. SUMMARY

[0004] In view of the deficiencies of the prior art, the purpose of the present application is to provide a high-intensity area thermal environment performance space design decision method based on reinforcement learning, which aims to dynamically adjust building layout and tree layout through reinforcement learning algorithm, generate building arrangement and tree arrangement scheme that meet the thermal comfort requirements under the constraints of random plot shape and plot ratio, and quickly optimize the thermal comfort indicators of urban thermal environment, providing scientific support for urban planning and built environment optimization.

[0005] The technical solution of the present application is:

[0006] A reinforcement learning-based decision-making method for high-intensity area thermal environment performance space design includes the following steps:

[0007] Define the state space, action space, and reward function of the building agent and the state space, action space, and reward function of the tree agent respectively;

[0008] The building agent is trained using the proximal policy optimization algorithm. Based on the training results of the building agent, the tree agent is trained using the proximal policy optimization algorithm.

[0009] The trained building agent and the trained tree agent are used to collaboratively generate the optimal layout of buildings and trees in the area to be optimized.

[0010] Furthermore, according to the reinforcement learning-based high-intensity area thermal environment performance space design decision method, the state space of the building agent represents the input information of the building agent, including the geometric characteristics of the plot 1, the target volume ratio FAR of the plot {target} , the geometric parameters of the building and the thermal environment state 1; the geometric features 1 of the plot include: the coordinates of each vertex of the plot boundary, the plot area; the geometric parameters of the building include: the coordinates of the center point of each building bottom surface (x i ,y i ), length l i 、Width w i 、Building height h i and the orientation angle θ i If n buildings have been placed in the plot, the geometric parameters of the buildings are expressed as [x1,y1,l1,w1,h1,θ1,...,x n ,y n ,l n ,w n ,h n ,θ n ]; The thermal environment state 1 includes the following parameters: Comfort area ratio Comfort {ratio} , ArchOverlap {ratio} 、Current Floor Area Ratio FAR {current} UTCI maximum value UTCI max , Building Boundary Violation Index {out} 1. Building size violation indicator Size {out} 1;

[0011] The state space of the tree agent represents the input information of the tree agent, including the geometric features of the plot 2, the target tree coverage rate Cover {current}, geometric parameters of trees and thermal environment status 2; the geometric features 2 of the plot include: the coordinates of each vertex of the plot boundary, the floor area of ​​the building placed in the plot, and the geometric parameters of the building; the geometric parameters of the trees include the coordinates of each tree (x j ,y j ), height h j and crown diameter d j If m trees have been placed in the plot, the geometric parameters of the trees are expressed as [x1,y1,h1,d1,...,x m ,y m ,h m ,d m ]; The thermal environment state 2 includes the following parameters: Comfort area ratio Comfort {ratio} , TreeOverlap {ratio} 、Current tree coverage {current} UTCI maximum value UTCI max , Boundary {out} 2. Tree size violation indicator Size {out} 2.

[0012] Furthermore, according to the reinforcement learning-based high-intensity area thermal environment performance space design decision method, the parameters of thermal environment state 1 and thermal environment state 2 are both obtained through UTCI simulation calculation:

[0013] (1) Comfort area ratio {ratio} , which represents the ratio of the area within the grid to the area of ​​the plot when the UTCI is 9 to 26:

[0014]

[0015] Where K is the total number of grids; UTCI k is the UTCI value of the kth grid; I(9≤UTCI k <26) is the indicator function, if UTCI k In the range of 9 to 26, I (9 ≤ UTCI k <26) is 1, otherwise I(9≤UTCI k <26) is 0; A grid is the area of ​​a single grid, A plot is the total area of ​​the plot;

[0016] (2) ArchOverlap {ratio} :

[0017]

[0018] Where AOverlap,i,j represents the overlapping floor area of the ith building and the jth building; A i j are the floor area of the ith building and the floor area of the jth building, respectively; I Overlap,i is an indicator function, I Overlap,i = 1 if the Ratio exceeds 0.6, otherwise I Overlap,i = 0; N is the total number of buildings in the current plot;

[0019] (3) Current FAR {current} :

[0020]

[0021] where F i is the floor height; is the building area of the ith building; A plot is the site area;

[0022] (4) Maximum UTCI UTCI max , the highest UTCI value in the grid:

[0023]

[0024] where UTCI k is the UTCI value of the kth grid obtained by UTCI simulation;

[0025] (5) Boundary violation index Boundary {out} 1, the proportion of the number of buildings exceeding the plot boundary line to the total number:

[0026]

[0027] where I boundary,i is the exponential function;

[0028] (6) Size violation index Size {out} 1, the proportion of the number of buildings exceeding the size limit to the total number:

[0029]

[0030] where I size,i is the exponential function;

[0031] (7) Tree overlap rate TreeOverlap {ratio} :

[0032]

[0033] ​In the formula, for the newly placed i th tree, check whether its tree crown overlaps with the existing tree crown; A Overlap,i,j represents the overlapping base area of the i th tree and the j th tree; A i j , A Overlap,i Overlap,i , A Overlap,i Overlap,i,j is the overlapping area of the i th tree and the j th tree; M is the total number of trees in the current plot.

[0034] (8) Current tree coverage Cover {current} :

[0035]

[0036] In the formula, d i is the crown diameter of the j th tree; l i · w i is the building area of the i th building, M is the total number of trees in the current plot, and N is the total number of buildings in the current plot.

[0037] (9) Tree boundary violation index Boundary {out} 2, the proportion of the number of trees exceeding the plot boundary to the total number:

[0038]

[0039] In the formula, I boundary,i is an exponential function, and I boundary,i = 1 if the i th tree exceeds the plot boundary, otherwise I boundary,i = 0.

[0040] (10) Tree size violation index Size {out} 2, the proportion of the number of trees exceeding the size limit to the total number:

[0041]

[0042] In the formula, I size,i is an exponential function.

[0043] Further, according to the high-intensity area thermal environment performance space design decision method based on reinforcement learning, the action space of the building Agent represents the output action of the building Agent, including new building placement and existing building adjustment; the new building placement refers to placing a new building in the plot; the existing building adjustment refers to specifying an adjustment amount for the i th existing building: [Δx i ,Δy​​​i ,Δl i ,Δw i ,Δh i ,Δθ i ], where Δx i ,Δy i is the coordinate change of the center point of the bottom surface of the building; Δl i ,Δw i is the length and width change of the building bottom surface; Δθ i The building's orientation angle changes;

[0044] The action space of the tree agent represents the output action of the tree agent, including new tree placement and existing tree adjustment; the new tree placement refers to placing new trees in the plot; the existing tree adjustment is the specified adjustment amount of the j-th existing tree: [Δx j ,Δy j ,Δh j ,Δd j ], where Δx j ,Δy j is the coordinate change of the tree; Δh j is the height change of the tree; Δd j is the change in crown diameter of the tree.

[0045] Furthermore, according to the high-intensity area thermal environment performance space design decision method based on reinforcement learning, the building agent's reward function r arch for:

[0046]

[0047]

[0048] Where, ω1·Comfort {ratio} Encourage maximization of comfort area ratio; ω2·ArchOverlap {ratio} Penalize new buildings that overlap with existing buildings by more than 60%; Penalty for floor area ratio deviation; ω4· max (0,UTCI max -26) Penalty UTCI Maximum UTCI max More than 26; ω5·Boundary {out}1 Penalize buildings that exceed the boundary; ω6·Size {out} 1 is the penalty size exceeding the limit; ω is the weight of different rewards;

[0049] The reward function r of the tree agent tree for:

[0050]

[0051] where τ1·Comfort {ratio} encourages maximizing the comfort area ratio; τ2·TreeOverlap {ratio} penalizes new trees overlapping with existing trees more than 60%; penalizes number coverage bias;

[0052] τ4· max (0, UTCI max -26) penalizes UTCI maximum UTCI max exceeding 26; τ5·Boundary {out} 2 penalizes trees exceeding the boundary; τ6·Size {out} 2 penalizes tree size exceeding the limit; τ is the weight of different rewards.

[0053] Further, according to the high-intensity area thermal environment performance spatial design decision-making method based on reinforcement learning, the building Agent is trained by using a proximal policy optimization algorithm, and the method comprises the following steps:

[0054] S2.1.1: thermal environment simulation of the plot: a batch of empty plots is randomly generated, geometric features 1 of each plot are extracted, a target volume rate of each plot is given, a site status layout for representing the positions of the layoutable buildings in the plot is generated based on the geometric features 1 of the plot, and an initial thermal environment state 1 is obtained by performing UTCI simulation on each plot based on historical meteorological data;

[0055] S2.1.2: based on the state space s arch of the current building Agent and the site status layout, an output action distribution parameter of the building Agent is generated by an Actor network to obtain a current policy π θ (a arch |s arch );

[0056] S2.1.3: based on the reward function of the building Agent, a state value V φ (s arch ) is predicted by a Critic network, and an advantage A arch of the action is calculated;

[0057] S2.1.4: based on the advantage A arch , the state value V φ (s arch ) and the policy π θ (a arch |s arch ), a total loss L is calculated by a loss function;

[0058] S2.1.5: Update the Actor network parameters θ using the Adam optimizer based on the total loss L arch and Critic network parameter φ arch ;

[0059] S2.1.6: Repeat steps S2.1.2 to S2.1.5 to iteratively update θ arch and φ arch , until the preset maximum number of training rounds is reached, thereby obtaining the final parameters of the Actor network of the building agent and the final parameters of the Critic network.

[0060] Furthermore, according to the reinforcement learning-based high-intensity area thermal environment performance space design decision method, the tree agent training using the proximal strategy optimization algorithm includes the following steps:

[0061] S2.2.1: Perform thermal environment simulation on the plot: Extract the geometric features of the plot after the building agent training is completed2, the geometric parameters of the building, and set the target greening ratio. Then, based on the geometric features2 of the plot and the geometric parameters of the building, generate a site layout map that represents the locations where trees can be placed within the plot. UTCI simulation is performed on each plot based on historical meteorological data to obtain the initial thermal environment state2.

[0062] S2.2.2: Based on the state space s of the current tree agent tree And the current layout of the site, generate the output action distribution parameters of the tree agent through the Actor network, and obtain the current strategy π θ (a tree |s tree );

[0063] S2.2.3: Based on the reward function of the tree agent, the state value V is predicted through the critic network φ (s tree ) and calculate the advantage A of the action tree ;

[0064] S2.2.4: Based on Advantage A tree 、Status value V φ (s tree ) and strategy π θ (a tree |s tree ), calculate the total loss L' through the loss function;

[0065] S2.2.5: Update the Actor network parameters θ using the Adam optimizer based on the total loss L' tree and Critic network parameter φ tree ;

[0066] S2.2.6: repeat S2.2.2 to S2.2.5, iteratively update θ tree and φ tree , until reaching a preset maximum training round, thereby obtaining the final parameters of the Actor network and the final parameters of the Critic network of the building Agent.

[0067] Further, according to the high-intensity area thermal environment performance spatial design decision method based on reinforcement learning, the trained building Agent and the trained tree Agent are applied to cooperatively generate an optimal layout of buildings and trees in the area to be designed and decided, comprising the following steps:

[0068] S3.1: obtain the state space s of the area to be optimized, i.e. the plot, according to the method of S2.1.1 arch , containing the geometric characteristics 1 of the plot, the thermal environment state 1, and the site current layout map; set the target FAR of the area {target} ; load the Actor network parameters θ arch of the trained building Agent;

[0069] S3.2: obtain the state space s of the area to be optimized, i.e. the plot, according to the method of S2.2.1 tree , containing the geometric characteristics 2 of the plot, the thermal environment state 2, and the site current layout map; set the target tree coverage Cover of the area {target} ; load the Actor network parameters θ tree of the trained tree Agent;

[0070] S3.3: take the current geometric characteristics 1 of the plot, the thermal environment state 1, and the site current layout map as inputs, and apply the trained building Agent to optimize the building layout of the area under the constraint conditions of meeting the target FAR and the preliminary thermal environment requirements, output the current optimal building layout of the area and the updated geometric characteristics 1 and thermal environment state 1;

[0071] S3.4: take the current geometric characteristics 2 of the plot, the thermal environment state 2, and the site current layout map as inputs, and apply the trained tree Agent to optimize the tree layout of the area under the constraint conditions of meeting the target tree coverage Cover and the preliminary thermal environment requirements, based on the building layout generated in S3.3, output the current optimal tree layout of the area and the updated geometric characteristics 2 and thermal environment state 2;

[0072] S3.5: repeat S3.3 to S3.4 until the preset termination condition is met;

[0073] S3.6: Output the final optimized layout of the plot to be optimized and its thermal environment state: According to the target volume rate and the indicators in the reward function of the building Agent, output the geometric characteristics 1 and thermal environment state 1 corresponding to the best building layout scheme in each indicator, and according to the target tree coverage rate and the indicators in the reward function of the tree Agent, output the geometric characteristics 2 and thermal environment state 2 corresponding to the best tree layout scheme in each indicator.

[0074] Further, according to the high-intensity plot thermal environment performance spatial design decision method based on reinforcement learning, the site status layout map is a 256x256 pixel gray image, and the heights of buildings and / or trees are mapped by the depth of color.

[0075] Compared with the prior art, the present application has the following beneficial effects:

[0076] (1) High-efficiency automatic optimization: the fast decision mechanism based on reinforcement learning, intelligent self-adaptive adjustment, and full use of modern computing power for parallel computing dynamic adjustment of building layout and tree layout, compared with traditional manual design or numerical simulation method, the efficiency is improved by tens of times, which meets the needs of rapid iteration of urban planning.

[0077] (2) Comprehensive optimization: the method of the present application simultaneously satisfies multiple constraints such as plot boundary, volume rate, overlap rate and thermal comfort, and combines with tree layout optimization, fully considers the adjusting effect of greening on thermal environment, and ensures the practicality and feasibility of the optimization result.

[0078] (3) Strong universality: random plots are used for algorithm training, so that the method of the present application can adapt to plots of different shapes and volume rates, and has strong generalization ability. BRIEF DESCRIPTION OF DRAWINGS

[0079] Figure 1 The flowchart of the high-intensity plot thermal environment performance spatial design decision method based on reinforcement learning of the present embodiment;

[0080] Figure 2 The structural diagram of the proximal policy optimization algorithm;

[0081] Figure 3 The structural diagram of the Actor network in the building Agent of the present embodiment;

[0082] Figure 4 The structural diagram of the Critic network in the building Agent of the present embodiment. DETAILED DESCRIPTION

[0083] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the related drawings.

[0084] Figure 1 is a flow chart of the high-intensity plot thermal environment performance spatial design decision method based on reinforcement learning of the present embodiment.

[0085] As shown in Figure 1 , the high-intensity plot thermal environment performance spatial design decision method based on reinforcement learning comprises the following steps:

[0086] S1: define the state space, action space and reward function of the building Agent and the state space, action space and reward function of the tree Agent respectively;

[0087] S1.1: define the state space of the building Agent and the state space of the tree Agent respectively;

[0088] The state space is a set of features used to describe the current state of the environment in reinforcement learning, which represents the input information of the Agent.

[0089] In the present embodiment, the state space s arch of the building Agent is represented by the input information of the building Agent, which includes the geometric characteristics 1 of the plot, the target FAR {target} of the plot, the geometric parameters of the building and the thermal environment state 1; the geometric characteristics 1 of the plot include the coordinates of each vertex of the plot boundary and the plot area; the target FAR {target} is randomly obtained within the design specification range, which is used to constrain the total floor area of the building; the geometric parameters of the building include the coordinates (x i ,y i ), length l i , width w i , building height h i and orientation angle θ i of the center point of each building bottom surface, if n buildings have been placed in the plot, the geometric parameters of the building are represented as [x1,y1,l1,w1,h1,θ1,...,x n ,y n ,l n ,w n ,h n ,θ n ]; the thermal environment state 1 includes the following parameters: Comfort {ratio} , ArchOverlap {ratio} , current FAR {current} , UTCI max , Boundary {out} 1, Size {out}1. In this embodiment, the parameters of thermal environment state 1 are obtained through UTCI simulation calculation, as follows:

[0090] (1) Comfort area ratio {ratio} , which represents the ratio of the area within the grid to the plot area when UTCI is 9 to 26:

[0091]

[0092] Where K is the total number of grids; UTCI k is the UTCI value of the kth grid; I(9≤UTCI k <26) is the indicator function, if UTCI k In the range of 9 to 26, I (9 ≤ UTCI k <26) is 1, otherwise I(9≤UTCI k <26) is 0; A grid is the area of ​​a single grid, A plot is the total area of ​​the plot.

[0093] (2) ArchOverlap {ratio} :

[0094]

[0095] Where, for the newly placed i-th building in the plot, check whether its base area overlaps with the existing buildings (1st to i-1) in the plot; this implementation uses Ratio to calculate the ratio of the overlapping area to the base area of ​​the overlapping existing buildings; A Overlap,i,j A represents the overlapping bottom area of ​​the i-th building and the j-th building; i ,A j are the base areas of the i-th building and the j-th building respectively; I Overlap,i is the indicator function. If Ratio exceeds 0.6, I Overlap,i =1, otherwise I Overlap,i =0; N is the total number of buildings in the current plot.

[0096] (3) Current Floor Area Ratio (FAR) {current} :

[0097]

[0098] Where, F i is the floor height, in this implementation, F i =3; is the building area of ​​the i-th building; A plot For the site area.

[0099] (4) UTCI maximum value UTCImax , the highest UTCI value in the grid:

[0100]

[0101] where UTCI k is the UTCI value of the kth grid obtained by UTCI simulation using Ladybug.

[0102] (5) Building boundary violation indicator Boundary {out} 1, the proportion of the number of buildings exceeding the plot boundary line to the total number:

[0103]

[0104] where I boundary,i is an exponential function, I boundary,i = 1 if the ith building exceeds the plot boundary, otherwise I boundary,i = 0; N is the number of buildings in the current plot.

[0105] (6) Building size violation indicator Size {out} 1, the proportion of the number of buildings exceeding the size limit to the total number:

[0106]

[0107] where I size,i is an exponential function, I i = 1 if the length l i > 20m or the width w size,i > 20m of the ith building, otherwise I size,i = 0.

[0108] In the embodiment, the state space s tree of the tree agent represents the input information of the tree agent, including the geometric characteristics 2 of the plot, the target tree coverage Cover {current} , the geometric parameters of the trees and the thermal environment state 2; the geometric characteristics 2 of the plot include: the coordinates of each vertex of the plot boundary, the floor area of the buildings placed in the plot, the geometric parameters of the buildings; the geometric parameters of the trees include the coordinates (x j , y j ), height h j and crown diameter d j of each tree, if m trees have been placed in the plot, the geometric parameters of the trees are represented as [x1, y1, h1, d1,..., xm, ym, hm, dm]. m m m m ​​​]; The thermal environment state 2 includes the following parameters: Comfort area ratio Comfort {ratio} , TreeOverlap {ratio} 、Current tree coverage {current} UTCI maximum value UTCI max , Boundary {out} 2. Tree size violation indicator Size {out} 2. The parameters of thermal environment state 2 in this embodiment are also obtained through UTCI simulation calculation, and the comfort area ratio is Comfort {ratio} and UTCI maximum value UTCI max The formula for is given above, and the calculation formulas for the other four parameters are as follows:

[0109] I. TreeOverlap {ratio} :

[0110]

[0111] Where, for the newly placed i-th tree, check whether its crown overlaps with the existing crowns (1st to i-1st); this embodiment calculates the ratio of the overlapping base area to the base area of ​​the overlapping existing trees through Ratio; A Overlap,i,j represents the overlapping base area between the i-th tree and the j-th tree; A i ,A j are the base areas of the i-th tree and the j-th tree respectively; I Overlap,i is the indicator function. If Ratio exceeds 0.6, I Overlap,i =1, otherwise I Overlap,i =0;A Overlap,i,j is the overlapping area between the i-th tree and the j-th tree; M is the total number of trees in the current plot.

[0112] II. Current tree cover {current} :

[0113]

[0114] Where, d i is the crown diameter of the jth tree; l i w i is the building area of ​​the i-th building, M is the total number of trees in the current plot, and N is the total number of buildings in the current plot.

[0115] III. Boundary {out} 2. Count the ratio of the number of trees that extend beyond the boundary of the plot to the total number of trees:

[0116]

[0117] where I boundary,i is an indicator function, I boundary,i = 1 if the i-th tree exceeds the plot boundary, otherwise I boundary,i = 0; M is the number of trees in the current plot.

[0118] IV. Tree size violation indicator Size {out} 2, the ratio of the number of trees exceeding the size limit to the total number of trees:

[0119]

[0120] where I size,i is an indicator function, I i = 1 if the i-th tree’s crown diameter d i < 2m or d i > 10m or the tree’s height h i < 3m or h size,i > 30m, otherwise I size,i = 0.

[0121] S1.2: Define the action space a arch for building agents and a tree for tree agents, respectively.

[0122] Action space is the set of operations that an agent can perform in reinforcement learning, representing the output action of the agent. In this embodiment, the action space a arch for building agents represents the output action of the building agents, including two types of actions: new building placement and existing building adjustment. New building placement specifically refers to placing the n-th new building in the plot, and the geometric parameters of the building are [x n , y n , l n , w n , h n , θ n ]. The center point coordinates (x n , y n ) of the building bottom are limited within the plot boundary; the length l n and width w n of the building bottom are set according to actual conditions, and in this embodiment, they are within the range of 5m to 20m; the orientation angle of the building is within the range of 0° to 180°. Existing building adjustment refers to specifying the adjustment amount for the i-th existing building: [Δx i , Δy i , Δl i , Δw i, Δh i , Δθ i ], wherein Δx i , Δy i are the coordinate changes of the center point of the building base, the step size is L1, and after adjustment, it is still within the plot; Δl i , Δw i are the length and width changes of the building base, the step size is L2, and in the embodiment, after adjustment, l i , w i are in the range of 5m to 20m; Δθ i is the orientation angle change of the building, the step size is L3, and after adjustment, θ i is in the range of 0° to 180°;

[0123] In the embodiment, the action space of the tree Agent represents the output actions of the tree Agent, including new tree placement and existing tree adjustment. New tree placement specifically refers to placing the mthnew tree within the plot, and the geometric parameters of the tree are [x m , y m , h m , d m ]. The coordinates (x m , y m ) of the tree are limited within the plot boundary and do not coincide with the building coverage area; the height h m of the tree is set according to the actual situation, and in the embodiment, it is in the range of 3m to 30m; the crown diameter d m of the tree is set according to the actual situation, and in the embodiment, it is in the range of 2m to 10m. Existing tree adjustment refers to specifying adjustment amounts [Δx j , Δy j , Δh j , Δd j ] for the jthexisting tree, wherein Δx j , Δy j are the coordinate changes of the tree, the step size is L4, and after adjustment, it is still within the plot and does not coincide with the building coverage area; Δh j is the height change of the tree, the step size is L5, and after adjustment, h j is in the range of 3m to 30m; Δd j is the crown diameter change of the tree, the step size is L6, and after adjustment, d j is in the range of 2m to 10m.

[0124] S1.3: Define the reward function r arch of the building Agent and the reward function r tree of the tree Agent, respectively;

[0125] The reward function is an indicator used in reinforcement learning to evaluate the effectiveness of the agent's actions and guide the agent to optimize its goals.

[0126] The reward function r of the building agent designed in this embodiment is arch for:

[0127]

[0128] Where, ω1·Comfort {ratio} Encourage maximization of comfort area ratio; ω2·ArchOverlap {ratio} Penalize new buildings that overlap with existing buildings by more than 60%; Penalty floor area ratio deviation; ω4·max(0,UTCI max -26) Penalty UTCI Maximum UTCI max More than 26; ω5·Boundary {out} 1 Penalty for buildings beyond the boundary; ω6·Size {out} 1 is the penalty for exceeding the size limit; ω is the weight of different rewards.

[0129] The reward function r of the tree agent designed in this embodiment is tree for:

[0130]

[0131] Where, τ1·Comfort {ratio} Encourage maximization of the comfort area ratio; τ2·TreeOverlap {ratio} Penalize cases where new trees overlap with existing trees by more than 60%; Penalize number coverage deviation;

[0132] τ4·max(0,UTCI max -26) Penalty UTCI Maximum UTCI max More than 26; τ5·Boundary {out} 2 Penalize trees beyond the boundary; τ6·Size {out} 2 penalizes tree size exceeding the limit; τ is the weight of different rewards.

[0133] S2: Use the proximal policy optimization algorithm to train the building agent and the tree agent;

[0134] S2.1: Training building agents using proximal policy optimization algorithms;

[0135] Figure 2is the structural diagram of the proximal policy optimization algorithm, the proximal policy optimization (PPO) algorithm is based on the Actor-Critic framework, which uses a policy network (Actor) to select actions and a value network (Critic) to evaluate action values, and uses a double-network collaborative optimization mechanism to realize the stability control of policy update. The PPO algorithm controls the amplitude of policy update to ensure the stability of the training process and finally generates the optimal building layout scheme. The following are the specific steps of training the building Agent using the PPO algorithm, including all the parameters and operation processes that need to be set.

[0136] S2.1.1: Thermal environment simulation of the plot: randomly generate a batch of empty plots, extract the geometric features 1 of each plot, and give the target FAR of each plot, then generate the site status layout based on the geometric features 1 of the plot, and perform UTCI simulation on each plot to obtain the initial thermal environment state 1.

[0137] In this embodiment, first, a batch of random quadrilateral empty plots is generated by Rhino software, and the geometric features 1 of each plot are extracted. Then, based on the geometric features 1 of the plot, the layout of the site status is generated, indicating the location of the buildable buildings within the plot. Then, a 256x256 pixel grayscale image is generated using Rhino, and the height of the building is mapped by the depth of the color.

[0138] Then, based on the historical meteorological data, UTCI simulation is performed on the generated empty plots to obtain the initial thermal environment state 1;

[0139] In this embodiment, the historical meteorological data is mainly provided by the online meteorological database of the research area, and the required meteorological data is exported in the ".epw" format to ensure that the meteorological data of the research area can be reflected. UTCI simulation is performed using the Ladybug plugin in the Grasshopper parameterization platform of Rhino, and the grid size of the simulation calculation is set to 5m x 5m, and the monitoring height is 1.8m, to capture the urban microclimate changes in detail. All input parameters are set according to the standard of the annual regular climate data of the research area to ensure the accuracy of the simulation results.

[0140] Finally, the state space s arch of the plot is composed of the geometric features 1 of the plot, the target FAR of the plot {target} , the geometric parameters of the building, and the thermal environment state 1, which serves as the input for subsequent Agent training.

[0141] S2.1.2: Based on the current state space s archAnd the current layout of the site, generate the output action distribution parameters of the building agent through the Actor network, and obtain the current strategy π θ (a arch |s arch );

[0142] Figure 3 The Actor network structure diagram of the building agent in this embodiment is shown below. The Actor network parameters θ of the building agent are initialized. arch , and input the state space s of the building agent generated in S2.1.1 arch , including the geometric characteristics of the plot 1, the target volume ratio FAR of the plot {target} , the geometric parameters of the building and the thermal environment state 1. Initialize θ using the Xavier method arch The weights of the Actor network are initialized to 0 (the Xavier method is a common initialization method to ensure that the initial output of the network is not too large or too small), and the bias is initialized to 0. The input layer of the Actor network receives the land state s arch The current layout map of the site is processed using a convolutional neural network (CNN) with a dimension of 256×256×1. Then, spatial features are extracted through three convolutional layers. The convolution kernel size of each layer is 3×3, the stride is 2, and the number of output channels is 16, 32, and 64 respectively. Each layer is followed by a ReLU activation function, and finally a fully connected layer is used to output the plot image feature 1×256; the geometric parameters of the buildings are processed using a long short-term memory network (LSTM), with 128 hidden units set to process the 6-dimensional parameters of each building step by time, and output the hidden state T of the last time step. n 1×128; Thermal environment state 1 (dimension 1×6, including [Comfort {ratio} ,ArchOverlap {ratio} ,FAR {current} ,UTCI max ,Boundary {out} 1, Size {out} 1]), through the fully connected layer, the output dimension is 1×64. Then the current layout of the site (1×256), LSTM hidden state T n (1×128) is concatenated with the thermal environment state 1 (1×64), and the output dimension is 1×448. Next, the feature vector enters two layers of fully connected layers with an input dimension of 1×448 and an output dimension of 1×128, and is activated by ReLU. Finally, the output layer of the network generates the action distribution parameter π θ (a arch |s arch), where the mean μ represents the expected value of the action ("most probable action"), is mapped from input 1 x 128 to output 1 x 6 by a fully connected layer and scaled to the action range by a tanh activation function; the standard deviation σ represents the variance of the action distribution ("exploration range of the action"), is also mapped from input 1 x 128 to output 1 x 6 by a fully connected layer and ensured to be positive by a softplus activation function. The actor network outputs the action distribution parameter π θ (a arch |s arch ), where π θ (a arch |s arch ) is the current policy, defined as the probability distribution of action a arch at time t given state s arch , a arch is sampled from the normal distribution N(μ, σ). The action a arch is passed to the environment, which updates the plot layout, i.e. adds new buildings or moves existing buildings based on the site status layout, re-performs the UTCI simulation (grid size 5 m x 5 m, monitoring height 1.8 m), calculates the reward r arch and generates the next state s arch+1 (including updated geometric features 1, thermal environment state 1 and site status layout), and saves (s arch , a arch , r arch , s arch+1 ) to the experience storage module, and repeats the process until the maximum decision step number N is collected or the preset termination condition is met.

[0143] S2.1.3: Based on the reward function of the building agent, the critic network predicts the state value V φ (s arch ) and calculates the advantage of the action A arch ;

[0144] Figure 4 is the Critic network structure diagram in the building agent of the embodiment. The Critic network parameters φ arch of the building agent are initialized, and the stored experience data is used, the Critic network predicts the state value, and calculates the advantage of the action. Traverse each data (s arch , a arch , r arch , s arch+1 ) in the experience storage module, and the state s arch and s arch+1Input Critic network. The structure of the Critic network is the same as that of the Actor network, containing a multi-branch structure, and after processing by the fusion layer, a feature vector with a dimension of 1x448 is obtained. Then, the feature vector enters two fully connected layers with an input dimension of 1x448 and an output dimension of 1x128, and is activated by ReLU. Finally, the input 1x128 is mapped to the output 1x1 through the fully connected layer to generate the state value V φ (s arch ), where V φ (s arch ) is the value predicted by the Critic network representing the expected value of the future discounted reward that can be obtained from the state s arch according to the current policy, and V φ (s arch+1 ) is also calculated.

[0145] The cumulative value R arch is calculated, which is the training target of the Critic network, representing the current reward plus the discounted value of the next state, and the formula is as follows:

[0146] R arch =r arch +γV φ (r arch+1 )

[0147] Where γ is the discount factor (representing the importance of future rewards, the closer the value is to 1, the more the Agent focuses on long-term earnings), and in this embodiment, γ = 0.99.

[0148] The time difference error δ t is calculated, which represents the error of the Critic network prediction, and is used to calculate the advantage, and the formula is as follows:

[0149] δ t =r arch +γV φ (s arch+1 )-V φ (s arch )

[0150] Where δ t is the time difference error (TD error), representing the difference between the predicted value at time step t and the actual return.

[0151] The advantage A arch is calculated, which represents the relative goodness of the action, and the larger the value, the better the action, which is used to guide the Actor network to improve the policy, and the formula is as follows:

[0152]

[0153] Where A archis the advantage function, which means that in state s arch Next select action a arch Compared with the average level; λ is the GAE parameter used to calculate the advantage of the action, in this embodiment λ = 0.95; q is a non-negative integer (q = 0, 1, 2, ..., ∞), representing the offset from the current time step t to the future q step.

[0154] S2.1.4: Based on Advantage A arch 、Status value V φ (s arch ) and strategy π θ (a arch |s arch ), calculate the total loss L through the loss function;

[0155] Use Advantage A arch 、Status value V φ (s arch ) and strategy π θ (a arch |s arch ), calculate the loss function and prepare for network update. Calculate the policy gradient target L CLIP , used to measure the improvement direction of the strategy and ensure the stability of the update. The formula is as follows:

[0156]

[0157] Among them, π θ (a arch |s arch ) is the current strategy; For the old strategy; arch (θ) is the probability ratio of the new and old strategies, indicating the magnitude of change in the current strategy; ∈ is a trimming parameter, and in this embodiment, ∈=0.2.

[0158] Calculate the value loss L VF , which is used to optimize the Critic network to make it more accurately predict the state value. The formula is as follows:

[0159]

[0160] in, Indicates expectation, R arch is the cumulative value, V φ (s arch ) is the value predicted by Critic.

[0161] Calculate the entropy loss S[π θ ](s arch), which represents the randomness of the policy (the greater the entropy value, the more random the policy, the more the Agent tends to explore, the total loss L combines policy improvement, value prediction and exploration, and is used to guide network parameter updates, as follows:

[0162] L = L CLIP - c1L VF + c2S[πθ](s arch )

[0163] wherein c1 is a value loss weight (the greater the value, the greater the training intensity of the Critic network); c2 is an entropy loss weight (the greater the value, the stronger the exploration of the Agent).

[0164] S2.1.5: Update the Actor network parameters θ arch and the Critic network parameters θ arch using the Adam optimizer according to the total loss L;

[0165] The Actor network parameters θ arch are updated using the Adam optimizer in the present embodiment. Since θ arch determines the policy of the Actor network, the updated Agent can select better actions, as follows:

[0166]

[0167] wherein η is the learning rate, and η = 0.0003 in the present embodiment; is the gradient of the loss with respect to θ arch , which reflects the direction and amplitude of adjustment of the parameter θ arch .

[0168] The Critic network parameters φ arch are updated. Since φ arch determines the value prediction of the Critic network, the updated Critic can more accurately evaluate actions, as follows:

[0169]

[0170] wherein, is the gradient of the loss with respect to φ arch , which reflects the direction and amplitude of adjustment of the parameter φ arch .

[0171] After updating the network parameters, the experience storage module is emptied, and the next round of experience collection is prepared.

[0172] S2.1.6: Repeat the training until the maximum number of rounds is reached

[0173] Repeat steps S2.1.2 to S2.1.5 to iteratively update θ arch and θ arch until a preset maximum number of training rounds (1000 rounds in this embodiment) is reached. The final θ arch and φ arch are the final parameters of the Actor and Critic networks of the building Agent, and θ arch of the Actor network corresponds to the optimal policy π θ , which is the optimal building layout strategy that meets the target plot ratio and thermal environment requirements. After training, the Actor network parameters θ arch are saved for subsequent generation of building layouts.

[0174] S2.2: Training of tree Agent using proximal policy optimization algorithm

[0175] The proximal policy optimization (PPO) algorithm is based on the Actor-Critic framework, which uses a policy network (Actor) to select actions and a value network (Critic) to evaluate action values, and uses a double-network collaborative optimization mechanism to achieve stable control of policy updates. The PPO algorithm controls the amplitude of policy updates to ensure stability during training and ultimately generates the optimal tree layout scheme. The following are the specific steps for training the tree Agent using the PPO algorithm, including all parameters and operation procedures that need to be set.

[0176] S2.2.1: Thermal environment simulation of the plot: extract the geometric features 2 of the plot and the geometric parameters of the buildings after optimization in S2.1, and give the target green rate, then generate the site status layout based on the geometric features 2 of the plot and the geometric parameters of the buildings, and perform UTCI simulation on each plot.

[0177] In this embodiment, first, the geometric features 2 of the plot and the geometric parameters of the buildings are modeled using Rhino software according to S2.2. Then, based on the geometric features 2 of the plot and the geometric parameters of the buildings, a layout map of the site status is generated to represent the locations within the plot where trees can be arranged. Then, a 256x256 pixel grayscale image is generated using Rhino, and the height of the trees is mapped by the depth of the color.

[0178] Then, based on historical meteorological data, UTCI simulation is performed on the generated plot to obtain the initial thermal environment state 2;

[0179] In this embodiment, historical meteorological data is mainly provided by studying the local online meteorological database of the study area, and the required meteorological data is exported in the ".epw" format to ensure that the meteorological conditions of the study area can be reflected. Ladybug plug-in in Rhino's Grasshopper parameterization platform is used for UTCI simulation, the grid size of simulation calculation is set to 5 meters x 5 meters, and the monitoring height is 1.8 meters, so as to capture the urban microclimate change in detail, and all input parameters are set according to the standard of the annual regular climate data of the study area to ensure the accuracy of the simulation results.

[0180] Finally, the plot state space is composed of the geometric characteristics 2 of the plot, the target FAR of the plot FAR {target} , the geometric parameters of the building, and the thermal environment state 2, which are used as the input of the subsequent Agent training.

[0181] S2.2.2: Based on the current tree Agent state space s tree and the site status layout, the output action distribution parameters of the tree Agent are generated through the Actor network to obtain the current strategy π θ (a tree |s tree );

[0182] The Actor network parameters θ tree of the tree Agent are initialized, and the state space s tree generated based on S2.2.1 is input, which includes the geometric characteristics 2, the target FAR of the plot FAR {target} , the geometric parameters of the tree, and the thermal environment state 2. The weights of θ tree are initialized using the Xavier method (the Xavier method is a common initialization method to ensure that the initial output of the network is not too large or too small), and the bias is initialized to 0. The input layer of the Actor network receives the plot state s tree , the site status layout is processed using a convolutional neural network (CNN), and the dimension is 256x256x1. Then, spatial features are extracted through three convolutional layers, the convolution kernel size of each layer is 3x3, the stride is 2, the output channel number is 16, 32, and 64 respectively, each layer is followed by a ReLU activation function, and finally a fully connected layer is used to output the plot image features 1x256; the geometric parameters of the tree are processed using a long short-term memory network (LSTM), 128 hidden units are set, and the 4-dimensional parameters of each tree are processed step by step, and the hidden state T n of the last time step is output as 1x128; the thermal environment state 2 (dimension 1x6, including

[0183] [Comfort {ratio} ,ArchOverLap {ratio} ,FAR {current} ,UTCI max ,Boundary {out} ,Size {out} ]) through a fully connected layer with output dimension 1x64. Then the site status layout (1x256), LSTM hidden state T n (1x128) and thermal environment state (1x64) are concatenated, with output dimension 1x448. Next, the feature vector goes through two fully connected layers with input dimension 1x448 and output dimension 1x128, activated by ReLU. Finally, the output layer of the network generates the action distribution parameter π θ (a tree |s tree ), where the mean μ represents the expected value of the action (“the most likely action”), mapped from input 1x128 to output 1x6 through a fully connected layer and scaled to the action range by a tanh activation function; the standard deviation σ represents the variance of the action distribution (“the exploration range of the action”), also mapped from input 1x128 to output 1x6 through a fully connected layer and ensured to be positive by a softplus activation function. The actor network outputs the action distribution parameter π θ (a tree |s tree ), where π θ (a tree |s tree ) is the current policy, defined as the probability distribution of action a tree given state s tree , from which action a tree is sampled from a normal distribution N(μ,σ). The action a tree is passed to the environment, which updates the plot layout (adding new trees or moving existing trees on the site status layout), re-performs the UTCI simulation (simulation parameters are consistent with S2.2.1, grid size is 5m x 5m, monitoring height is 1.8m), calculates the reward r tree and generates the next state s arch+1 (including updated geometric features, thermal environment state and site status layout), and saves (s tree , a tree , r tree , s tree+1 ) to the experience storage module, and repeats this process until the maximum decision step number N is collected or the preset termination condition is met.

[0184] S2.2.3: Reward function based on tree agent, state value V predicted by Critic networkφ (s tree ) and calculate the advantage of action A tree ;

[0185] Initialize the Critic network parameters φ of the tree Agent tree , and use the stored experience data, the Critic network predicts the state value, and calculates the advantage of action. Traverse each data (s tree ,a tree ,r tree ,s tree+1 ) in the experience storage module, input the state s tree and s tree+1 into the Critic network. The structure of the Critic network is the same as that of the Actor network, which contains a multi-branch structure, and after processing by the fusion layer, a feature vector (dimension 1x448) is obtained. Then, the feature vector enters two fully connected layers, with an input dimension of 1x448 and an output dimension of 1x128, and is activated by ReLU. Finally, the input 1x128 is mapped to the output 1x1 through the fully connected layer to generate the state value V φ (s tree ), where V φ (s tree ) is the value predicted by the Critic network (representing the expected value of the future discounted reward that can be obtained starting from state s tree according to the current policy), and V φ (s tree+1 ) is also calculated.

[0186] Calculate the cumulative value R arch , which is the training target of the Critic network, representing the current reward plus the discounted value of the next state, and the formula is as follows:

[0187] R tree =r tree +γV φ (r tree+1 )

[0188] Where γ is the discount factor (representing the importance of future rewards, the closer the value is to 1, the more the Agent focuses on long-term earnings), and in this embodiment, γ = 0.99.

[0189] Calculate the time difference error δ t , which represents the error of the Critic network prediction, used to calculate the advantage, and the formula is as follows:

[0190] δ t =r tree +γV φ (s tree+1 )-V φ (s tree)

[0191] Among them, δ t is the temporal difference error (TD error), which represents the difference between the predicted value and the actual return at time step t.

[0192] Computational Advantage A tree , which indicates the relative quality of the action. The larger the value, the better the action. It is used to guide the Actor network improvement strategy. The formula is as follows:

[0193]

[0194] Among them, A tree is the advantage function, which means that in state s tree Next select action a tree Compared with the average level, λ is the advantage of the GAE parameter used to calculate the action. In this embodiment, λ = 0.95; q is a non-negative integer (k = 0, 1, 2, ..., ∞), which represents the offset from the current time step t to the future q steps.

[0195] S2.2.4: Based on Advantage A tree 、Status value V φ (s tree ) and strategy π θ (a tree |s tree ), calculate the total loss L' through the loss function;

[0196] like Figure 2 As shown, using advantage A tree 、Status value V φ (s tree ) and strategy π θ (a tree |s tree ), calculate the loss function and prepare for network update. Calculate the policy gradient target L CLIP , used to measure the improvement direction of the strategy and ensure the stability of the update. The formula is as follows:

[0197]

[0198] Among them, π θ (a tree |s tree ) is the current strategy, For the old strategy, r arch (θ) is the probability ratio of the new and old strategies, indicating the magnitude of change in the current strategy, and ∈ is a trimming parameter. In this embodiment, ∈=0.2.

[0199] Calculate the value loss L VF, to optimize the Critic network to make it more accurate in predicting state values, as follows:

[0200]

[0201] where, represents expectation, R tree is cumulative value, V φ (s tree ) is the value predicted by the Critic.

[0202] Calculate the entropy loss S[π θ ](s tree ), which represents the randomness of the policy. The greater the entropy value, the more random the policy, and the more the Agent tends to explore. The total loss l' combines policy improvement, value prediction, and exploration, which is used to guide network parameter updates, as follows:

[0203] l' = L CLIP -c1L VF +c2S[π θ ](s tree )

[0204] where c1 is the value loss weight (the greater this value, the greater the training intensity of the Critic network), and c2 is the entropy loss weight (the greater this value, the stronger the exploration of the Agent).

[0205] S2.2.5: Update the Actor network parameters θ tree and the Critic network parameters φ tree using the Adam optimizer according to the total loss L';

[0206] In this embodiment, the Adam optimizer is used to update the Actor network parameters θ tree . Because θ tree determines the policy of the Actor network, the Agent can select better tree actions after updating, as follows:

[0207]

[0208] where η is the learning rate, and in this embodiment, η = 0.0003; is the gradient of the loss with respect to θ tree , which reflects the direction and amplitude of the adjustment of the parameters θ tree .

[0209] Update the Critic network parameters φ tree . Because φ tree determines the value prediction of the Critic network, the Critic can more accurately evaluate actions after updating, as follows:

[0210]

[0211] where, is the gradient of the loss with respect to φ tree , reflecting the direction and magnitude of the adjustment of parameter φ tree .

[0212] After updating the network parameters, the experience storage module is emptied, preparing for the next round of experience collection.

[0213] S2.2.6: Repeat training until the maximum number of rounds is reached

[0214] Repeat steps S2.2.2 to S2.2.6, iteratively update θ tree and φ tree , until the maximum number of training rounds is reached (set to 1000 in this embodiment). The final θ tree and φ tree are the final parameters of the Actor and Critic networks of the tree Agent, and the θ tree of the Actor network corresponds to the optimal policy π θ , which is the optimal tree layout strategy that meets the target FAR and thermal environment requirements. After training, save the Actor network parameters θ tree for subsequent tree layout generation.

[0215] S3: Apply building Agent and tree Agent to cooperatively generate optimal layout

[0216] This step uses the building Agent and tree Agent trained in S2.1 and S2.2 to cooperatively generate optimal building layout and tree layout on the same plot, dynamically balancing the target FAR and thermal environment performance requirements. Building Agent and tree Agent alternate actions, gradually optimizing the layout until the termination condition is met. The following are the specific steps, including operation process and result evaluation.

[0217] S3.1: Input the state and target parameters of the plot to be optimized.

[0218] In this step, the state space of the plot to be optimized, the target FAR {target} , and the target tree coverage Cover {target} are input to provide initial conditions for layout optimization. The plot state space is generated based on S2.1.1, including the geometric characteristics 1, thermal environment state 1, and site status layout of the plot. At the same time, the Actor network parameters of the building Agent trained in S2.1 are loaded.

[0219] S3.2: Obtain the state space s of the plot to be optimized according to the method of S2.2.1tree , including geometric features 2, thermal environment state 2, and site status layout map; set the target tree coverage rate Cover of the plot {target} ; load the trained tree Agent actor network parameters θ tree ;

[0220] Next, the building Agent and the tree Agent alternately perform actions to cooperatively optimize the building layout and the tree layout of the plot. In each iteration, the building layout is first generated or adjusted by the building Agent, and then the tree layout is generated or adjusted by the tree Agent.

[0221] S3.3: Take the current geometric features 1, thermal environment state 1, and site status layout map of the plot as input, and apply the trained building Agent to optimize the building layout of the plot with the target volume rate and preliminary thermal environment requirements as constraints, output the current optimal building layout of the plot and updated geometric features 1 and thermal environment state 1;

[0222] In this step, the trained building Agent in S2.1 is used to optimize the building layout of the plot to meet the target volume rate and preliminary thermal environment requirements. Take s arch (contains geometric features 1, thermal environment state 1, and site status layout map) as input to generate an optimized building layout until the preset termination condition is met. Output the updated geometric features 1 and thermal environment state 1 for evaluating thermal comfort and building layout effect.

[0223] S3.4: Take the current geometric features 2, thermal environment state 2, and site status layout map of the plot as input, and apply the trained tree Agent to optimize the tree layout of the plot based on the building layout generated in S3.3 with the target tree coverage rate and preliminary thermal environment requirements as constraints, output the current optimal tree layout of the plot and updated geometric features 2 and thermal environment state 2;

[0224] In this step, the trained tree Agent in S2.2 is used to optimize the tree layout based on the building layout generated in S3.3 to further improve the thermal environment performance. Take s tree (contains geometric features 2, thermal environment features 2, and site status layout map) as input to generate an optimized tree layout until the preset termination condition is met. Output the updated geometric features 2 and thermal environment state 2 for evaluating thermal comfort and tree layout effect.

[0225] S3.5: Repeat S3.3 to S3.4 until the preset termination condition is met;

[0226] S3.6: Output the final optimized layout of the plot to be optimized and its thermal environment state.

[0227] In this step, according to the target volume rate and the indicators in the building agent reward function, the building geometric characteristics 1 and the thermal environment state 1 of the best building layout scheme in each indicator are output. Then according to the target tree coverage rate and the indicators in the tree agent reward function, the tree geometric characteristics 2 and the thermal environment state 2 of the best tree layout scheme in each indicator are output for the user to use for scheme screening or further optimization of the arrangement scheme.

[0228] It should be understood that, under the inspiration of the technical concept of the present application, those skilled in the art can make various improvements or changes according to the above content without departing from the content of the present application, which still falls within the protection scope of the present application.

Claims

1. A high-intensity area thermal environment performance space design decision method based on reinforcement learning, characterized by: The method comprises the following steps: Define the state space, action space, and reward function of the building agent and the state space, action space, and reward function of the tree agent respectively; The building agent is trained using the proximal policy optimization algorithm. Based on the training results of the building agent, the tree agent is trained using the proximal policy optimization algorithm. Use the trained building agent and the trained tree agent to collaboratively generate the optimal layout of buildings and trees in the area to be optimized; The state space of the building agent represents the input information of the building agent, including the geometric features of the plot 1, the target volume ratio FAR of the plot {target} , geometric parameters of the building and thermal environment status 1; The geometric features 1 of the plot include: the coordinates of each vertex of the plot boundary and the plot area; the geometric parameters of the building include: the coordinates of the center point of each building bottom surface (x i ,y i ), length l i 、Width w i 、Building height h i and the orientation angle θ i If n buildings have been placed in the plot, the geometric parameters of the buildings are expressed as [x1,y1,l1,w1,h1,θ1,...,x n ,y n ,l n ,w n ,h n ,θ n ]; The thermal environment state 1 includes the following parameters: Comfort area ratio Comfort {ratio} , ArchOverlap {ratio} 、Current Floor Area Ratio FAR {current} UTCI maximum value UTCI max , Building Boundary Violation Index {out} 1. Building size violation indicator Size {out} 1; The state space of the tree agent represents the input information of the tree agent, including the geometric features of the plot 2, the target tree coverage rate Cover {current} , geometric parameters of trees and thermal environment status 2; the geometric features 2 of the plot include: the coordinates of each vertex of the plot boundary, the floor area of ​​the building placed in the plot, and the geometric parameters of the building; the geometric parameters of the trees include the coordinates of each tree (x j ,y j ), height h j and crown diameter d j If m trees have been placed in the plot, the geometric parameters of the trees are expressed as [x1,y1,h1,d1,...,x m ,y m ,h m ,d m ]; The thermal environment state 2 includes the following parameters: Comfort area ratio Comfort {ratio} , TreeOverlap {ratio} 、Current tree coverage {current} UTCI maximum value UTCI max , Boundary {out} 2. Tree size violation indicator Size {out} 2; The action space of the building agent represents the output action of the building agent, including new building placement and existing building adjustment; the new building placement refers to placing a new building in the plot; the existing building adjustment refers to specifying the adjustment amount for the i-th existing building: [Δx i ,Δy i ,Δl i ,Δw i ,Δh i ,Δθ i ], where Δx i ,Ay i is the coordinate change of the center point of the bottom surface of the building; Δl i ,Δw i is the length and width change of the building bottom surface; Δθ i The building's orientation angle changes; The action space of the tree agent represents the output action of the tree agent, including new tree placement and existing tree adjustment; the new tree placement refers to placing new trees in the plot; the existing tree adjustment is the specified adjustment amount of the j-th existing tree: [Δx j ,Δy j ,Δh j ,Δd j ], where Δx j ,Δy j is the coordinate change of the tree; Δh j is the height change of the tree; Δd j is the change in crown diameter of the tree.

2. The high-intensity area thermal environment performance space design decision method based on reinforcement learning according to claim 1 is characterized in that: The parameters of thermal environment state 1 and thermal environment state 2 are obtained through UTCI simulation calculation: (1) Comfort area ratio {ratio} , which represents the ratio of the area within the grid to the area of ​​the plot when the UTCI is 9 to 26: Where K is the total number of grids; UTCI k is the UTCI value of the kth grid; I(9≤UTCI k <26) is the indicator function, if UTCI k In the range of 9 to 26, I (9 ≤ UTCI k <26) is 1, otherwise I(9≤UTCI k <26) is 0; A grid is the area of ​​a single grid, A plot is the total area of ​​the plot; (2) ArchOverlap {ratio} : Where A Overlap,i,j A represents the overlapping bottom area of ​​the i-th building and the j-th building; i ,A j are the base areas of the i-th building and the j-th building respectively; I Overlap,i is the indicator function. If Ratio exceeds 0.6, I Overlap,i =1, otherwise I Overlap,i =0; N is the total number of buildings in the current plot; (3) Current Floor Area Ratio (FAR) {current} : Where, F i is the floor height; is the building area of ​​the i-th building; A plot is the site area; (4) UTCI maximum value UTCI max , the highest UTCI value in the grid: Where, UTCI k is the UTCI value of the kth grid obtained through UTCI simulation; (5) Building boundary violation indicator Boundary {out} 1. Calculate the ratio of the number of buildings that extend beyond the boundary of the plot to the total number of buildings: Where, I boundary,i is an exponential function. If the i-th building exceeds the boundary of the plot, I boundary,i =1, otherwise I boundary,i =0; N is the number of buildings in the current plot; (6) Building size violation indicator Size {out} 1. Calculate the ratio of the number of buildings exceeding the size limit to the total number of buildings: Where, I size,i is an exponential function, if the length l of the i-th building i or width w i If it is not within the preset value range, I size,i =1, otherwise I size,i =0; (7) TreeOverlap {ratio} : In the formula, for the newly placed i-th tree, check whether its crown overlaps with the existing crowns; A Overlap,i,j represents the overlapping base area between the i-th tree and the j-th tree; A i ,A j are the base areas of the i-th tree and the j-th tree respectively; I Overlap,i is the indicator function. If Ratio exceeds 0.6, I Overlap,i =1, otherwise I Overlap,i =0;A Overlap,i,j is the overlapping area between the i-th tree and the j-th tree; M is the total number of trees in the current plot; (8) Current tree cover {current} : Where, d i is the crown diameter of the jth tree; l i w i is the building area of ​​the i-th building, M is the total number of trees in the current plot, and N is the total number of buildings in the current plot; (9) Boundary {out} 2. Count the ratio of the number of trees that extend beyond the boundary of the plot to the total number of trees: Where, I boundary,i is an exponential function. If the i-th tree exceeds the boundary of the plot, I boundary,i =1, otherwise I boundary,i =0; (10) Tree size violation indicator Size {out} 2. Calculate the ratio of the number of trees exceeding the size limit to the total number of trees: Where, I size,i is an exponential function, if the crown diameter d of the i-th tree i Or the tree height h i If it is not within the preset value range, I size,i =1, otherwise I size,i =0.

3. The high-intensity area thermal environment performance space design decision method based on reinforcement learning according to claim 2 is characterized in that: The reward function r of the building agent arch for: Where, ω1·Comfort {ratio} Encourage maximization of comfort area ratio; ω2·ArchOverlap {ratio} Penalize new buildings that overlap with existing buildings by more than 60%; Penalty for floor area ratio deviation; ω4· max (0,UTCI max -26) Penalty UTCI Maximum UTCI max More than 26; ω5·Boundary {out}1 Penalize buildings that exceed the boundary; ω6·Size {out} 1 is the penalty size exceeding the limit; ω is the weight of different rewards; The reward function r of the tree agent tree for: Where, τ1·Comfort {ratio} Encourage maximization of the comfort area ratio; τ2·TreeOverlap {ratio} Penalize cases where new trees overlap with existing trees by more than 60%; Penalize number coverage deviation; τ4· max (0,UTCI max -26) Penalty UTCI Maximum UTCI max More than 26; τ5·Boundary {out} 2 Penalize trees beyond the boundary; τ6·Size {out} 2 penalizes tree size exceeding the limit; τ is the weight of different rewards.

4. The high-intensity area thermal environment performance space design decision method based on reinforcement learning according to claim 3 is characterized in that: The method of training the building agent using the proximal strategy optimization algorithm includes the following steps: S2.1.1: Perform thermal environment simulation on the plots: Randomly generate a batch of empty plots, extract the geometric features 1 of each plot and assign a target floor area ratio to each plot. Then, based on the geometric features 1 of the plots, generate a site layout diagram that represents the locations of buildings that can be arranged within the plots. Furthermore, perform UTCI simulation on each plot based on historical meteorological data to obtain an initial thermal environment state 1. S2.1.2: Based on the state space s of the current building agent arch And the current layout of the site, generate the output action distribution parameters of the building agent through the Actor network, and obtain the current strategy π θ (a arch |s arch ); S2.1.3: Based on the reward function of the building agent, the state value V is predicted through the critic network φ (s arch ) and calculate the advantage A of the action arch ; S2.1.4: Based on Advantage A arch 、Status value V φ (s arch ) and strategy π θ (a arch |s arch ), calculate the total loss L through the loss function; S2.1.5: Update the Actor network parameters θ using the Adam optimizer based on the total loss L arch and Critic network parameter φ arch ; S2.1.6: Repeat steps S2.1.2 to S2.1.5 to iteratively update θ arch and φ arch , until the preset maximum number of training rounds is reached, thereby obtaining the final parameters of the Actor network of the building agent and the final parameters of the Critic network.

5. The high-intensity area thermal environment performance space design decision method based on reinforcement learning according to claim 4 is characterized in that: The method of training the tree agent using the proximal strategy optimization algorithm includes the following steps: S2.2.1: Perform thermal environment simulation on the plot: Extract the geometric features of the plot after the building agent training is completed2, the geometric parameters of the building, and set the target greening ratio. Then, based on the geometric features2 of the plot and the geometric parameters of the building, generate a site layout map that represents the locations where trees can be placed within the plot. UTCI simulation is performed on each plot based on historical meteorological data to obtain the initial thermal environment state2. S2.2.2: Based on the state space s of the current tree agent tree And the current layout of the site, generate the output action distribution parameters of the tree agent through the Actor network, and obtain the current strategy π θ (a tree |s tree ); S2.2.3: Based on the reward function of the tree agent, the state value V is predicted through the critic network φ (s tree ) and calculate the advantage A of the action tree ; S2.2.4: Based on Advantage A tree 、Status value V φ (s tree ) and strategy π θ (a tree |s tree ), calculate the total loss L' through the loss function; S2.2.5: Update the Actor network parameters θ using the Adam optimizer based on the total loss L' tree and Critic network parameter φ tree ; S2.2.6: Repeat S2.2.2 to S2.2.5, iteratively updating θ tree and φ tree , until the preset maximum number of training rounds is reached, thereby obtaining the final parameters of the Actor network of the building agent and the final parameters of the Critic network.

6. The high-intensity area thermal environment performance space design decision method based on reinforcement learning according to claim 5 is characterized in that: The method of using the trained building agent and the trained tree agent to collaboratively generate the optimal layout of buildings and trees in the design decision area includes the following steps: S3.1: Obtain the state space s of the area to be optimized, i.e., the land parcel, according to the method in S2.1.1 arch , including the geometric characteristics of the plot 1, thermal environment status 1, and the current layout of the site; set the target volume ratio FAR of the area {target} ; Load the trained building agent's Actor network parameters θ arch ; S3.2: Obtain the state space s of the area to be optimized, i.e., the land parcel, according to the method in S2.2.1 tree , including the geometric characteristics of the plot 2, the thermal environment status 2, the current layout of the site; set the target tree coverage rate Cover {target} ; Load the trained tree agent's Actor network parameters θ tree ; S3.3: Taking the current plot's geometric features 1, thermal environment status 1, and the current site layout as input, and subject to the target volume ratio and preliminary thermal environment requirements as constraints, the trained building agent is applied to optimize the area's building layout, outputting the area's current optimal building layout along with the updated geometric features 1 and thermal environment status 1. S3.4: Taking the current plot's geometric features 2, thermal environment status 2, and the current site layout as input, and subject to the target tree coverage and preliminary thermal environment requirements as constraints, the trained tree agent is applied to optimize the tree layout of the area based on the building layout generated in S3.

3. The trained tree agent then outputs the optimal tree layout for the area, along with the updated geometric features 2 and thermal environment status 2. S3.5: Repeat S3.3 to S3.4 until the preset termination condition is met; S3.6: Output the final optimized layout of the area to be optimized, i.e., the plot, and its thermal environment status: Based on the target floor area ratio and the indicators in the reward function of the building agent, output the geometric features 1 and thermal environment status 1 corresponding to the best building layout scheme among each indicator; and based on the target tree coverage rate and the indicators in the reward function of the tree agent, output the geometric features 2 and thermal environment status 2 corresponding to the best tree layout scheme among each indicator.

7. The high-intensity area thermal environment performance space design decision method based on reinforcement learning according to any one of claims 4 to 6, characterized in that: The current site layout map is a 256×256 pixel grayscale image, and the height of buildings and / or trees is mapped by the depth of color.