Green building design multi-objective optimization method and system based on deep reinforcement learning
Through deep reinforcement learning technology, deep neural networks and DDPG algorithms based on Bayesian optimization are constructed to optimize green building design parameters, solving the problems of low efficiency and poor adaptability in the existing technology, achieving multi-objective optimization of building performance, and improving energy efficiency and environmental sustainability.
Patent Information
- Application Number
- CN202510657278.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-29
AI Technical Summary
The existing green building design methods are inefficient and poorly adaptable, making it difficult to meet the dual-carbon targets of energy consumption and carbon emissions, and the traditional multi-objective optimization algorithm is inefficient and poorly adaptable.
Using deep reinforcement learning technology, the architectural design parameters are optimized and multi-objective optimization is achieved by constructing a deep neural network predictive meta model and a deep deterministic strategy gradient (DDPG) algorithm based on Bayesian optimization.
It significantly improves the energy efficiency, environmental sustainability and living comfort of the building. The model stably outputs reliable solutions in complex environments, has strong adaptability and reliability, and has performed well in long-term and multi-regional testing.
Smart Images

Figure CN120387222A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of architectural design and artificial intelligence technology, and particularly to a multi-objective optimization method and system for green building design based on deep reinforcement learning. Background Art
[0002] Under the background of global sustainable development, the green development of the construction industry is of great importance. With the increase in newly built buildings in China, energy consumption and carbon emissions are prominent. Only 4% of existing buildings implement energy-saving measures, posing challenges to the "dual carbon" goal. It is urgent to practice "green buildings". Green building design is of great significance. Although there are evaluation standards at home and abroad, it has high manpower requirements and relies on experience, with deficiencies. BIM technology has brought opportunities to green building design, but due to large amounts of data, limitations of traditional methods, and defects in machine learning models, it is difficult to meet the requirements. Classical multi-objective optimization algorithms also have problems such as low efficiency and poor adaptability. Developing an efficient and intelligent green building design optimization method has become a research hotspot.
[0003] Therefore, there is an urgent need for an efficient and intelligent green building design optimization scheme at present. Summary of the Invention
[0004] The present disclosure provides a multi-objective optimization method and system for green building design based on deep reinforcement learning, which realizes the optimization of multiple objectives in green building design by using deep reinforcement learning technology, and at least solves the technical problems of low efficiency and poor adaptability of existing design methods.
[0005] According to the first aspect of the present disclosure, there is provided a multi-objective optimization method for green building design based on deep reinforcement learning, including the following steps: Determine building performance objectives and related design parameters and construct an evaluation index system, establish a 3D BIM model and import it into a simulation software based on BIM technology and building performance, so as to obtain a sample data set; Construct a deep neural network prediction meta-model based on Bayesian optimization, and use the sample data set to train the meta-model to learn the non-linear relationship between the building performance objectives and related design parameters and building performance, and output the predicted values of the building performance objectives; Construct a framework based on DRL, and optimize the meta-model based on the framework, and optimize the building performance through iterative learning.
[0006] In the above aspect and any possible implementation manner, a further implementation manner is provided, where the evaluation index system includes expected objectives and influencing variables; The expected objectives include building energy consumption , carbon emissions , indoor thermal discomfort ; The influencing variables include: building envelope structure, building openings, and HVAC parameters; Among them, the building envelope structure includes the thickness of the wall insulation layer, the thickness of the roof insulation layer, and the building airtightness; The building openings include the solar heat gain coefficient of the windows, the window-to-wall area ratio, and the opening rate of the exterior wall windows; The HVAC parameters include the coefficient of performance of the air conditioning system, the thermal efficiency of the heating system, the fresh air volume of the ventilation system, the indoor design temperature, the indoor lighting power density, and the equipment power density.
[0007] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. The process of establishing a 3D BIM model and importing a simulation software based on BIM technology and building performance to obtain a sample data set is as follows: Use Autodesk Revit 2021 to establish a 3D BIM model of the building, save it as a gbXML format file, and then import it into the DesignBuilder software for building performance simulation and orthogonal testing to obtain a sample data set containing building performance results.
[0008] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. The prediction meta-model includes an input layer, an output layer, and a hidden layer; The input layer takes the building performance objectives and relevant design parameters as inputs, the output layer outputs the predicted values of the building performance objectives, and the hidden layer extracts and transforms the features of the inputs through non-linear transformation.
[0009] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. The backpropagation function of the prediction meta-model is: ; where is the learning rate, is the objective function, and are both weight parameters of the deep neural network DNN. Among them, is the original parameter, is the parameter after update iteration.
[0010] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. The DRL-based framework includes engineering problem conceptualization and a DDPG model; Conceptualizing the engineering problem regards the setting of the building performance objectives and related design parameters as action A, defines the state S as the predicted value of the building performance objective, and the state update depends on the execution of the action. Among them, the building performance objective value S includes the expected target predicted value; The DDPG model includes an actor network and a critic network.
[0011] For the above aspects and any possible implementation manners, a further implementation manner is provided. The predicted value of the building performance objective is specifically: ; Among them, 、 、 is a DNN model for predicting based on a given action combination.
[0012] For the above aspects and any possible implementation manners, a further implementation manner is provided. When the actor network selects an action, random noise is added , to complete policy exploration. Specifically: ; Among them, represents the building performance at the current time t , and the action selection obtained through the actor network at this time; after adding Gaussian noise , the real action selection is obtained, that is .
[0013] For the above aspects and any possible implementation manners, a further implementation manner is provided. The optimization process of the DDPG model for building performance is as follows: During the training process, the actor network generates an action function according to the predicted value S of the building performance objective ; The critic network calculates the Q value according to the action function and the state, and updates the actor network parameters and the critic network parameters by using the Q value through the gradient method, so as to maximize the Q value, that is, to obtain the maximized expected return and complete the performance optimization process.
[0014] According to the second aspect of the present disclosure, a multi-objective optimization system for green building design based on deep reinforcement learning is provided, which is used to implement the multi-objective optimization method for green building design based on deep reinforcement learning as described in the first aspect, including: a sample data set acquisition module, a building performance prediction module, and a building performance optimization module; The sample data set acquisition module is used to determine building performance goals and related design parameters, construct an evaluation index system, establish a 3D BIM model, and import it into a simulation software based on BIM technology and building performance, so as to obtain a sample data set; The building performance prediction module is used to construct a deep neural network prediction meta-model based on Bayesian optimization, and use the sample data set to train the meta-model to learn the non-linear relationship between the building performance goals, related design parameters and building performance, and output the building performance prediction result; The building performance optimization module is used to construct a DRL-based framework, optimize the meta-model based on the framework, and optimize the building performance through iterative learning.
[0015] Compared with the prior art, the present invention has the following technical effects: 1. The present invention constructs an optimization framework based on deep reinforcement learning, with the DDPG algorithm at its core, which has significant advantages in the multi-objective optimization of green buildings and can greatly improve the comprehensive performance of buildings; 2. The present invention conducts a comprehensive robustness verification on the DDPG algorithm, covering a variety of complex environmental factors, and the model can stably output reliable solutions, with strong adaptability and reliability; 3. The present invention has been tested for a long time and in multiple regions, and the model has good stability and generalization, providing double support in theory and practice for green building design.
[0016] It should be understood that the content described in the summary of the invention is not intended to limit the key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings
[0017] Combined with the drawings and referring to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. In the drawings, the same or similar reference numerals represent the same or similar elements, where: Figure 1 Shows a schematic flowchart of a multi-objective optimization method for green building design based on deep reinforcement learning according to an embodiment of the present disclosure; Figure 2 Shows a schematic flowchart of a multi-objective optimization method for green building design based on deep reinforcement learning according to an embodiment of the present disclosure, covering the integration of BIM models and simulation tools, DNN performance prediction, and DDPG optimization; Figure 3 Shows a schematic diagram of an evaluation index system for a multi-objective optimization method for green building design based on deep reinforcement learning according to an embodiment of the present disclosure. Figure 4 Shows a schematic structural diagram of a meta-model for predicting based on a deep neural network for a multi-objective optimization method for green building design according to an embodiment of the present disclosure; Figure 5 Shows a schematic diagram of the model prediction results for a multi-objective optimization method for green building design based on deep reinforcement learning according to an embodiment of the present disclosure; Figure 6 Shows a schematic structural diagram of the framework of the DRL for a multi-objective optimization method for green building design according to an embodiment of the present disclosure; Figure 7 Shows a schematic diagram of the DDPG learning curve for a multi-objective optimization method for green building design based on deep reinforcement learning according to an embodiment of the present disclosure; Figure 8 Shows a schematic diagram of building performance simulation for a multi-objective optimization method for green building design based on deep reinforcement learning according to an embodiment of the present disclosure; Figure 9 Shows a schematic structural diagram of a multi-objective optimization system for green building design based on deep reinforcement learning according to an embodiment of the present disclosure. Detailed implementation manners
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0019] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0020] Refer to Figure 1 and Figure 2 As shown, this embodiment provides a multi-objective optimization method for green building design based on deep reinforcement learning, including the following steps: S101. Determine the building performance objectives and relevant design parameters, construct an evaluation index system, establish a 3D BIM model, and import it into a simulation software based on BIM technology and building performance, so as to obtain a sample data set.
[0021] As Figure 3 shown, in this embodiment, the evaluation index system includes three expected objectives, specifically: building energy consumption , carbon emissions , and indoor thermal discomfort , and influencing variables, where the influencing variables include: (1) Building envelope structure: Wall insulation layer thickness; Roof insulation layer thickness; Building airtightness (permeability) (2) Building openings: Solar heat gain coefficient (SHGC) of windows; Window-to-wall area ratio; Exterior wall window opening rate (3) HVAC parameters: Coefficient of performance (COP) of the air conditioning system; Thermal efficiency of the heating system; Fresh air volume of the ventilation system; Indoor design temperature; Indoor lighting power density; Equipment power density.
[0022] The specific process is as follows: Use Autodesk Revit 2021 to establish a 3D BIM model of the building, save it as a gbXML format file, and then import it into the DesignBuilder software for building performance simulation and orthogonal testing to obtain a sample data set containing 384 groups of building performance results; The building performance simulation software DesignBuilder is based on the EnergyPlus program and has five core modules. It can consider the interactive effects of various factors on building performance and obtain the numerical values of building energy consumption, carbon emissions, and indoor thermal discomfort under different combinations of design parameters through simulation.
[0023] In this embodiment, the data set used is constructed based on a three-story teaching building located in Shanghai. By establishing a 3D BIM model of the teaching building in Autodesk Revit 2021, saving it as a gbXML format file, and then importing it into the DesignBuilder software for building performance simulation and orthogonal testing. During the simulation process, twelve predefined building design parameters are adjusted to obtain 384 groups of building performance simulation results covering the numerical values of building energy consumption, carbon emissions, and indoor thermal discomfort under different combinations of design parameters. These results constitute the data set for training the DNN meta-model and conducting DRL framework experiments. At the same time, calculate the Euclidean distance between the three objectives, and select the group of data with the smallest distance as the baseline for comparison with the optimization results.
[0024] S102. Construct a deep neural network (DNN) as a prediction meta-model, optimize the hyperparameters through Bayesian optimization (BO) technology, use the sample data set to train the DNN to learn the non-linear relationship between building design parameters and multi-objective building performance, and evaluate its prediction performance.
[0025] Such as Figure 4As shown, in this embodiment, the constructed meta-model for predicting based on a deep neural network consists of an input layer, an output layer, and several hidden layers. The input layer receives building design parameters as inputs, the output layer outputs the predicted value S of the building performance target, and the hidden layers perform feature extraction and transformation on the inputs through non-linear transformations. The hyperparameters of the DNN are tuned using Bayesian optimization (BO) technology to determine the optimal hyperparameter configuration, including the number of input units, the number of layers and nodes in each layer, the dropout rate, the optimizer, the learning rate, the activation function, etc. During the training process, the backpropagation algorithm is used to adjust the network weights, and the formula is: (1) Where is the learning rate, is the objective function, , are both weight parameters of the deep neural network DNN. Among them, is the original parameter, is the parameter after update and iteration. By continuously adjusting the weight parameters, the objective function is minimized, thereby achieving the purpose of optimizing the network performance.
[0026] The trained DNN is used to predict the building performance, and the prediction performance of the model is evaluated through R² and the mean absolute error (MAE). The specific prediction results are as Figure 5 shown.
[0027] Among them, the R² index is specifically: (2) Where TSS is the total sum of squared deviations, that is, the sum of the squares of the differences between the actual values and the mean of the actual values, and RSS is the residual sum of squares, that is, the sum of the squares of the differences between the actual values and the predicted values.
[0028] Generally speaking, the calculation result is generally in the range of 0 - 1. The closer it is to 1, the more accurate the model prediction is.
[0029] At the same time, in this embodiment, the performance is evaluated through the mean absolute value (MAE), specifically: (3) Where is the actual value, is the predicted value.
[0030] S103. Based on the established DNN meta-model, the deep deterministic policy gradient (DDPG) algorithm is used for optimization. The optimization process is treated as a multi-objective optimization (MOO) task, a framework based on DRL is constructed and trained, and the building performance is optimized through iterative learning.
[0031] As Figure 6As shown, specifically, in this embodiment, the DRL-based framework includes two components: engineering problem conceptualization and DDPG model establishment; Specifically: in the process of engineering problem conceptualization, setting the building design parameters is regarded as action A, and its value range is determined by the reasonable range of the building design parameters; defining the state S as the building performance target value, and the state update depends on the action execution. The formula is: (4) Among them, 、 、 is a DNN model for predicting based on a given action combination.
[0032] The DDPG model consists of an actor network and a critic network: Add random noise when selecting an action in the actor network, and the formula is to complete policy exploration. Among them, represents the building performance at the current t moment , and the action selection (i.e., parameter setting) obtained through the actor network at this time (with parameters ); adding Gaussian noise to obtain the real action selection, that is, .
[0033] Set the reward function (R) as (5) The reward function evaluates the pros and cons of transferring from state S by taking action A to state . It guides the model to find the optimal building design parameters, thereby achieving performance optimization.
[0034] The critic network is used to evaluate the performance of the current policy, that is, to calculate the Q value. Among them, the Q value is defined as the expected return of taking action A under the given state S, and the formula is: (6) Among them, is the immediate reward obtained when transferring from state S by taking action A to state ; is the discount factor, which is used to balance the weights of the immediate reward and the expected return; is the optimal action generated by the actor network under state ; represents the expectation of the return and reward R for all possible next states .
[0035] The critic network is trained by minimizing the mean squared error of the Q-value (i.e., maximizing the Q-value), and the formula is as follows: (7) The critic network parameters are updated using the gradient descent method , to minimize the mean squared error of the Q-value: (8) During the training process, the actor network predicts the value S of the building performance target (i.e., ) and generates the action function . The critic network calculates the Q-value based on the action function and the state, and uses the Q-value to update the actor network parameters through the gradient ascent method : (9) Thus, by maximizing the Q-value, that is, maximizing the expected return, the optimization of the building performance is achieved.
[0036] Specifically, as Figure 7 shown, during the training process of the DDPG model, the model performance is optimized by adjusting parameters such as the reward discount factor , learning rate, number of network layers, batch size, etc. It can be seen from the figure that the value of the reward function fluctuates greatly at the initial stage of the DDPG model training and gradually stabilizes as the number of iterations increases, reflecting the learning effect of the model in the process of optimizing the building performance. The setting of different parameters will affect the convergence speed and optimization effect of the model.
[0037] As Figure 8 shown, after being optimized by the DDPG model, the building is simulated for a 20-year time span. From the performance change curves of energy consumption, carbon emissions, and indoor thermal discomfort in the figure, it can be seen that the building performance remains stable in the long term, verifying the stability and effectiveness of the DDPG model in long-term optimization.
[0038] As Figure 9 shown, this embodiment also provides a multi-objective optimization system for green building design based on deep reinforcement learning, including: a sample data set acquisition module 1, a building performance prediction module 2, and a building performance optimization module 3; The sample data set acquisition module 1 is used to determine the building performance target and related design parameters and construct an evaluation index system, establish a 3D BIM model and import it into the simulation software based on BIM technology and building performance, so as to obtain the sample data set; The building performance prediction module 2 is used to construct a deep neural network prediction meta-model based on Bayesian optimization, and use the sample data set to train the meta-model to learn the non-linear relationship between the building performance target and related design parameters and the building performance, and output the building performance prediction result; The building performance optimization module 3 is used to construct a DRL-based framework, optimize the meta-model based on the framework, and optimize the building performance through iterative learning.
[0039] Before adopting the multi-objective optimization method for green building design based on deep reinforcement learning of the present invention, traditional design methods and existing optimization algorithms are difficult to achieve the efficient collaborative optimization of multiple building performance objectives, and have poor adaptability to complex environments. After adopting the method of the present invention, with the synergistic effect of the DDPG algorithm, BIM, and DNN, it is possible to effectively optimize building design parameters, significantly improve the performance of buildings in terms of energy consumption, carbon emissions, and indoor thermal comfort, and this method exhibits good stability and generalization under different climate regions and long time scales.
[0040] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited by the described action sequence, because according to the present disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present disclosure.
[0041] It should be understood that various forms of the processes shown above can be used, reordering, adding, or deleting steps. For example, the steps recorded in the present disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is made herein.
[0042] The above specific implementation manners do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A multi-objective optimization method for green building design based on deep reinforcement learning, characterized in that, The following steps are involved: Determine building performance objectives and related design parameters and construct an evaluation index system. Build a 3D BIM model and import it into simulation software based on BIM technology and building performance to obtain a sample data set. Constructing a deep neural network prediction meta-model based on Bayesian optimization, training the meta-model using the sample data set to learn the nonlinear relationship between the building performance target and related design parameters and building performance, and outputting a predicted value of the building performance target; A DRL-based framework is constructed, the meta-model is optimized based on the framework, and building performance is optimized through iterative learning.
2. The multi-objective optimization method for green building design based on deep reinforcement learning according to claim 1 is characterized in that: The evaluation index system includes expected goals and influencing variables; The expected goals include building energy consumption , carbon emissions , indoor thermal discomfort ; The influencing variables include: building maintenance structure, building openings and HVAC parameters; The building envelope includes the thickness of the wall insulation layer, the thickness of the roof insulation layer and the air tightness of the building; The building openings include the solar heat gain coefficient of windows, the window to wall area ratio, and the exterior wall window opening rate; The HVAC parameters include the cooling coefficient of the air conditioning system, the thermal efficiency of the heating system, the fresh air volume of the ventilation system, the indoor design temperature, the indoor lighting power density, and the equipment power density.
3. The multi-objective optimization method for green building design based on deep reinforcement learning according to claim 1, characterized in that The process of establishing a 3D BIM model and importing it into simulation software based on BIM technology and building performance to obtain a sample data set is as follows: A 3D BIM model of the building was created using Autodesk Revit 2021 and saved as a gbXML format file. The model was then imported into DesignBuilder software for building performance simulation and orthogonal testing to obtain a sample data set containing building performance results.
4. The multi-objective optimization method for green building design based on deep reinforcement learning according to claim 1, characterized in that, The prediction meta-model includes an input layer, an output layer and a hidden layer; The input layer takes the building performance target and related design parameters as input, the output layer outputs the building performance target prediction value, and the hidden layer extracts and converts the input features through nonlinear transformation.
5. The multi-objective optimization method for green building design based on deep reinforcement learning according to claim 4, wherein, The back propagation function of the prediction meta-model is: ; wherein, is the learning rate, is the objective function, , are both weight parameters of the deep neural network DNN, wherein, is the original parameter, is the parameter after update and iteration.
6. The multi-objective optimization method for green building design based on deep reinforcement learning according to claim 2, wherein, The DRL-based framework includes engineering problem conceptualization and DDPG model; The engineering problem conceptualization regards the setting of the building performance target and related design parameters as an action A, defines a state S as a predicted value of the building performance target, and the state update depends on the action execution, wherein the building performance target value S includes the expected target predicted value; The DDPG model includes an actor network and a critic network.
7. The multi-objective optimization method for green building design based on deep reinforcement learning according to claim 6, characterized in that, The building performance target prediction value is specifically: ; wherein, , , are DNN models for predicting based on a given action combination.
8. The multi-objective optimization method for green building design based on deep reinforcement learning according to claim 6, characterized in that, The actor network adds random noise when selecting actions , to complete the strategy exploration, specifically: ; wherein, Represents the building performance at the current time t Next, the action selection obtained by the actor network at this time; adding Gaussian noise Then we get the real action selection, namely .
9. The multi-objective optimization method for green building design based on deep reinforcement learning according to claim 6, wherein, The optimization process of the DDPG model for building performance is as follows: During the training process, the actor network predicts the value S according to the building performance target and generates an action function ; The critic network calculates the Q-value based on the action function and the state, and updates the actor network parameters by the gradient method using the Q-value and the critic network parameters , so as to maximize the Q-value, that is, to obtain the maximized expected return and complete the performance optimization process.
10. A multi-objective optimization system for green building design based on deep reinforcement learning, which is used to implement the multi-objective optimization method for green building design based on deep reinforcement learning according to any one of claims 1-9, characterized in that, include: Sample dataset acquisition module (1), building performance prediction module (2), and building performance optimization module (3); The sample data set acquisition module (1) is used to determine building performance goals and related design parameters and construct an evaluation index system, establish a 3D BIM model and import it into simulation software based on BIM technology and building performance, thereby obtaining a sample data set; The building performance prediction module (2) is used to construct a deep neural network prediction meta-model based on Bayesian optimization, and use the sample data set to train the meta-model to learn the non-linear relationship between the building performance objectives, relevant design parameters and building performance, and output the building performance prediction result; The building performance optimization module (3) is used to construct a DRL-based framework, optimize the meta-model based on the framework, and optimize the building performance through iterative learning.