Dormitory building indoor environment optimization method and system based on deep reinforcement learning combination
Through the combined method of deep reinforcement learning, the design parameters and layout matrix of dormitory buildings are optimized, which solves the problem that the existing technology is difficult to efficiently find a comfortable indoor environment optimization solution in complex environments, and achieves efficient and comfortable indoor environment optimization and has good generalization.
Patent Information
- Application Number
- CN202510270393.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing building performance optimization methods are difficult to efficiently find comfortable and reasonable indoor environment optimization solutions in environments with many control parameters and complex changes, and ignore the impact of indoor furniture placement on the physical environment.
Using a method of combining deep reinforcement learning, the pre-trained deep reinforcement learning model is used to optimize the indoor environment by adjusting the design parameters and layout matrix of dormitory buildings, combining the indoor average wind speed reward and the average effective daylight illumination reward, balanced exploration and efficient search are achieved.
The efficiency and comfort of indoor environment optimization have been improved, and the generated optimization scheme is more suitable for human living, has good generalization and adaptability, and can still perform well in non-training scenarios.
Smart Images

Figure CN120197261A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of indoor environment optimization, and particularly to an indoor environment optimization method and system for dormitory buildings based on the combination of deep reinforcement learning. Background Art
[0002] Existing research on building performance optimization mainly uses classical optimization algorithms for building environment optimization, such as Bayesian optimization, genetic algorithms, and annealing algorithms. Evolutionary algorithms such as genetic algorithms are based on population update or iterative search and are usually used to find the Pareto front solution set of multi-objective optimization problems. However, these algorithms still exhibit two significant limitations: First, finding a set of Pareto front solutions necessarily involves a time-consuming process related to population update or iterative search, but for tasks with more control parameters, it will lead to quite low operating efficiency; Second, when the optimization problem requires some minor adjustments to the building design or even the addition of some new optimization objectives, the existing algorithms cannot flexibly adapt to the changing conditions and exhibit poor generalization performance. Or use machine learning, such as the Chinese patent application 《CN109740243A》, which provides a furniture layout method based on piecewise reinforcement learning technology to achieve indoor environment optimization. An environmental scoring neural network model is constructed for each type of furniture, and then the indoor furniture placement is optimized in combination with reinforcement learning. Although the furniture layout optimization is achieved, during the optimization process, it only focuses on the static rationality of the furniture layout, such as the placement logic or functionality, and ignores the impact of its placement on the physical environment, and cannot ensure the comfort of the optimization plan.
[0003] Therefore, providing a method that can maximize the comfort of the indoor environment is a technical problem to be solved. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a method for optimizing the indoor environment of dormitories using deep reinforcement learning, which uses the deep reinforcement learning method to balance exploration and efficiently search for the optimal solution to obtain the indoor environment optimization plan with the highest comfort.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] According to the first aspect of the present invention, an indoor environment optimization method for dormitory buildings based on the combination of deep reinforcement learning is provided. The method uses a pre-trained deep reinforcement learning model to optimize the indoor environment of dormitory buildings; sets the building design parameters and layout matrix of the indoor environment of dormitory buildings as the state of the deep reinforcement learning model, sets the adjustment of the building design parameters as the action of the deep reinforcement learning model, and the reward of the deep reinforcement learning model includes the indoor average wind speed reward and the average effective daily illuminance reward.
[0007] As a preferred technical solution, the architectural design parameters include room width, room depth, sunshade depth, window width, bathroom depth, bathroom width, entrance door position, bathroom window width, and bathroom window height.
[0008] As a preferred technical solution, the layout matrix is an N×N integer matrix, and each element in the matrix represents a type of building component. The types of building components include indoor and outdoor spaces of the room, balcony railing, window, door, and wall.
[0009] As a preferred technical solution, the method for obtaining the layout matrix includes: encoding the types of building components, where the indoor and outdoor spaces of the room are encoded as 0, the balcony railing is encoded as 1, the window is encoded as 2, the door is encoded as 3, and the wall is encoded as 4.
[0010] As a preferred technical solution, the expression for calculating the reward is:
[0011] R = αR a - βR astd + γ(R1R d - R2R dstd )
[0012] where R a represents the average wind speed reward for the natural ventilation environment; R d represents the average effective daily illuminance reward for the sunlight environment; R astd represents the standard deviation of the wind speed; R dstd represents the standard deviation of the effective daily illuminance; α and β respectively represent the average wind speed weight and the average effective daily illuminance weight; γ represents the effective daily illuminance reward coefficient.
[0013] As a preferred technical solution, the calculation method for the average wind speed reward of the natural ventilation environment is:
[0014] R a = r a1 v1 + r a2 v2 + r a3 v3
[0015] where v1 represents the average wind speed in the living area; v2 represents the average wind speed in the balcony area; v3 represents the average wind speed in the bathroom area; r a1 , r a2 and r a3 respectively represent the weight values of the living area, the balcony area, and the bathroom area, and r a1 > r a2 > r a3 .
[0016] As a preferred technical solution, the calculation method of the average effective daily illuminance reward for the sunlight environment is as follows:
[0017] R d = r d1 D1 + r d2 D2,
[0018] where D1 is the average effective daily illuminance of the living area; D2 represents the effective daily illuminance of the balcony area; r a1 and r a2 respectively represent the weight values of the living area and the balcony area, and r d1 > r d2 .
[0019] As a preferred technical solution, the pre-training includes:
[0020] Parametrize according to the dormitory floor plan, use the building design parameters in the dormitory as the form control indicators, construct a three-dimensional model based on the form control indicators, obtain the initial state based on the three-dimensional model, and repeat the following steps until the number of iterations reaches the maximum number of iterations:
[0021] Generate the actions of the agent in the deep reinforcement learning model based on the state, and obtain the new state after executing the action; among them, the state in the first iteration is the initial state, and the state in the Nth iteration is the new state after the previous iteration;
[0022] Calculate the corresponding reward in the new state;
[0023] Update the parameters of the deep reinforcement learning model using the deep deterministic gradient algorithm based on the new state and the reward.
[0024] According to the second aspect of the present invention, a dormitory building indoor environment optimization system based on deep reinforcement learning combination is provided. The system is used to implement the above method and includes: a visualization platform for parametric dormitory modeling based on the dormitory floor plan and updating the dormitory modeling based on the actions generated by the deep reinforcement learning model; a deep reinforcement learning algorithm end for generating the action with the maximum reward value in the current state according to the state feedback by the visualization platform.
[0025] As a preferred technical solution, the visualization platform and the deep reinforcement learning algorithm end use the socket method for data transmission.
[0026] Compared with the prior art, the present invention has the following advantages:
[0027] 1). The present invention calculates rewards based on environmental parameters such as indoor wind speed and effective daily illumination, sets the architectural design parameters of the dormitory as the state, and the adjustment of the architectural design parameters as the action. It uses a deep reinforcement learning model to explore the coupling relationship between dynamic environmental factors including wind speed and effective daily illumination and the architectural design parameters, eliminating unreasonable schemes generated only based on the rationality of the building space (such as the maximum indoor space utilization rate), so as to generate an optimized indoor environment scheme with higher comfort and more suitable for human habitation. For example, when generating the optimized scheme, the effective daily illumination of the main living area is considered, etc.
[0028] 2). The present invention uses a pre-trained deep reinforcement learning model for indoor environment optimization design. Compared with traditional optimization algorithms, the method provided by the present invention generates optimized schemes with higher efficiency, and the method provided by the present invention has good generalization, and the optimization effect in non-training scenarios is still good. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a flowchart of the method of the present invention;
[0030] Figure 2 is a schematic diagram of the parametric modeling result of dormitory type 1 of the present invention;
[0031] Figure 3 is a schematic diagram of the parametric modeling result of dormitory type 2 of the present invention;
[0032] Figure 4 is a schematic diagram of the parametric modeling result of dormitory type 3 of the present invention;
[0033] Figure 5 is a curve graph of the training result of the deep reinforcement learning model of the present invention;
[0034] Figure 6 is an environmental performance optimization diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0036] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "an", "one kind", "the" and the like involved in this application do not indicate a limitation in quantity and may represent a singular or plural number. The terms "including", "comprising", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The similar words such as "connected", "coupled" and "linked" involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the front and rear associated objects. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0037] Embodiment 1
[0038] Deep reinforcement learning (DRL), which combines reinforcement learning and deep neural networks, excels in automatically creating novel artificial intelligence agents to learn from iterative interactions with the environment and adjust strategies through a trial-and-error process. Compared with genetic algorithms, reinforcement learning uses a policy or value function to approximate complex high-dimensional environments and can effectively explore high-dimensional spaces through a large amount of experience learning. Moreover, the pre-trained model of reinforcement learning can be transferred to similar tasks without the need to optimize from scratch. In addition, many traditional optimization algorithms do not have an inherent exploration mechanism, are prone to falling into local optima or have low exploration efficiency for global optima. DRL achieves efficient search by balancing exploration and exploitation, especially in unknown or changing environments. Nowadays, in the construction field, although a large number of studies have applied DRL to the building energy and environmental control stages, such as the reinforcement learning research for the design stage mainly focuses on the urban building design and outdoor environment optimization fields including indoor space layout and furniture layout design, there are few studies for the early building design stage. To address the above problems, this application proposes an indoor environment optimization method for dormitory buildings based on the combination of deep reinforcement learning. The process of this method is as Figure 1 shown and includes:
[0039] S1. Parametric modeling.
[0040] Obtain real design cases of school dormitories in a certain area, and select three typical dormitory floor types. These three types of dormitory floors are all composed of living areas, bathrooms, and balconies, but their layouts are different. Based on the three types of floors, parametric modeling is carried out based on Grasshopper shape grammar as Figure 2 , Figure 3 and Figure 4 shown. And batch collect the architectural design parameter data of parametric dormitory modeling. The architectural design parameters include room width, room depth, sunshade depth, window width, bathroom depth, bathroom width, position of the entrance door, width of the bathroom window, and height of the bathroom window. The detailed parameters are shown in Table 1.
[0041] Table 1 Architectural design parameters
[0042] Design parameter Value range Room frontage (RW) 3.3~3.9m Residential area depth (BedD) 5.1~7.3m Balcony depth (BalD) 1.2~1.8m Balcony door width (BalDW) 1.0~2.9m Bathroom depth (BathD) 1.8~3.0m Bathroom frontage (BathW) 1.1~1.6m Entrance door position (DP) (0.0~1.0)*RW Bathroom window width (BathWW) 0.6~1.1m Bathroom window sill height (BathWH) 1.3~1.6m
[0043] S2. Construction of the deep reinforcement learning model.
[0044] Based on the dormitory model built in step S1 and the obtained architectural design parameters, build a deep reinforcement learning model, including setting states, actions, and rewards.
[0045] 1) States: The state represents the current environment in which the agent operates. As part of the interface between the agent's perception and environmental conditions, the state provides the set of information that the agent needs to know to make decisions. In this embodiment, the state includes architectural design parameters and a layout matrix. The architectural design parameters are shown in Table 1. The layout matrix encodes the dormitory model into a 100×100 integer matrix, and each value represents a specific type of building component. The types of building components include indoor and outdoor spaces in the room, balcony railings, windows, doors, and walls. Specifically, the encoding method is as follows: The indoor and outdoor spaces in the room are encoded as 0, the balcony railings are encoded as 1, the windows are encoded as 2, the doors are encoded as 3, and the walls are encoded as 4.
[0046] 2) Actions: The agent performs operations to change the current state to the next state. Different actions will inevitably produce different effects. The goal is to select the best action from the action space a∈A, aiming to maximize the reward in a given scenario. In this embodiment, the adjustment of the architectural design parameters is used as the action of the deep reinforcement learning model. As shown in Table 1, there are 9 variable design parameters in the dormitory model, and the action space of each parameter is the variable range of the parameter as shown in the "Value range" column in Table 1. The value of each parameter can be freely adjusted within a continuous range, and the numerical precision is reserved to one decimal place. Therefore, the entire action space consists of 9-dimensional continuous variables. The formula for a complete action combination A is as follows: A=(a1,a2,…,a9), a i ∈[l i ,h i(i = 1, 2, …, 9), where l i and h i represent the minimum and maximum values of the i-th parameter, respectively.
[0047] 3) Reward: The specific goal of the agent is to maximize the reward and finally return the best policy consisting of a series of operations. The formulation of the reward function is highly correlated with two environmental performance values and the uniformity of distribution. It is necessary to increase the indoor wind speed and the useful daylight illuminance (UDI) while reducing the standard deviations of the wind speed and UDI, that is, including the average indoor wind speed reward and the average useful daylight illuminance reward. Specifically, its calculation formula is: R = αR a - βR astd + γ(R1R d - R2R dstd ), R a represents the average wind speed reward in the natural ventilation environment; R d represents the average useful daylight illuminance reward in the sunlight environment; R astd represents the wind speed standard deviation; R dstd represents the useful daylight illuminance standard deviation; α and β represent the average wind speed weight and the average useful daylight illuminance weight, respectively; γ represents the useful daylight illuminance reward coefficient.
[0048] Among them, the calculation method of the average wind speed reward in the natural ventilation environment is: R a = r a1 v1 + r a2 v2 + r a3 v3, v1 represents the average wind speed in the living area; v2 represents the average wind speed in the balcony area; v3 represents the average wind speed in the bathroom area; r a1 , r a2 and r a3 represent the weight values of the living area, the balcony area and the bathroom area, respectively, and r a1 > r a2 > r a3 .
[0049] The calculation method of the average useful daylight illuminance reward in the sunlight environment is:
[0050] R d = r d1 D1 + r d2 D2,
[0051] D1 is the average useful daylight illuminance in the living area; D2 represents the useful daylight illuminance in the balcony area; r d1 and r d2 represent the weight values of the living area and the balcony area, respectively, and r d1 > r d2 .
[0052] In this embodiment, the average wind speed weight α = 0.7, the average effective daily illumination weight β = 0.3, the effective daily illumination reward coefficient γ = 0.001, and the wind speed weight r of the living area a1 = 0.7, the wind speed weight r of the balcony area a2 = 0.1, the wind speed weight r of the bathroom area a3 = 0.2, the effective daily illumination weight r of the living area d1 = 0.7, the effective daily illumination weight r of the balcony area d2 = 0.3.
[0053] S3. Pre-training of the deep reinforcement learning model.
[0054] Parametrize according to the dormitory floor plan, use the building design parameters in the dormitory as the form control indicators, construct a 3D model based on the form control indicators, obtain the initial state based on the 3D model, and repeat the following steps until the number of iterations reaches 500 times.
[0055] S31. Generate the actions of the agent in the deep reinforcement learning model based on the state, and obtain the new state after executing the action; among them, the state at the first iteration is the initial state, and the state at the Nth iteration is the new state after the previous iteration.
[0056] S32. Calculate the corresponding reward in the new state.
[0057] S33. Update the parameters of the deep reinforcement learning model using the deep deterministic gradient algorithm based on the new state and the reward.
[0058] The present invention also provides a dormitory building indoor environment optimization system based on deep reinforcement learning combination for implementing the above method, including: a visualization platform for parametric dormitory modeling based on the dormitory floor plan and updating the dormitory modeling based on the actions generated by the deep reinforcement learning model; a deep reinforcement learning algorithm terminal for generating the action with the maximum reward value in the current state according to the state feedback by the visualization platform.
[0059] Specifically, when the system works, the deep reinforcement learning algorithm terminal runs based on python, the state of the dormitory indoor environment and the corresponding building design parameters are sent from the dormitory modeling constructed by Grasshopper of the visualization platform to the deep reinforcement learning algorithm terminal. After receiving the state information, the deep reinforcement learning algorithm terminal drives the agent to select the next action according to the learned strategy and sends these actions back to the visualization platform for execution; after receiving the action data, the visualization platform updates the dormitory model according to the action, obtains the indoor wind environment and light environment performance through simulation, calculates the corresponding reward value, and selects the scheme with the maximum reward value as the final optimization scheme. Among them, the visualization platform and the deep reinforcement learning algorithm terminal use the socket method for data transmission.
[0060] Example 2
[0061] To test the generalization and implementability of the DRL model, for three different indoor environment layouts (Type 1, Type 2, Type 3), the reinforcement learning model was independently trained in each scenario to obtain their respective pre-trained models (Model 1, Model 2, Model 3). Subsequently, each pre-trained model was applied to all three scenarios respectively, and a total of nine tests were conducted to record the optimal solutions in different scenarios. By comparing the performance of the pre-trained models in the training scenarios and non-training scenarios, the generalization ability of the models in cross-scenario applications was evaluated.
[0062] Among them, the training process of the deep reinforcement learning models corresponding to the three types of scenarios is as Figure 5 shown. It can be seen that the reward values of the training models for the three scenarios all gradually tend to be stable after about 100 generations, indicating that the models have found relatively optimal strategies. Figure 6 is the best optimization plan generated after being optimized by the deep reinforcement learning model for the corresponding scenario. It can be seen that after optimization, the uniform distribution of the wind speed and effective daily illuminance in the indoor environment of the dormitory can be achieved, and the optimization effect is good.
[0063] To verify that the method provided in the above embodiment is feasible and superior, in this embodiment, the method provided by the present invention is compared with the traditional optimization algorithm (taking the genetic GA algorithm as an example) to test its optimization effect and optimization efficiency, as shown in Table 2.
[0064] Table 2 Optimization test results of DRL and GA for three types of dormitories
[0065]
[0066] It can be seen that in the three types of scenarios, the rewards of the method (DRL) provided by the present invention are all better than those of GA, especially the advantage of the wind environment performance is more significant. In Type 1, the overall performance of DRL is 1.55% better than that of the GA method, and the advantage of the wind environment performance is 9.33%; the overall advantage obtained in Type 2 is 0.44%, and the advantage of the wind environment performance is 2.13%; the overall advantage in Type 3 is 1.97%, the advantage of the wind environment performance is 2.35%, and the advantage of the light environment performance is 2.26%. Through experiments, it takes an average of 9.5h to optimize a dormitory type scenario with the genetic algorithm, while it only takes about 6h to train a scenario with the DRL model. It only takes an average of 15 minutes to optimize a scenario with the method of the present invention, and the method of the present invention has reusability and does not require retraining during secondary optimization, showing an obvious time advantage compared with GA.
[0067] To verify the generalization of the pre-trained models, three pre-trained models were tested in three scenarios respectively, and the optimal results obtained from the tests are shown in Table 3. In Scenario Type 1, the maximum reward values of the three pre-trained models are similar, and the relative difference does not exceed 2.41%. In Scenario Type 2, the relative difference of the maximum reward values of the three pre-trained models is less than 0.7%. In Scenario Type 3, the relative difference of the maximum reward values of the three pre-trained models is less than 2.17%. It shows that the three pre-trained models all have a certain degree of generalization in non-training scenarios.
[0068] Table 3 Optimization test results of three pre-trained models in three scenarios
[0069]
[0070] As mentioned above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A dormitory building indoor environment optimization method based on deep reinforcement learning, characterized in that: The method uses a pre-trained deep reinforcement learning model to optimize the indoor environment of dormitory buildings; The architectural design parameters and layout matrix of the indoor environment of the dormitory building are set as the state of the deep reinforcement learning model, the adjustment of the architectural design parameters is set as the action of the deep reinforcement learning model, and the reward of the deep reinforcement learning model includes the indoor average wind speed reward and the average effective daylight illumination reward.
2. According to claim 1, a dormitory building indoor environment optimization method based on deep reinforcement learning is characterized in that: The architectural design parameters include room width, room depth, shading depth, window width, bathroom depth, bathroom width, entrance door position, bathroom window width and bathroom window height.
3. According to claim 1, a dormitory building indoor environment optimization method based on deep reinforcement learning is characterized in that: The layout matrix is an N×N integer matrix, and each element in the matrix represents a type of building component, and the building component types include the interior and exterior spaces of the room, balcony railings, windows, doors and walls.
4. The method for optimizing the indoor environment of dormitory buildings based on deep reinforcement learning according to claim 3 is characterized in that: The method for obtaining the layout matrix includes: encoding the building component type, wherein the interior and exterior space of the room is encoded as 0, the balcony railing is encoded as 1, the window is encoded as 2, the door is encoded as 3, and the wall is encoded as 4.
5. The method for optimizing the indoor environment of dormitory buildings based on deep reinforcement learning according to claim 1 is characterized in that: The expression for calculating the reward is: R=αR a -βR astd +γ(R1R d -R2R dstd ), Among them, R a Represents the average wind speed reward for natural ventilation environment; R d Represents the average effective daylight illuminance reward of the sunshine environment; R astd represents the standard deviation of wind speed; R dstd represents the standard deviation of effective daylight illumination; α and β represent the average wind speed weight and the average effective daylight illumination weight respectively; γ represents the effective daylight illumination bonus coefficient.
6. The method for optimizing the indoor environment of dormitory buildings based on deep reinforcement learning according to claim 5 is characterized in that: The calculation method of the average wind speed bonus in the natural ventilation environment is: R a =r a1 v1+r a2 v2+r a3 v3, Among them, v1 represents the average wind speed in the living area; v2 represents the average wind speed in the balcony area; v3 represents the average wind speed in the bathroom area; r a1 、r a2 and r a3 Represent the weight values of the living area, balcony area and living area respectively, and r a1 >r a2 >r a3 .
7. The method for optimizing the indoor environment of dormitory buildings based on deep reinforcement learning according to claim 5 is characterized in that: The calculation method of the average effective daylight illuminance reward of the sunshine environment is: R d =r d1 D1+r d2 D2, Where D1 is the average effective daylight illumination in the residential area; D2 is the effective daylight illumination in the balcony area; a1 and r a2 Represent the weight values of the living area and balcony area respectively, and r d1 >r d2 .
8. The method for optimizing the indoor environment of dormitory buildings based on deep reinforcement learning according to claim 1 is characterized in that: The pre-training includes: According to the dormitory plane, parameterization is performed, and the architectural design parameters in the dormitory are used as morphological control indicators. A three-dimensional model is constructed based on the morphological control indicators, and an initial state is obtained based on the three-dimensional model. The following steps are repeated until the number of iterations reaches the maximum number of iterations: Generate the action of the agent in the deep reinforcement learning model based on the state, and obtain the new state after executing the action; the state at the first iteration is the initial state, and the state at the Nth iteration is the new state after the previous iteration; Calculate the corresponding reward under the new state; Based on the new state and reward, the deep deterministic gradient algorithm is used to update the deep reinforcement learning model parameters.
9. A dormitory building indoor environment optimization system based on deep reinforcement learning, characterized in that: The system is used to implement the method according to any one of claims 1 to 8, comprising: A visualization platform for parametric dormitory modeling based on dormitory planes and updating dormitory modeling based on actions generated by deep reinforcement learning models; The deep reinforcement learning algorithm is used to generate the action with the maximum reward value in the current state based on the state feedback from the visualization platform.
10. The dormitory building indoor environment optimization system based on deep reinforcement learning according to claim 9 is characterized in that: The visualization platform and the deep reinforcement learning algorithm end use a socket method for data transmission.
Citation Information
Patent Citations
A furniture layout method and system based on a part-by-part reinforcement learning technology
CN109740243A
Indoor space temperature and humidity regulation and control method and system
CN115717758A
Zero-carbon building layout optimization method and system based on deep learning technology
CN116663412A
Indoor thermal environment control method based on RC model and deep reinforcement learning
CN116734424A
Multi-target sunlight greenhouse ventilation decision-making method and device, electronic equipment and storage medium
CN119126889A
Cited By
Country residential area automatic layout method based on shape grammar and multi-objective optimization
CN121279065A