Building Energy Consumption Dynamic Optimization Method and System Based on BIM and Reinforcement Learning
The integration of BIM and reinforcement learning creates a dynamic digital twin for building energy management, addressing inefficiencies in static systems by optimizing energy use and user comfort through real-time adaptation.
Patent Information
- Application Number
- CN202510621855.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-15
AI Technical Summary
Traditional building energy consumption management methods cannot be dynamically adjusted, resulting in energy waste or sacrificing user comfort. Existing model-based methods cannot maintain optimization results for a long time, and the static model and actual building physical properties are seriously mismatched.
Using a method based on BIM and reinforcement learning, a dynamic building digital twin is constructed, physical constraints are embedded for reinforcement learning training, multi-objective optimization strategies are generated, and energy equipment control is optimized through closed-loop feedback.
It realizes accurate dynamic optimization of building energy consumption, improves energy utilization efficiency and user comfort, ensures the physical feasibility of control strategies and system stability, and adapts to environmental changes and building state evolution.
Smart Images

Figure CN120145880B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of building energy management and intelligent control, and particularly to a building energy consumption dynamic optimization method and system based on BIM and reinforcement learning. Background Art
[0002] Traditional building energy consumption management mostly adopts preset strategies or simple rule-based control, such as the start and stop of heating, ventilation, and air conditioning (HVAC) based on a fixed schedule. The core defect of such methods lies in their static nature, which cannot dynamically adjust according to real-time environmental changes, human activities, and equipment status, often resulting in unnecessary energy waste or sacrificing user comfort, and it is difficult to meet the requirements of refined management of modern buildings.
[0003] To improve the optimization effect, existing technologies have introduced model-based methods such as model predictive control (MPC). However, the building models relied on by these methods are often static or quasi-static, and it is difficult to accurately reflect the dynamic evolution of physical properties such as building material aging, exterior wall pollution, and equipment efficiency decay. This model mismatch problem leads to a decline in the accuracy of optimal control over time and cannot achieve long-term optimal energy consumption performance. Summary of the Invention
[0004] To solve the above problems, the present invention provides a building energy consumption dynamic optimization method and system based on BIM and reinforcement learning. By adopting a mechanism of constructing a dynamic building digital twin integrating real-time data, embedding physical constraints for reinforcement learning training, and closed-loop feedback optimization, it can generate a multi-objective optimization strategy for building energy consumption that is physically feasible and dynamically adapts to environmental changes, significantly improving energy utilization efficiency, system operation safety, and user comfort.
[0005] The above object can be achieved through the following solutions:
[0006] A building energy consumption dynamic optimization method based on BIM and reinforcement learning includes establishing a BIM containing physical property parameters of building components, integrating real-time monitoring data to generate a building digital twin reflecting the evolution of dynamic thermal properties; extracting spatial topological relationships and component physical property parameters from the building digital twin to construct a multi-dimensional state space of a preset reinforcement learning model; embedding physical constraint conditions from the building digital twin into the reinforcement learning model, and training the reinforcement learning model based on the physical constraint conditions and a preset composite reward function to generate a multi-objective optimization strategy for energy equipment; parsing the multi-objective optimization strategy into a device control instruction set, controlling the energy equipment to perform parameter adjustment, and feeding back the adjustment result of the energy equipment to the building digital twin for real-time physical property simulation.
[0007] Optionally, the generation of the building digital twin reflecting the evolution of dynamic thermal properties includes: obtaining an aging coefficient dataset of building materials and establishing a dynamic heat conduction model associated with seasonal changes; identifying the solar radiation exposure coefficient of outdoor facade components and generating a light reflectance attenuation curve by combining historical dust deposition monitoring data; and updating the thermal property parameters of components in the BIM according to the dynamic heat conduction model and the light reflectance attenuation curve.
[0008] Optionally, the construction of the multi-dimensional state space for reinforcement learning includes: partitioning the building structure topology network from the building digital twin and extracting the energy conversion efficiency matrix of HVAC equipment; collecting air pressure gradient data and historical equipment operation logs to generate a probability distribution model of environmental load disturbance terms; and fusing the topology network, the efficiency matrix, and the probability distribution model into a discrete parameter set.
[0009] Optionally, the collection of air pressure gradient data and historical equipment operation logs to generate a probability distribution model of environmental load disturbance terms includes: collecting air pressure gradient data and historical equipment operation logs; analyzing the dynamic characteristics of the building internal environment according to the air pressure gradient data and performing micro-environment zoning; calculating the equipment energy consumption fluctuation threshold in each zone according to the historical equipment operation logs; and triggering the discrete parameter set when the real-time energy consumption data exceeds the equipment energy consumption fluctuation threshold.
[0010] Optionally, embedding the physical constraint conditions of the building digital twin in the reinforcement learning model includes: generating a constraint rule set that prohibits equipment from operating overloaded based on the stress simulation results of the BIM; generating and evaluating action sequences representing candidate policies during the reinforcement learning training process and intercepting the candidate policies that violate the constraint rule set; and when a conflict is detected, correcting the weight allocation of the preset composite reward function to guide policy convergence.
[0011] Optionally, the correction of the weight allocation of the preset composite reward function includes: obtaining data on conflicts between historical optimized policies and physical constraints, analyzing and constructing a conflict type - correction coefficient mapping table based on the data; and when a new policy conflict occurs, calling the corresponding correction coefficient in the conflict type - correction coefficient mapping table to adjust the weight allocation related to comfort and energy consumption in the preset composite reward function.
[0012] Optionally, after feedback of the adjustment result to the building digital twin for real-time physical property simulation, the following steps are included: collecting actual operation parameters of the energy equipment, and calculating a deviation rate from the simulation result of the building digital twin; when the deviation rate exceeds a preset threshold, triggering an online fine-tuning signal of the reinforcement learning model, and performing fine-tuning of the reinforcement learning; re-injecting the optimized policy after fine-tuning into the device control instruction set until the deviation rate returns to the tolerance interval.
[0013] Optionally, triggering the online fine-tuning signal of the reinforcement learning model includes: switching a dynamic threshold of the deviation rate according to a building operation mode switching signal; reducing the dynamic threshold of the deviation rate during a low-energy consumption period at night; and expanding the dynamic threshold of the deviation rate during a peak electricity consumption period.
[0014] Optionally, parsing the multi-objective optimization policy into a device control instruction set includes: identifying communication interface types of different device protocols, and constructing an instruction format conversion template library; matching control protocol characteristics of a target device in the instruction format conversion template library, and generating an adapted binary instruction block; and dynamically queueing and sorting the sending order of the instruction blocks according to device response delay requirements.
[0015] Based on the same inventive concept, the present invention further provides a building energy consumption dynamic optimization system based on BIM and reinforcement learning. The system includes: a digital twin construction module, configured to establish a BIM including physical property parameters of building components, and fuse real-time monitoring data to generate a building digital twin reflecting the evolution of dynamic thermal properties; a state space construction module, configured to extract spatial topological relationships and component physical property parameters based on the building digital twin, and construct a multi-dimensional state space of a preset reinforcement learning model; a policy generation module, configured to embed physical constraint conditions from the building digital twin into the reinforcement learning model, and train the reinforcement learning model based on the physical constraint conditions and a preset composite reward function to generate a multi-objective optimization policy for energy equipment; and a control and feedback module, configured to parse the multi-objective optimization policy into a device control instruction set, control the energy equipment to perform parameter adjustment, and feedback the adjustment result of the energy equipment to the building digital twin for real-time physical property simulation.
[0016] Compared with the prior art, the present invention has the following advantages:
[0017] 1. The accuracy and long-term effectiveness of energy consumption optimization are improved. By constructing a dynamic building digital twin that fuses real-time monitoring data, the dynamic evolution of the physical properties of the building can be reflected in real time, overcoming the mismatch problem brought by traditional static models, providing a more accurate environmental model for reinforcement learning, and thus improving the accuracy and long-term effect of the optimization policy.
[0018] 2. It ensures the physical feasibility of the control strategy and the system security. By embedding the physical constraint conditions (such as equipment operation limits, structural safety limits, etc.) from BIM into the reinforcement learning training process, it ensures that the generated energy consumption optimization strategy is physically feasible, avoids damage to energy equipment and potential safety risks, and improves the stability and reliability of the system.
[0019] 3. It realizes the effective balance and adaptive optimization of multiple objectives. By using the decision-making ability of reinforcement learning and the preset composite reward function mechanism, it can dynamically balance multiple optimization objectives such as energy conservation, user comfort, and equipment health according to actual needs; combined with the closed-loop feedback and online fine-tuning mechanism, the system can continuously self-optimize according to the actual operation effect and dynamically adapt to environmental changes and the evolution of building states.
[0020] 4. It expands the application value of BIM in the building operation stage, elevates BIM from a traditional static information database to the core basis for dynamic optimization control, and gives full play to the potential of BIM in the whole life cycle management of buildings, especially in the intelligent and refined operation stage, through the deep integration with real-time data and intelligent algorithms.
[0021] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a schematic flowchart of the building energy consumption dynamic optimization method and system based on BIM and reinforcement learning according to an embodiment of the present invention.
[0024] Figure 2 It is a regional temperature heat map according to an embodiment of the present invention.
[0025] Figure 3 It is a three-dimensional temperature scatter plot according to an embodiment of the present invention.
[0026] Figure 4 It is a deviation rate and dynamic threshold graph according to an embodiment of the present invention.
[0027] Figure 5It is a schematic structural diagram of the building energy consumption dynamic optimization method and system based on BIM and reinforcement learning according to an embodiment of the present invention. Detailed implementation manners
[0028] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] Referring to Figure 1 , an embodiment of the present invention proposes a building energy consumption dynamic optimization method based on BIM and reinforcement learning, aiming to construct a building digital twin by integrating building information model and reinforcement learning technology, and on this basis, realize real-time, dynamic, and closed-loop optimization of the operation strategy of energy equipment, so as to minimize building energy consumption while meeting the comfort requirements.
[0030] The method of this embodiment specifically includes the following steps:
[0031] Establish a BIM containing physical attribute parameters of building components, and integrate real-time monitoring data to generate a building digital twin reflecting the evolution of dynamic thermal properties;
[0032] Specifically, the building information model not only includes the geometric information of the building, such as the positions and dimensions of walls, floors, and doors and windows, but also needs to input the physical attribute parameters of key components, such as the thermal conductivity, specific heat capacity, density, solar radiation absorption rate, reflectivity, etc. of materials. Integrating real-time monitoring data means inputting the real-time data streams obtained from sensors deployed inside and outside the building, the real-time data streams obtained from external sensors such as temperature, humidity, light intensity, carbon dioxide concentration, and personnel counting sensors, as well as the weather forecast data from meteorological services, into a preset digital model to construct a building digital twin. The building digital twin is a dynamic and real-time mirror of the physical building, and its core lies in simulating the internal environment parameters of the building, such as the temperature field and humidity field distribution, through physical models and data-driven models, such as the models established based on the energy balance equation and heat transfer principles, as Figure 2 shown, and the thermal properties of components, such as the thermal conductivity and reflectivity considering the influence of aging and dust accumulation, and the real state changing with time and external conditions. Its purpose is to provide a high-fidelity virtual environment for subsequent energy consumption simulation and optimization strategy formulation, as Figure 3 shown.
[0033] Extract the spatial topological relationship and component physical property parameters based on the building digital twin, and construct the multi-dimensional state space of the preset reinforcement learning model;
[0034] Specifically, construct the multi-dimensional state space of reinforcement learning based on the building digital twin. This process includes extracting key spatial topological relationships, such as functional zoning, connectivity, and equipment service relationships, as well as core component physical property parameters, such as dynamic heat transfer coefficient, air permeability, and equipment energy efficiency. Combine these extracted attributes with real-time environmental parameters such as indoor and outdoor temperature and humidity, personnel and air quality indicators, and equipment operating status such as switch status, set point, and power consumption, and jointly fuse them into a multi-dimensional state vector or tensor. This state space aims to comprehensively characterize the key factors affecting building energy consumption and comfort, serving as the basis for the reinforcement learning agent to perceive the environment.
[0035] Embed the physical constraint conditions from the building digital twin into the reinforcement learning model, and train the reinforcement learning model based on the physical constraint conditions and the preset composite reward function to generate a multi-objective optimization strategy for energy equipment;
[0036] Specifically, the physical constraint conditions are used to guide the reinforcement learning model to explore within a safe and reasonable range, and they are derived from digital twin physical simulations or equipment specifications. For example, physical constraints may include equipment operation limits such as start-stop frequency and power range, safety thresholds such as pipeline pressure and supply air temperature and humidity, and temperature requirements or structural load-bearing limits in specific areas such as data rooms. The preset composite reward function (Reward Function) is the key to guiding the reinforcement learning goal, and it is usually the weighted sum of multiple sub-goals. For example:
[0037] ,
[0038] Among them, represents the immediate total reward obtained by the agent, represents the energy consumption cost, is the user comfort score, is the indoor air quality score. , , are the weight coefficients corresponding to these three items respectively, used to balance the importance of different goals, It is the penalty value for behaviors that violate physical constraints or cause adverse consequences, such as frequent start-stop of equipment. The comfort score can be calculated based on the Predicted Mean Vote (PMV) model, the Predicted Percentage of Dissatisfied (PPD) model, or user feedback, and the air quality score can be evaluated according to indicators such as CO2 concentration. During the training process, the agent interacts with the digital twin environment, using deep reinforcement learning algorithms to learn to maximize the long-term cumulative reward under physical constraints, and accordingly generates optimized control strategies for energy equipment such as air conditioners, lighting, and fresh air, including but not limited to recommended temperature setpoints, air volumes, or switching times.
[0039] Parse the multi-objective optimization strategy into a device control instruction set, control the energy device to perform parameter adjustment, and feedback the adjustment result of the energy device to the building digital twin for real-time physical property simulation.
[0040] Specifically, the high-level optimization strategy output by the reinforcement learning model needs to be parsed into a device control instruction set for execution. This process involves converting the policy requirements into low-level commands and determining specific instruction parameters such as register values. The generated instructions are sent to the energy device actuator, such as a valve or a frequency converter, through the Building Automation System (BAS). After execution, the actual operating states collected by sensors, such as power consumption, temperature, and environmental change data, are used on the one hand to update the digital twin to keep it synchronized, and on the other hand as the real feedback of reinforcement learning, including the next state and the actual reward. This feedback drives the agent to continuously make decisions and learn, forming a closed-loop optimization process of "perception - decision - control - feedback - learning".
[0041] By integrating the high-precision BIM, real-time data, and reinforcement learning algorithms, constructing the building digital twin and performing closed-loop optimization, it is possible to achieve refined and intelligent management of building energy consumption, significantly improving energy utilization efficiency and user comfort experience.
[0042] Optionally, the generation of the building digital twin reflecting the evolution of dynamic thermal properties includes:
[0043] Obtain the aging coefficient dataset of building materials and establish a dynamic heat conduction model associated with seasonal changes;
[0044] Specifically, to reflect the evolution of dynamic thermal properties, the method includes obtaining an aging coefficient dataset of building materials and establishing a dynamic thermal conductivity model associated with seasonal variations. The aging coefficient can be sourced from material suppliers, industry standard databases, or accelerated aging experiments, and it quantifies the laws of changes in key material properties such as thermal conductivity and airtightness over time and environmental factors such as temperature, humidity, radiation, etc. When establishing the dynamic thermal conductivity model based on this, the actual service life of the material and the annual average temperature representing seasonal environmental influencing factors and the annual average relative humidity etc. can be used as input variables of the model, and these input variables are used to correct the basic thermal conductivity of the material in the initial state . For example, the dynamic thermal conductivity of a specific insulation material can be expressed as a function of service life , annual average temperature and annual average relative humidity . An exemplary model formula is as follows:
[0045] ,
[0046] where, represents the value of the dynamic thermal conductivity under specific conditions, is the basic thermal conductivity of the material, is the service life, and are the annual average environmental temperature and relative humidity respectively, and are empirical aging coefficients related to the material type, jointly describing the change trend of performance over time (where the coefficient is usually less than 1), and are correction coefficients considering the influence of temperature and humidity respectively, and and are the reference temperature and humidity values used in the calculation. Using such a dynamic thermal conductivity model that includes the effects of aging and environmental factors can enable the building digital twin to more accurately simulate and reflect the dynamic changes in the actual thermal insulation performance of the building envelope during long-term use, thereby providing a more reliable basis for energy consumption analysis and optimization.
[0047] Identify the solar radiation exposure coefficient of outdoor facade components and generate a light reflectance attenuation curve in combination with historical dust deposition monitoring data;
[0048] Specifically, the solar radiation exposure coefficients of each component are calculated using the geographical information and solar trajectory tools of the BIM, and this calculation needs to consider the occlusion effect. At the same time, the cumulative effective dust deposition amount per unit area of the facade is estimated. , which can be achieved by analyzing the data of the deployed dust sensors or by combining the local air quality historical data, such as PM2.5 and PM10 concentrations, with rainfall records, for example, by integrating the difference between the dust deposition and washing rates per unit time. Based on the estimated , an attenuation model of the surface light reflectivity is established. For example, it is assumed that the reflectivity decreases exponentially with the dust deposition amount:
[0049]
[0050] where R(t) represents the real-time surface light reflectivity considering the influence of dust deposition, is the initial light reflectivity of the surface of this component in the clean state, is the natural exponential function, is a light reflectivity attenuation constant related to the surface characteristics of the material and the type of dust, and is the cumulative effective dust deposition amount corresponding to the time. Applying such a light reflectivity attenuation model to update the surface properties of the component can enable the building digital twin to more realistically simulate the change of the heat absorption of the exterior wall due to dust coverage over time, thereby improving the accuracy of the overall building energy consumption simulation.
[0051] Update the thermal property parameters of the components in the BIM according to the dynamic heat conduction model and the light reflectivity attenuation curve.
[0052] Specifically, this parameter update process can be executed at a preset cycle, such as quarterly or annually, or triggered when the monitored aging / dust deposition index reaches the threshold. After being triggered, the dynamic model is called to calculate the actual thermal properties of the current component, such as the dynamic heat conduction coefficient Kt and the reflectivity Rt. Subsequently, through the BIM software API or directly modifying the digital twin database, the calculated new parameter values are written into the corresponding component attributes. This dynamic update mechanism ensures that the digital twin accurately reflects the long-term changes of the building physical properties and improves the accuracy of subsequent energy consumption simulation and prediction.
[0053] Exemplarily, for an office building located in the city center, a certain low-emissivity coated glass is used for its glass curtain wall. Initially, the reflectivity is set to 0.6 in the BIM. By analyzing the PM10 monitoring data and rainfall data over the past three years, the annual average effective dust accumulation amount is calculated and denoted as Deff(3 years). Substituting this value and the assumed light reflectivity decay constant Kd = 0.05 into the light reflectivity decay model, the actual reflectivity R(3 years) of the glass surface after three years has decayed from 0.6 to approximately 0.52. At the same time, the aging model of its sealant strip shows that the thermal conductivity has increased by 5%. Before conducting the annual energy consumption simulation and optimization strategy training, the system automatically updates the reflectivity of the corresponding curtain wall components in the digital twin to 0.52, and the thermal conductivity of the sealed part is also adjusted accordingly. This enables the digital twin to more accurately predict the solar radiation heat entering through the curtain wall in summer and the cold air infiltration or heat loss in winter, thereby allowing the reinforcement learning model to formulate a more practical air-conditioning operation strategy and avoid energy waste caused by using outdated parameters.
[0054] Optionally, the multi-dimensional state space for constructing the reinforcement learning includes:
[0055] Partition the building structure topology network from the building digital twin and extract the energy conversion efficiency matrix of the HVAC equipment;
[0056] Specifically, according to the spatial data in the BIM and the HVAC system drawings, the building is divided into several thermodynamically or control-wise relatively independent zones. At the same time, identify the air flow and heat transfer paths between the zones, as well as the connection service relationships between specific HVAC equipment and each zone, so as to construct a building topology network in the form of a graphical structure or an adjacency matrix. In addition, it is necessary to collect the technical specifications of the main HVAC equipment or fit through measured data to obtain its energy conversion efficiency under different operating conditions. The operating conditions involve conditions such as load rate, ambient temperature, water temperature, etc. For example, the performance coefficient of a variable-frequency chiller can be expressed as the chilled water outlet temperature , the cooling water inlet temperature and the load rate function ,
[0057] ,
[0058] Similarly, the boiler thermal efficiency can be modeled as a function of the load rate and the return water temperature function ,
[0059] ,
[0060] The pump efficiency can be modeled as flow rate and head functions ,
[0061] .
[0062] Herein represents the model function for calculating the chiller , represents the model function for calculating the boiler thermal efficiency represents the model function for calculating the pump efficiency is the coefficient of performance of refrigeration , are the boiler thermal efficiency and the pump efficiency respectively , , are the chilled water outlet temperature, the cooling water inlet temperature and the boiler return water temperature respectively is the load percentage is the pump flow rate is the pump head. These efficiency models can be in the form of matrices, polynomials or neural networks, etc., and are an important part of the state space for evaluating the potential energy consumption of different control strategies
[0063] Collect the air pressure gradient data and the historical equipment operation logs to generate a probability distribution model of the environmental load disturbance term
[0064] Specifically, collect the air pressure difference data by installing air pressure sensors at key positions inside and outside the building, such as different floors and the main entrance. Combine the historical operation logs containing equipment switch, duration and energy consumption information and the corresponding meteorological data such as wind speed and wind direction to analyze the influence of the air pressure gradient change on infiltration, ventilation and equipment load. Use statistical methods such as kernel density estimation, Gaussian mixture model or machine learning models such as hidden Markov model to establish a probability distribution model. This model is used to describe the occurrence probability of the disturbance term and its influence degree on the load, such as outputting the magnitude and probability of the increased load due to infiltration under specific conditions
[0065] Fuse the topological network, the efficiency matrix and the probability distribution model into a discrete parameter set
[0066] Specifically, integrate the aforementioned information into a multi-dimensional state vector S that can be understood by the agent. This vector includes key environmental parameters such as indoor and outdoor temperature and humidity, CO2 and radiation intensity; equipment status such as switch, mode, output percentage and the current efficiency calculated by the efficiency model; and the current or predicted disturbance information obtained from the disturbance model. Continuous variables can be discretized as needed. The state vector S comprehensively describes the current environment and operation status and constitutes the basis for the agent's decision-making
[0067] Exemplarily, consider the energy consumption optimization scenario of a large shopping mall building. The constructed topological network model shows that there is a significant heat exchange relationship between the atrium area and the surrounding store areas, and these two types of areas are served by two large air handling units (AHUs). The extracted equipment efficiency model, such as the model represented by the function Fchiller or Fboiler, further indicates that when the outdoor temperature is below 10 degrees Celsius, the heating efficiency of the AHU will decrease by 15%. At the same time, the probability distribution model of the environmental load disturbance term generated based on historical data shows that during the peak pedestrian flow period on weekend afternoons, there is an 80% probability that the carbon dioxide (CO2) concentration in the area near the main entrance of the mall will exceed the safety or comfort threshold of 1000 ppm, which will lead to a significant increase in the required fresh air load. Therefore, the reinforcement learning state space S constructed for this mall contains information in multiple dimensions, such as: the current temperature and humidity readings in the atrium and each store area, the real-time outdoor temperature, the current operating modes and output air volumes of the two AHUs, the current actual heating efficiency of the AHU queried from the efficiency model based on the real-time outdoor temperature, and the probability value of CO2 concentration exceeding the standard in the entrance area output by the disturbance term model. Based on this comprehensive and dynamic state information set S, the reinforcement learning model can make predictive decisions. For example, before predicting the upcoming peak of pedestrian flow, it appropriately increases the introduction volume of the fresh air system in advance and intelligently selects to perform preheating or precooling operations during periods when the equipment operating efficiency is relatively high, so as to actively cope with the upcoming personnel load disturbance and the challenge of equipment efficiency fluctuating with changes in the external environment.
[0068] Optionally, the generating of the probability distribution model of the environmental load disturbance term by collecting the barometric gradient data and historical equipment operation logs includes:
[0069] Collecting the barometric gradient data and historical equipment operation logs;
[0070] Specifically, the barometric gradient data is collected by deploying barometric sensors at representative measuring points inside and outside the building, recording the differential pressure data at a preset frequency to form a time series of the barometric field reflecting the effects of wind pressure, thermal pressure, etc. The historical equipment operation logs are collected by extracting the operation records of key energy equipment from the BAS, etc., including timestamps, status, set values, and measured performance parameters, providing a data basis for subsequent analysis.
[0071] Analyzing the dynamic characteristics of the building internal environment based on the barometric gradient data and performing microenvironment zoning;
[0072] Specifically, in addition to monitoring the pressure gradient, micro wind speed sensors can be deployed in key areas within the building, or calibrated CFD simulations in the digital twin can be used to monitor or simulate the internal air flow velocity and direction. Based on characteristics such as air flow stability, temperature uniformity, and pollutant diffusion patterns, a large area is subdivided into microenvironment zones with different dynamic characteristics. For example, an open office area can be divided into a near-window zone, an internal stable zone, and a downstream zone of the air outlet.
[0073] According to the historical equipment operation logs, calculate the equipment energy consumption fluctuation thresholds within each zone.
[0074] Specifically, for each microenvironment zone, screen the operation logs of the service equipment. Analyze the temporal fluctuation statistical characteristics of the equipment energy consumption in the zone under similar external conditions and personnel patterns. For example, calculate the standard deviation or interquartile range. Set the dynamic energy consumption fluctuation threshold based on the statistic. For example, calculate it by adding or subtracting several times the standard deviation from the recent average energy consumption to define the normal energy consumption fluctuation range of the zone.
[0075] When the real-time energy consumption data exceeds the equipment energy consumption fluctuation threshold, trigger the discrete parameter set.
[0076] Specifically, monitor the equipment energy consumption of each microenvironment zone in real time and compare it with the preset dynamic threshold. If the energy consumption continuously exceeds the threshold, for example, for several consecutive sampling periods, it is determined that an unexpected disturbance has occurred, such as a window being opened, a sudden increase in personnel, or a precursor to equipment failure. At this time, trigger the state space dynamic reorganization: an abnormal flag bit can be temporarily added to increase the update frequency or accuracy of the state parameters of this zone, or activate a fine-grained local disturbance sub-model to replace / supplement the global model prediction and incorporate the results into the state space. This reorganization aims to enable the reinforcement learning agent to perceive local anomalies faster and more accurately and quickly adjust the strategy to respond.
[0077] Exemplarily, consider an application scenario of an intelligent ward floor. Among them, the ward near the balcony door is identified and divided into an independent micro - environment partition. According to the analysis of historical operation data, under normal circumstances where there is no personnel activity at night and the doors and windows are confirmed to be closed, the standard deviation of the energy consumption fluctuation of the fan - coil unit serving this ward is 5 watts. Based on this, the dynamic fluctuation threshold of the energy consumption of this partition is set to float 10 watts above and below the normal average value. One night, real - time monitoring found that the energy consumption of the fan - coil unit in this ward suddenly exceeded its night - time average level by 30 watts, and this high - energy - consumption state lasted for five minutes, significantly exceeding the upper limit of the preset positive 10 - watt fluctuation threshold. Accordingly, it is judged that abnormal situations such as the balcony door not being tightly closed may have occurred in this micro - environment partition, thus triggering the dynamic reorganization mechanism of the state space. The specific reorganization operation may be to specifically mark the state of this ward as "suspected high - penetration load" in the current state vector of the reinforcement learning model, and temporarily increase the weight of the temperature reading of this ward in the calculation of the state vector or the reward function. After receiving this reorganized state information containing the abnormal indication, the reinforcement learning agent may preferentially select control actions such as increasing the air supply volume to this ward or adjusting the temperature set point to try to maintain the temperature stability in the ward. At the same time, it may also send a prompt message about this potential abnormal situation to the building operation and maintenance management platform. This dynamic reorganization and response mechanism enables the overall optimization to respond more sensitively to local and sudden disturbance events.
[0078] Optionally, the physical constraint conditions of embedding the building digital twin in the reinforcement learning model include:
[0079] Based on the stress simulation results of the BIM, generate a constraint rule set that prohibits equipment from operating over - load;
[0080] Specifically, use the equipment location, weight, and structure data provided by the BIM, combined with finite - element analysis software, to simulate the influence of static and dynamic loads of large or vibration - sensitive equipment under different working conditions on the structure. According to the simulation results, design specifications, and safety standards, determine the safety boundary conditions, such as the maximum allowable vibration amplitude of the floor slab or the maximum instantaneous stress of the pipeline. Convert these boundary conditions into clear constraint rule sets, such as specifying the maximum operating frequency or maximum starting acceleration of the equipment.
[0081] During the reinforcement learning training process, generate and evaluate action sequences representing candidate strategies, and intercept the candidate strategies that violate the constraint rule set;
[0082] Specifically, in the reinforcement learning training process, it is usually carried out in the digital twin simulation environment. When the agent generates a candidate action sequence, it is reviewed through the "safety layer" or "constraint check module". This module predicts the future state after the execution sequence, may call the digital twin simulation, and compare it with the constraint rule set. If the prediction result shows that a constraint will be violated, such as vibration exceeding the limit, the action sequence or policy will be intercepted, such as preventing its execution or giving a large negative reward, to prevent the agent from learning dangerous policies.
[0083] When a conflict is detected, the weight allocation of the preset composite reward function is corrected to guide the policy convergence.
[0084] Specifically, when a candidate action is intercepted due to a predicted violation of physical constraints, it indicates that the policy tends to explore dangerous areas. To guide learning, the weights of the preset composite reward function can be temporarily adjusted. If the conflict is related to equipment overload, the energy efficiency weight Wenergy can be temporarily reduced, and at the same time, the penalty strength of the penalty term Penalty for violating this constraint can be increased. This dynamic weight adjustment aims to emphasize that avoiding specific constraints is more critical than energy conservation at present, guiding the agent to select safe actions away from the constraint boundary in the future, and making the policy converge to a safe and optimal area.
[0085] Optionally, the correction of the weight allocation of the preset composite reward function includes:
[0086] Obtain the data of the conflict between the historical optimized policy and physical constraints, and analyze and construct a conflict type - correction coefficient mapping table based on the data;
[0087] Specifically, during the training and long-term operation of the reinforcement learning model, record every physical constraint conflict event that occurs or is attempted, including the type of conflict, such as equipment overclocking, pipeline overpressure, minimum temperature limit, or maximum start-stop times, etc., as well as the state at the time, the attempted actions, and the relevant environmental background. Offline analysis of these historical cases can identify which types of conflicts are more frequent or more likely to occur under specific conditions. According to the severity, frequency of the conflict, and its potential impact on the overall goals, such as energy consumption and comfort, set a set of correction coefficients for each identified conflict type. These coefficients clearly define how each item in the preset composite reward function, such as the energy efficiency weight Wenergy, the comfort weight Wcomfort, or a specific penalty term Penaltytype, should be adjusted when this type of conflict occurs. Finally, store these conflict types and their corresponding correction coefficients, which can be a set of multipliers or addends, in a lookup table or rule base to form a mapping table.
[0088] When a new policy conflict occurs, call the corresponding correction coefficient in the conflict type - correction coefficient mapping table to adjust the weight allocation related to comfort and energy consumption in the preset composite reward function.
[0089] Specifically, during the online decision-making or offline training of the reinforcement learning model, when the safety layer detects an impending policy conflict, that is, when the agent attempts to execute an action that violates physical constraints, the specific type of the conflict is first identified. Then, according to the identified conflict type, the pre-constructed mapping table is queried to obtain the corresponding correction coefficients. These coefficients are used to dynamically and temporarily adjust the weights of the terms in the preset composite reward function currently used for calculating rewards, such as the weight ratio of the comfort factor Wcomfort and the energy consumption factor Wenergy, as well as the intensity of possible specific penalty terms. For example, if the mapping table indicates that for a "pipeline overpressure" conflict, Wenergy should be multiplied by 0.9, that is, the energy consumption weight is reduced, and the penalty term Penaltypressure related to pressure should be multiplied by 1.5, that is, the penalty intensity is increased, then when calculating the reward generated by this attempted violation action, these adjusted weights will be used. This adjustment is usually temporary and aims to give a more accurate feedback signal to the current violation attempt to guide the agent to correct its behavior.
[0090] Optionally, after the adjustment result is fed back to the building digital twin for real-time physical property simulation, it includes:
[0091] Collect the actual operating parameters of the energy equipment and calculate the deviation rate from the simulation results of the building digital twin;
[0092] Specifically, when the device executes the control instruction, through the deployed sensors, such as power meters, flow meters, or thermometers, etc., collect the actual operating parameters of the device, such as the actual power consumption , the actual flow rate , or the actual temperature of the area . At the same time, in the building digital twin, use the same control instruction for simulation calculation to obtain the corresponding simulation results, such as the simulated power , the simulated flow rate , or the simulated temperature . Then, calculate the deviation between the actual value and the simulated value, such as calculating the relative deviation rate of power and the absolute deviation of temperature . The example formulas are as follows:
[0093] ,
[0094] ,
[0095] In these formulas refers to the actually measured power, refers to the power predicted by simulation, is the calculated relative deviation rate of power; Refers to the actually measured temperature, Refers to the temperature predicted by simulation, is the calculated absolute temperature deviation. By calculating such deviations for key parameters, the current prediction accuracy of the digital twin model can be effectively quantified, or the gap between the physical actual response and the simulation expectation can be reflected.
[0096] When the deviation rate exceeds the preset threshold, trigger the online fine-tuning signal of the reinforcement learning model and perform the fine-tuning of the reinforcement learning;
[0097] Specifically, a reasonable deviation threshold needs to be set for each key parameter to be monitored. The setting of this threshold should comprehensively consider the measurement accuracy of the sensors used, the inherent errors that may be brought about by the simplification of the digital twin model, and the random fluctuation range under normal operating conditions. For example, for the relative deviation rate related to energy consumption, its threshold can be set to ten percent, while for indicators such as regional temperature, its absolute deviation threshold can be set to 0.5 degrees Celsius. These deviation values will be continuously monitored. When it is detected that the deviation of one or several key parameters continuously exceeds their corresponding preset thresholds, for example, this over-threshold state lasts for a set time window, such as thirty minutes, then it is judged that there is a significant and continuous deviation between the digital twin model and the actual operating state of the physical building, as Figure 4 shown. This deviation may be due to the inaccuracy of the model parameters themselves, the existence of dynamic influencing factors not fully considered by the model, or the change of the performance of physical devices over time. In this case, an online fine-tuning signal will be generated and sent to the reinforcement learning model, indicating that model adjustment is required.
[0098] Inject the fine-tuned optimization strategy back into the device control instruction set until the deviation rate returns to the tolerance interval.
[0099] Specifically, after receiving the online fine-tuning signal, the reinforcement learning model fine-tunes its network parameters using the real interaction data accumulated recently. This adjustment usually adopts an optimization algorithm, but with a small learning rate and a limited number of iterations, aiming to absorb real information to correct the model's understanding without destroying the learned strategy. The newly generated optimization strategy after fine-tuning is parsed into instructions and executed. Continuously monitor the deviation rate: if the deviation returns to the tolerance interval, the fine-tuning pauses; if the deviation continues to exceed the standard, further fine-tuning may be continued or deeper model calibration may be triggered. This process forms a high-level adaptive closed loop, continuously correcting the model to adapt to the dynamics of the real world.
[0100] Optionally, the online fine-tuning signal for triggering the reinforcement learning model includes:
[0101] According to the building operation mode switching signal, switch the dynamic threshold of the deviation rate;
[0102] Specifically, building operation mode signals are received, such as working hours, vacant mode, pre-treatment, high-density activities, holidays, etc. Different deviation rate thresholds are preset for each mode. For example, in the "working hours mode", in order to ensure stability and comfort, the tolerance for energy consumption deviation can be set relatively high to avoid frequent fine-tuning. While in the "vacant mode", in order to utilize the low-interference period to finely calibrate the model, the threshold can be set relatively low. When the operation mode is switched, the preset threshold standard of the new mode is automatically adopted to determine whether the deviation exceeds the standard and decide whether to trigger online fine-tuning.
[0103] During the low-energy consumption period at night, lower the dynamic threshold of the deviation rate;
[0104] Specifically, it is first necessary to be able to identify that the building has entered the low-energy consumption operation mode at night, which can be achieved by, for example, querying a preset schedule or determining that the current overall energy load level of the building is lower than a certain predetermined value. Once it is confirmed to enter this mode, the deviation rate thresholds of the key parameters used to trigger online fine-tuning will be automatically lowered to a more stringent level. For example, if the absolute deviation threshold of temperature during working hours is set to plus or minus 0.5 degrees Celsius, then in the low-energy consumption mode at night, this threshold can be reduced to plus or minus 0.3 degrees Celsius. Similarly, the deviation thresholds of other parameters such as energy consumption may also be lowered accordingly. The purpose of adopting this strategy is to make full use of the favorable conditions that the internal load of the building is relatively stable at night and the disturbance factors from the external environment are also less, and to calibrate the reinforcement learning model and its dependent digital twin model more precisely. By capturing and correcting more subtle prediction deviations, the prediction accuracy of the model under low-load conditions can be effectively improved, which can also lay a more solid foundation for the optimization control under complex working conditions during the day.
[0105] During the peak electricity consumption period, expand the dynamic threshold of the deviation rate.
[0106] Specifically, it is necessary to be able to identify whether the current period belongs to the peak electricity consumption period of the power grid, for example, it can be judged based on a preset time period or the peak-time electricity price signal received from the power company. During this peak period, in order to give priority to ensuring the stable operation of electricity, avoid triggering a high demand electricity charge mechanism, and reduce the disturbance that frequent control adjustments may cause to energy equipment and indoor users, the deviation rate thresholds used to trigger online fine-tuning will be raised to a relatively loose level. For example, if the energy consumption deviation threshold during normal working hours is set to 15%, it may be temporarily raised to 25% during the peak electricity consumption period. This means that unless there is a very significant deviation between the prediction of the digital twin model and the physical reality, the triggering of online fine-tuning will tend to be suppressed, allowing the current optimization strategy being executed to operate stably. At this time, the priority is to ensure "no mistakes" rather than pursuing the ultimate "optimal" performance.
[0107] Optionally, the parsing of the multi-objective optimization strategy into a device control instruction set includes:
[0108] Identifying the communication interface types of different device protocols and constructing an instruction format conversion template library;
[0109] Specifically, first, a device asset information library needs to be established to record the detailed information of all controllable energy devices in the building, such as chillers, pumps, fans, valves, variable air volume (VAV) terminals, and intelligent lighting fixtures. The content of the information library should include the model, manufacturer, and supported communication protocols of each device, such as common building automation protocols like BACnet, Modbus, KNX, or DALI. It also needs to contain the network address information of the device, such as IP address, device ID, or port number, as well as the specific point information required for control and status monitoring, such as the object identifier of BACnet or the register address of Modbus. Based on this information library, for each supported communication protocol and common device control actions, such as reading or writing binary or analog values, changing setpoints, or switching operation modes, standardized instruction format templates are pre-created. For example, a template for writing an analog value for the BACnet protocol may contain placeholders for fields such as protocol data unit (PDU) type, service type, object identifier, attribute identifier, priority, and the value to be written. This pre-constructed instruction format conversion template library is the basis for achieving unified control and interoperability of various heterogeneous devices in the building.
[0110] Matching the control protocol characteristics of the target device in the instruction format conversion template library to generate an adapted binary instruction block;
[0111] Specifically, when the reinforcement learning model outputs a high-level optimization strategy, such as the instruction "Set the VAV box damper opening in area Z01 to 60%", the parsing module first identifies the controlled target device, i.e., the VAV box in area Z01, its communication protocol, assumed to be BACnet / IP, and the point to be written, such as the analog output object AO-1 corresponding to the damper opening. Then, a template applicable to "BACnet / IP writing analog value" is selected from the template library. Information such as the specific device address, object identifier such as AO-1, attribute identifier such as PresentValue, priority such as 8, and the target value 60.0 to be written is filled into the corresponding fields of the template. Finally, according to the selected protocol specification, such as the BACnet protocol, the filled template is encoded into a binary data packet that conforms to the protocol standard, i.e., an application protocol data unit (APDU), for network transmission.
[0112] Dynamically queue-sort the sending order of instruction blocks according to the device response latency requirements.
[0113] Specifically, considering that not all energy devices respond to control instructions at the same speed, and some complex control logics may have strict timing requirements. For example, it is necessary to open the water valve first and then start the associated fan. An information library should be maintained, which contains typical response latency information of different device types or specific device instances, or set priorities for them. When there are multiple control instructions to be sent to different devices simultaneously or within a very short time, it is not necessarily simply sent in the order of generation. Instead, the generated binary instruction blocks are first placed in a queue to be sent. Then, the queue manager dynamically sorts the instruction blocks in the queue according to a series of factors. These factors include, for example, the expected response latency of the target devices of each instruction. The sorting strategy may prioritize sending to devices with fast responses to observe the effects as soon as possible, or prioritize sending to devices with slow responses to ensure their early start, depending on the control objective at that time, the urgency of the instruction itself. For example, safety-related instructions have the highest priority, and the possible control logic dependencies between instructions. Finally, the binary instruction blocks are sent to the target devices for execution one by one in the order of the dynamic sorting.
[0114] Exemplarily, assume that at a certain moment, the optimization strategy determines that it is necessary to simultaneously adjust the chilled water outlet temperature setpoint of the central chiller and the damper openings of VAV boxes distributed on multiple floors. The former is controlled, for example, through the Modbus TCP protocol, and its response is usually slow, and instructions need to be sent in advance; the latter is controlled, for example, through the BACnet / IP protocol, and the response is relatively fast. The parsing module generates a Modbus instruction block and ten BACnet instruction blocks accordingly. Subsequently, the instruction queue manager places the Modbus instruction block at the head of the sending queue according to the preset sorting rules. For example, the rules are set to require that instructions for controlling central system-level devices be sent first to guide the overall operation trend, and the ten BACnet instruction blocks are arranged behind it. Finally, the instructions are sent in this sorted order. This can ensure that the central chiller starts to adjust its outlet temperature first, and then the VAV boxes on each floor can accurately adjust the air volume according to the changed water temperature and the real-time demands of their respective areas, thus achieving a smoother and more coordinated control state transition process.
[0115] Based on the same inventive concept, the present invention also provides a building energy consumption dynamic optimization system based on BIM and reinforcement learning, as Figure 5 shown, the system includes:
[0116] A digital twin construction module for establishing a BIM containing physical property parameters of building components, integrating real-time monitoring data, and generating a building digital twin reflecting the evolution of dynamic thermal properties;
[0117] A state space construction module for extracting spatial topological relationships and component physical property parameters based on the building digital twin and constructing a multi-dimensional state space of a preset reinforcement learning model;
[0118] A policy generation module for embedding physical constraint conditions from the building digital twin into the reinforcement learning model and training the reinforcement learning model based on the physical constraint conditions and a preset composite reward function to generate a multi-objective optimization policy for energy equipment;
[0119] A control and feedback module for parsing the multi-objective optimization policy into a device control instruction set, controlling the energy equipment to execute parameter adjustment, and feeding back the adjustment result of the energy equipment to the building digital twin for real-time physical property simulation.
[0120] It should be noted that the functional division and information interaction among the above-mentioned modules are logical, and physically they can be integrated on the same software platform or distributedly deployed. The connections between them represent data flow and control flow, aiming to collaboratively achieve the goal of dynamic optimization of building energy consumption in the present invention. The above is only an exemplary embodiment of the present invention, and the protection scope of the present invention cannot be limited thereby.
Claims
1. A building energy consumption dynamic optimization method based on BIM and reinforcement learning, characterized in that, The method includes: Establishing a BIM containing physical property parameters of building components and integrating real-time monitoring data to generate a building digital twin reflecting the evolution of dynamic thermal properties; wherein, generating the building digital twin reflecting the evolution of dynamic thermal properties includes: obtaining an aging coefficient data set of building materials, establishing a dynamic heat conduction model associated with seasonal changes; identifying the solar radiation exposure coefficient of outdoor facade components, and generating a light reflectance attenuation curve in combination with historical dust deposition monitoring data; updating the thermal property parameters of the components in the BIM according to the dynamic heat conduction model and the light reflectance attenuation curve; Extracting spatial topological relationships and component physical property parameters from the building digital twin to construct a multi-dimensional state space of a preset reinforcement learning model; wherein, constructing the multi-dimensional state space of the preset reinforcement learning model includes: partitioning a building structure topological network from the building digital twin, and extracting an energy conversion efficiency matrix of heating, ventilation, and air conditioning equipment; collecting air pressure gradient data and historical equipment operation logs to generate a probability distribution model of environmental load disturbance terms; fusing the topological network, the efficiency matrix, and the probability distribution model into a discrete parameter set; wherein, collecting air pressure gradient data and historical equipment operation logs to generate a probability distribution model of environmental load disturbance terms includes: collecting air pressure gradient data and historical equipment operation logs; analyzing the dynamic characteristics of the building internal environment according to the air pressure gradient data and performing micro-environment zoning; calculating the equipment energy consumption fluctuation threshold in each zone according to the historical equipment operation logs; if the real-time energy consumption data continuously exceeds the equipment energy consumption fluctuation threshold, triggering dynamic reorganization of the state space, temporarily adding an abnormal flag bit, and increasing the update frequency or accuracy of the state parameters in this zone; Embedding physical constraint conditions from the building digital twin into the reinforcement learning model, and training the reinforcement learning model based on the physical constraint conditions and a preset composite reward function to generate a multi-objective optimization strategy for energy equipment; Parsing the multi-objective optimization strategy into a device control instruction set, controlling the energy equipment to perform parameter adjustment, and feeding back the adjustment result of the energy equipment to the building digital twin for real-time physical property simulation.
2. The building energy consumption dynamic optimization method based on BIM and reinforcement learning according to claim 1, wherein, The embedding the physical constraint conditions of the building digital twin into the reinforcement learning model includes: Generating a constraint rule set prohibiting equipment from operating overloaded based on the stress simulation results of the BIM; During the reinforcement learning training process, generating and evaluating an action sequence representing a candidate strategy, and intercepting the candidate strategy that violates the constraint rule set; When a conflict is detected, modifying the weight allocation of the preset composite reward function to guide strategy convergence.
3. The building energy consumption dynamic optimization method based on BIM and reinforcement learning according to claim 2, characterized in that, The modifying the weight allocation of the preset composite reward function includes: Obtaining data on conflicts between historical optimization strategies and physical constraints, and analyzing and constructing a conflict type - correction coefficient mapping table based on the data; When a new strategy conflict occurs, calling the corresponding correction coefficient in the conflict type - correction coefficient mapping table to adjust the weight allocation related to comfort and energy consumption in the preset composite reward function.
4. The building energy consumption dynamic optimization method based on BIM and reinforcement learning according to claim 1, characterized in that, The feedback of the adjustment result of the energy device to the building digital twin for real-time physical property simulation includes: Collect the actual operating parameters of the energy device and calculate the deviation rate from the simulation result of the building digital twin; When the deviation rate exceeds the preset threshold, trigger the online fine-tuning signal of the reinforcement learning model and perform the fine-tuning of the reinforcement learning; Re-inject the optimized strategy after fine-tuning into the device control instruction set until the deviation rate returns to the tolerance interval.
5. The building energy consumption dynamic optimization method based on BIM and reinforcement learning according to claim 4, characterized in that The triggering of the online fine-tuning signal of the reinforcement learning model includes: According to the building operation mode switching signal, switch the dynamic threshold of the deviation rate; During the low-energy consumption period at night, reduce the dynamic threshold of the deviation rate; During the peak electricity consumption period, expand the dynamic threshold of the deviation rate.
6. The building energy consumption dynamic optimization method based on BIM and reinforcement learning according to claim 1, characterized in that The parsing of the multi-objective optimization strategy into the device control instruction set includes: Identify the communication interface types of different device protocols and construct an instruction format conversion template library; Match the control protocol characteristics of the target device in the instruction format conversion template library and generate an adapted binary instruction block; Dynamically queue and sort the sending order of the instruction blocks according to the device response delay requirement.
7. A building energy consumption dynamic optimization system based on BIM and reinforcement learning, which is applied to execute the building energy consumption dynamic optimization method based on BIM and reinforcement learning according to any one of claims 1 to 6, characterized in that, The system includes: A digital twin construction module for establishing a BIM containing the physical property parameters of building components and integrating real-time monitoring data to generate a building digital twin reflecting the evolution of dynamic thermal properties; among them, the generation of the building digital twin reflecting the evolution of dynamic thermal properties includes: obtaining the aging coefficient data set of building materials, establishing a dynamic heat conduction model associated with seasonal changes; identifying the solar radiation exposure coefficient of outdoor facade components, and generating a light reflectance attenuation curve in combination with historical dust deposition monitoring data; updating the thermal property parameters of the components in the BIM according to the dynamic heat conduction model and the light reflectance attenuation curve; A state space construction module for extracting the spatial topological relationship and component physical property parameters based on the building digital twin to construct a multi-dimensional state space of a preset reinforcement learning model; among them, the construction of the multi-dimensional state space of the preset reinforcement learning model includes: dividing the building structure topological network from the building digital twin, and extracting the energy conversion efficiency matrix of the HVAC equipment; collecting the air pressure gradient data and historical device operation logs, and generating a probability distribution model of the environmental load disturbance term; fusing the topological network, the efficiency matrix and the probability distribution model into a discrete parameter set; among them, the collection of the air pressure gradient data and historical device operation logs, and the generation of the probability distribution model of the environmental load disturbance term includes: collecting the air pressure gradient data and historical device operation logs; analyzing the dynamic characteristics of the building internal environment according to the air pressure gradient data and performing micro-environment zoning; calculating the device energy consumption fluctuation threshold in each zone according to the historical device operation logs; if the real-time energy consumption data continuously exceeds the device energy consumption fluctuation threshold, trigger the dynamic reorganization of the state space, temporarily add an abnormal flag bit, and increase the update frequency or accuracy of the state parameters in this zone; A strategy generation module, configured to embed physical constraint conditions from the building digital twin into the reinforcement learning model, and train the reinforcement learning model based on the physical constraint conditions and a preset composite reward function to generate a multi-objective optimization strategy for energy equipment; A control and feedback module, configured to parse the multi-objective optimization strategy into a device control instruction set, control the energy equipment to perform parameter adjustment, and feedback the adjustment result of the energy equipment to the building digital twin for real-time physical property simulation.
Citation Information
Patent Citations
Building thermal performance intelligent energy-saving analysis method based on BIM
CN118981816A
Intelligent supervision method for digital power system based on cloud computing
CN119727137A