Building energy consumption dynamic optimization method and system based on BIM and reinforcement learning

By applying BIM and reinforcement learning methods in building energy consumption management, dynamic building digital twins are constructed and optimized, the problems of staticity and model mismatch of energy consumption management in traditional methods are solved, and efficient, safe and comfortable building energy consumption management is achieved.

CN120145880AActive Publication Date: 2025-06-13ZHONGQI JIAOJIAN GRP

Patent Information

Application Number
CN202510621855.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The static nature of traditional building energy consumption management methods cannot be adjusted dynamically, resulting in energy waste or a decrease in user comfort, and existing model-based methods are difficult to accurately reflect the dynamic evolution of building physical properties.

Method used

Using a method based on BIM and reinforcement learning, a dynamic building digital twin with real-time data is constructed, and through reinforcement learning training and closed-loop feedback optimization, a multi-objective optimization strategy for building energy consumption that is physically feasible and dynamically adapted to environmental changes is generated.

Benefits of technology

It significantly improves energy utilization efficiency, system operation safety and user comfort, achieves the accuracy and long-term effectiveness of energy consumption optimization, and ensures the physical feasibility of control strategies and system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145880A_ABST
    Figure CN120145880A_ABST
Patent Text Reader

Abstract

The invention discloses a building energy consumption dynamic optimization method and system based on BIM and reinforcement learning, and belongs to the technical field of building energy management and intelligent control, and the method comprises the steps: building a BIM containing building component physical attribute parameters, and generating a building digital twinborn body with dynamic thermal attribute evolution; extracting the spatial topological relation and the physical property parameters of the components, and constructing a multi-dimensional state space of a preset reinforcement learning model; embedding physical constraint conditions, and training the reinforcement learning model to generate a multi-objective optimization strategy of the energy equipment; and analyzing the multi-objective optimization strategy into an equipment control instruction set, and feeding back the equipment control instruction set to the building digital twin for real-time physical attribute simulation. According to the method, the physical accuracy of the BIM and the self-adaptive decision-making ability of reinforcement learning are combined, adversarial training under physical constraints is introduced, an energy consumption optimization strategy which conforms to actual operation limitation and dynamically adapts to environmental changes can be generated, and the energy utilization efficiency and the system response speed are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of building energy management and intelligent control, and particularly to a building energy consumption dynamic optimization method and system based on BIM and reinforcement learning. Background Art

[0002] Traditional building energy consumption management mostly adopts preset strategies or simple rule-based control, such as the start and stop of heating, ventilation, and air conditioning (HVAC) based on a fixed schedule. The core defect of such methods lies in their static nature, which cannot dynamically adjust according to real-time environmental changes, human activities, and equipment status, often resulting in unnecessary energy waste or sacrificing user comfort, and it is difficult to meet the requirements of refined management of modern buildings.

[0003] To improve the optimization effect, existing technologies have introduced model-based methods such as model predictive control (MPC). However, the building models relied on by these methods are often static or quasi-static, and it is difficult to accurately reflect the dynamic evolution of physical properties such as building material aging, exterior wall pollution, and equipment efficiency decay. This model mismatch problem leads to a decline in the accuracy of optimal control over time and cannot achieve long-term optimal energy consumption performance. Summary of the Invention

[0004] To solve the above problems, the present invention provides a building energy consumption dynamic optimization method and system based on BIM and reinforcement learning. By adopting the mechanisms of constructing a dynamic building digital twin integrating real-time data, embedding physical constraints for reinforcement learning training, and closed-loop feedback optimization, it can generate a multi-objective optimization strategy for building energy consumption that is physically feasible and dynamically adapts to environmental changes, significantly improving energy utilization efficiency, system operation safety, and user comfort.

[0005] The above objectives can be achieved through the following solutions: A building energy consumption dynamic optimization method based on BIM and reinforcement learning, including establishing a BIM containing physical property parameters of building components and integrating real-time monitoring data to generate a building digital twin reflecting the evolution of dynamic thermal properties; extracting spatial topological relationships and component physical property parameters from the building digital twin to construct a multi-dimensional state space of a preset reinforcement learning model; embedding physical constraint conditions from the building digital twin into the reinforcement learning model, and training the reinforcement learning model based on the physical constraint conditions and a preset composite reward function to generate a multi-objective optimization strategy for energy equipment; parsing the multi-objective optimization strategy into a device control instruction set, controlling the energy equipment to perform parameter adjustment, and feeding back the adjustment results of the energy equipment to the building digital twin for real-time physical property simulation.

[0006] Optionally, the generation of the building digital twin reflecting the evolution of dynamic thermal properties includes: obtaining an aging coefficient dataset of building materials and establishing a dynamic heat conduction model associated with seasonal changes; identifying the solar radiation exposure coefficient of outdoor facade components and generating a light reflectance attenuation curve by combining historical dust deposition monitoring data; and updating the thermal property parameters of components in the BIM according to the dynamic heat conduction model and the light reflectance attenuation curve.

[0007] Optionally, the construction of the multi-dimensional state space for reinforcement learning includes: partitioning the building structure topology network from the building digital twin and extracting the energy conversion efficiency matrix of HVAC equipment; collecting air pressure gradient data and historical equipment operation logs to generate a probability distribution model of environmental load disturbance terms; and fusing the topology network, the efficiency matrix, and the probability distribution model into a discrete parameter set.

[0008] Optionally, the collection of air pressure gradient data and historical equipment operation logs to generate a probability distribution model of environmental load disturbance terms includes: collecting air pressure gradient data and historical equipment operation logs; analyzing the dynamic characteristics of the building internal environment according to the air pressure gradient data and performing micro-environment zoning; calculating the equipment energy consumption fluctuation threshold in each zone according to the historical equipment operation logs; and triggering the discrete parameter set when the real-time energy consumption data exceeds the equipment energy consumption fluctuation threshold.

[0009] Optionally, the embedding of the physical constraint conditions of the building digital twin in the reinforcement learning model includes: generating a constraint rule set that prohibits equipment from operating overloaded based on the stress simulation results of the BIM; generating and evaluating an action sequence representing a candidate policy during the reinforcement learning training process and intercepting the candidate policy that violates the constraint rule set; and when a conflict is detected, correcting the weight assignment of the preset composite reward function to guide policy convergence.

[0010] Optionally, the correction of the weight assignment of the preset composite reward function includes: obtaining data on the conflict between the historical optimized policy and the physical constraint, analyzing and constructing a conflict type - correction coefficient mapping table based on the data; and when a new policy conflict occurs, calling the corresponding correction coefficient in the conflict type - correction coefficient mapping table to adjust the weight assignment related to comfort and energy consumption in the preset composite reward function.

[0011] Optionally, after the adjustment result is fed back to the building digital twin for real-time physical property simulation, it includes: collecting the actual operation parameters of the energy equipment and calculating the deviation rate from the simulation results of the building digital twin; when the deviation rate exceeds the preset threshold, triggering an online fine-tuning signal of the reinforcement learning model and performing fine-tuning of the reinforcement learning; and re-injecting the fine-tuned optimized policy into the equipment control instruction set until the deviation rate returns to the tolerance interval.

[0012] Optionally, the online fine-tuning signal for triggering the reinforcement learning model includes: switching the dynamic threshold of the deviation rate according to the building operation mode switching signal; reducing the dynamic threshold of the deviation rate during the low-energy consumption period at night; and expanding the dynamic threshold of the deviation rate during the peak electricity consumption period.

[0013] Optionally, parsing the multi-objective optimization strategy into a device control instruction set includes: identifying the communication interface types of different device protocols and constructing an instruction format conversion template library; matching the control protocol characteristics of the target device in the instruction format conversion template library to generate an adapted binary instruction block; and dynamically queueing the sending order of the instruction blocks according to the device response delay requirements.

[0014] Based on the same inventive concept, the present invention also provides a building energy consumption dynamic optimization system based on BIM and reinforcement learning. The system includes: a digital twin construction module for establishing a BIM containing physical attribute parameters of building components and fusing real-time monitoring data to generate a building digital twin reflecting the evolution of dynamic thermal properties; a state space construction module for extracting spatial topological relationships and component physical property parameters based on the building digital twin to construct a multi-dimensional state space of a preset reinforcement learning model; a policy generation module for embedding physical constraint conditions from the building digital twin into the reinforcement learning model and training the reinforcement learning model based on the physical constraint conditions and a preset composite reward function to generate a multi-objective optimization strategy for energy equipment; and a control and feedback module for parsing the multi-objective optimization strategy into a device control instruction set, controlling the energy equipment to execute parameter adjustments, and feeding back the adjustment results of the energy equipment to the building digital twin for real-time physical property simulation.

[0015] Compared with the prior art, the present invention has the following advantages: 1. It improves the accuracy and long-term effectiveness of energy consumption optimization. By constructing a dynamic building digital twin that fuses real-time monitoring data, it can reflect the dynamic evolution of building physical properties in real time, overcoming the mismatch problem caused by traditional static models, providing a more accurate environment model for reinforcement learning, and thus improving the accuracy and long-term effect of the optimization strategy.

[0016] 2. It ensures the physical feasibility and system security of the control strategy. By embedding physical constraint conditions (such as equipment operation limitations, structural safety limitations, etc.) from BIM into the reinforcement learning training process, it ensures that the generated energy consumption optimization strategy is physically feasible, avoids damage to energy equipment and potential safety risks, and improves the stability and reliability of the system.

[0017] 3. Achieved effective balance and adaptive optimization of multiple objectives. By leveraging the decision-making ability of reinforcement learning and the mechanism of a preset composite reward function, it can dynamically balance multiple optimization objectives such as energy conservation, user comfort, and equipment health according to actual needs. Combining the closed-loop feedback and online fine-tuning mechanism enables the system to continuously self-optimize based on the actual operation effect, dynamically adapting to environmental changes and the evolution of building states.

[0018] 4. Expanded the application value of BIM in the building operation stage, upgraded BIM from a traditional static information database to the core foundation for dynamic optimization control, and fully exploited the potential of BIM in the whole life cycle management of buildings, especially in the intelligent and refined operation stage, through the deep integration with real-time data and intelligent algorithms.

[0019] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 is a schematic flowchart of the building energy consumption dynamic optimization method and system based on BIM and reinforcement learning according to an embodiment of the present invention.

[0022] Figure 2 is a regional temperature heat map according to an embodiment of the present invention.

[0023] Figure 3 is a three-dimensional temperature scatter plot according to an embodiment of the present invention.

[0024] Figure 4 is a deviation rate and dynamic threshold graph according to an embodiment of the present invention.

[0025] Figure 5 is a schematic structural diagram of the building energy consumption dynamic optimization method and system based on BIM and reinforcement learning according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0027] Referring to Figure 1 , an embodiment of the present invention proposes a building energy consumption dynamic optimization method based on BIM and reinforcement learning, aiming to construct a building digital twin by integrating building information models and reinforcement learning technologies, and on this basis, realize real-time, dynamic, and closed-loop optimization of the operation strategies of energy equipment, so as to minimize building energy consumption while meeting comfort requirements.

[0028] The method of this embodiment specifically includes the following steps: Establish a BIM containing the physical property parameters of building components, and integrate real-time monitoring data to generate a building digital twin reflecting the evolution of dynamic thermal properties; Specifically, the building information model not only includes the geometric information of the building, such as the positions and dimensions of walls, floors, and doors and windows, but also needs to input the physical property parameters of key components, such as the thermal conductivity, specific heat capacity, density, solar radiation absorption rate, reflectivity, etc. of materials. Integrating real-time monitoring data means inputting the real-time data streams obtained from sensors deployed inside and outside the building, such as temperature, humidity, light intensity, carbon dioxide concentration, and occupancy sensors, as well as weather forecast data from meteorological services, into a preset digital model to construct a building digital twin. The building digital twin is a dynamic and real-time mirror of the physical building, and its core lies in simulating the internal environmental parameters of the building, such as the temperature field and humidity field distribution, through physical models and data-driven models, such as models established based on the energy balance equation and heat transfer principles, as Figure 2 shown, as well as the thermal properties of components, such as the thermal conductivity and reflectivity considering the effects of aging and dust accumulation, and the real state changing with time and external conditions. Its purpose is to provide a high-fidelity virtual environment for subsequent energy consumption simulation and optimization strategy formulation, as Figure 3 shown.

[0029] Extract the spatial topological relationship and component physical property parameters based on the building digital twin to construct the multi-dimensional state space of a preset reinforcement learning model; Specifically, a multi-dimensional state space for reinforcement learning is constructed based on the building digital twin. This process includes extracting key spatial topological relationships, such as functional partitions, connectivity, and equipment service relationships, as well as physical property parameters of core components, such as dynamic heat transfer coefficients, air permeabilities, and equipment energy efficiencies. These extracted attributes, combined with real-time environmental parameters such as indoor and outdoor temperature and humidity, personnel and air quality indicators, and equipment operating states such as switch states, set points, and power consumption, are jointly integrated to construct a multi-dimensional state vector or tensor. This state space aims to comprehensively characterize the key factors affecting building energy consumption and comfort, serving as the basis for the reinforcement learning agent to perceive the environment.

[0030] Embed physical constraint conditions from the building digital twin into the reinforcement learning model, and train the reinforcement learning model based on the physical constraint conditions and a preset composite reward function to generate a multi-objective optimization strategy for energy equipment; Specifically, the physical constraint conditions are used to guide the reinforcement learning model to explore within a safe and reasonable range, and they are derived from digital twin physical simulations or equipment specifications. For example, physical constraints may include equipment operation limits such as start-stop frequencies and power ranges, safety thresholds such as pipeline pressures and supply air temperature and humidity, and temperature requirements or structural load-bearing limits in specific areas such as data rooms. The preset composite reward function (Reward Function) is the key to guiding the reinforcement learning goal, and it is usually a weighted sum of multiple sub-goals. For example: , where, represents the instant total reward obtained by the agent, represents the energy consumption cost, is the user comfort score, is the indoor air quality score. , , are the weight coefficients corresponding to these three items respectively, used to balance the importance of different goals, is the penalty value for behaviors that violate physical constraints or cause adverse consequences, such as frequent equipment start-stop. The comfort score can be calculated based on the Predicted Mean Vote (PMV) model, the Predicted Percentage of Dissatisfied (PPD) model, or user feedback, and the air quality score can be evaluated based on indicators such as CO2 concentration. During the training process, the agent interacts with the digital twin environment, using deep reinforcement learning algorithms to learn to maximize the long-term cumulative reward while satisfying physical constraints, and accordingly generates optimized control strategies for energy equipment such as air conditioners, lighting, and fresh air, including but not limited to recommended temperature set points, air volumes, or switch times.

[0031] Parse the multi-objective optimization strategy into a device control instruction set, control the energy device to perform parameter adjustment, and feed back the adjustment result of the energy device to the building digital twin for real-time physical property simulation.

[0032] Specifically, the high-level optimization strategy output by the reinforcement learning model needs to be parsed into a device control instruction set for execution. This process involves converting the policy requirements into low-level commands and determining specific instruction parameters such as register values. The generated instructions are sent to the energy device actuator such as a valve or a frequency converter through the Building Automation System (BAS). After execution, the actual operating status collected by the sensor such as power consumption, temperature, and environmental change data is used on the one hand to update the digital twin to keep it synchronized, and on the other hand as the real feedback of the reinforcement learning, including the next state and the actual reward. This feedback drives the agent to continuously make decisions and learn, forming a closed-loop optimization process of "perception - decision - control - feedback - learning".

[0033] By integrating the high-precision BIM, real-time data, and reinforcement learning algorithm, constructing the building digital twin and performing closed-loop optimization can achieve refined and intelligent management of building energy consumption, and significantly improve energy utilization efficiency and user comfort experience.

[0034] Optionally, the generation of the building digital twin reflecting the evolution of dynamic thermal properties includes: Obtain the aging coefficient data set of building materials and establish a dynamic heat conduction model associated with seasonal changes; Specifically, to reflect the evolution of dynamic thermal properties, the method includes obtaining the aging coefficient data set of building materials and establishing a dynamic heat conduction model associated with seasonal changes. The aging coefficient can be obtained from material suppliers, industry standard databases, or accelerated aging experiments, which quantifies the laws of changes in key material properties such as thermal conductivity and airtightness over time and environmental factors such as temperature, humidity, and radiation. When establishing a dynamic heat conduction model based on this, the actual service life of the material and the annual average temperature representing seasonal environmental impact factors and the annual average relative humidity etc. can be used as input variables of the model, and these input variables are used to correct the basic thermal conductivity of the material in the initial state . For example, the dynamic thermal conductivity of a specific insulation material can be expressed as a function of service life , annual average temperature and annual average relative humidity . An exemplary model formula is as follows: , wherein, represents the dynamic thermal conductivity value under specific conditions, is the basic thermal conductivity of the material, is the service life, and are the annual average ambient temperature and relative humidity respectively, and are empirical aging coefficients related to the material type, jointly describing the change trend of performance over time (where the coefficient is usually less than 1), and are correction coefficients considering the effects of temperature and humidity respectively, while and are the reference temperature and humidity values used in the calculation. Using such a dynamic thermal model that includes the effects of aging and environmental factors can enable the building digital twin to more accurately simulate and reflect the actual dynamic change of the thermal insulation performance of the building envelope during long-term use, thus providing a more reliable basis for energy consumption analysis and optimization.

[0035] Identify the solar radiation exposure coefficient of outdoor facade components, and generate a light reflectance attenuation curve in combination with historical dust deposition monitoring data; Specifically, use the geographic information and solar trajectory tools of the BIM to calculate the solar radiation exposure coefficient of each component, and this calculation needs to consider the occlusion effect. At the same time, estimate the cumulative effective dust deposition per unit area of the facade , which can be achieved by analyzing the data of the deployed dust sensors or combining the local historical air quality data, such as PM2.5 and PM10 concentrations, with rainfall records, for example, by integrating the difference between the dust deposition and scouring rates per unit time. Based on the estimated , establish an attenuation model of the surface light reflectance, for example, assume that the reflectance decreases exponentially with the dust deposition:

[0036] wherein, R(t) represents the real-time surface light reflectance considering the influence of dust deposition, is the initial light reflectance of the surface of the component in the clean state, is the natural exponential function, is a light reflectance attenuation constant related to the surface characteristics of the material and the type of dust, while is the cumulative effective dust deposition amount corresponding to the time. Applying such a light reflectance attenuation model to update the surface properties of components can enable the building digital twin to more realistically simulate the change of the heat absorption of the exterior wall due to dust coverage over time, thereby improving the accuracy of the overall building energy consumption simulation.

[0037] Update the thermal property parameters of the components in the BIM according to the dynamic heat conduction model and the light reflectance attenuation curve.

[0038] Specifically, this parameter update process can be executed at a preset cycle, such as quarterly or annually, or triggered when the monitored aging / dust deposition index reaches a threshold. After being triggered, call the dynamic model to calculate the actual thermal properties of the current component, such as the dynamic thermal conductivity Kt and the reflectance Rt. Subsequently, write the calculated new parameter values into the corresponding component properties through the BIM software API or directly modify the digital twin database. This dynamic update mechanism ensures that the digital twin accurately reflects the long-term changes in the physical properties of the building, improving the accuracy of subsequent energy consumption simulation and prediction.

[0039] Exemplarily, for an office building located in the city center, its glass curtain wall uses a certain low-emissivity coated glass. Initially, the reflectance is set to 0.6 in the BIM. By analyzing the PM10 monitoring data and rainfall data in the past three years, the annual average effective dust deposition amount is calculated and denoted as Deff(3 years). Substitute this value and the assumed light reflectance attenuation constant Kd = 0.05 into the light reflectance attenuation model, and it is obtained that the actual reflectance R(3 years) of the glass surface has decayed from 0.6 to approximately 0.52 after three years. At the same time, the aging model of its sealant strip shows that the thermal conductivity has increased by 5%. Before conducting the annual energy consumption simulation and optimization strategy training, the system automatically updates the reflectance of the corresponding curtain wall component in the digital twin to 0.52, and the thermal conductivity of the sealed part is also adjusted accordingly. This enables the digital twin to more accurately predict the solar radiation heat entering through the curtain wall in summer and the cold air infiltration or heat loss in winter, so that the reinforcement learning model can formulate a more practical air-conditioning operation strategy and avoid energy waste caused by using outdated parameters.

[0040] Optionally, the multi-dimensional state space for constructing the reinforcement learning includes: Divide the building structure topology network from the building digital twin and extract the energy conversion efficiency matrix of the HVAC equipment; Specifically, according to the spatial data in the BIM and the HVAC system drawings, the building is divided into several thermodynamically or control-wise relatively independent zones. Meanwhile, the air flow and heat transfer paths between the zones, as well as the connection and service relationships between specific HVAC equipment and each zone, are identified, thereby constructing a building topology network in the form of a graphical structure or an adjacency matrix. In addition, the technical specifications of the main HVAC equipment need to be collected or fitted through measured data to obtain their energy conversion efficiencies under different operating conditions. The operating conditions involve, for example, conditions such as load rate, ambient temperature, and water temperature. For example, the coefficient of performance of a variable-frequency chiller can be expressed as a function of the chilled water outlet temperature , the cooling water inlet temperature , and the load rate . , , Similarly, the boiler thermal efficiency can be modeled as a function of the load rate and the return water temperature . , , The pump efficiency can then be modeled as a function of the flow rate and the head . , .

[0041] Here, represents the model function used to calculate the chiller , represents the model function used to calculate the boiler thermal efficiency, represents the model function used to calculate the pump efficiency, is the coefficient of performance for refrigeration, , are the boiler thermal efficiency and the pump efficiency respectively, , , are the chilled water outlet, cooling water inlet, and boiler return water temperatures respectively, is the load percentage, is the pump flow rate, is the pump head. These efficiency models can be in the form of matrices, polynomials, or neural networks, etc., and are an important part of the state space for evaluating the potential energy consumption of different control strategies.

[0042] Collect the air pressure gradient data and the historical equipment operation logs to generate a probability distribution model of the environmental load disturbance term; ​Specifically, air pressure difference data is collected by installing air pressure sensors at key positions inside and outside the building, such as on different floors and at the main entrances. Combining the historical operation logs containing equipment switch, duration, and energy consumption information, and the corresponding meteorological data such as wind speed and wind direction, the impact of air pressure gradient changes on infiltration, ventilation, and equipment load is analyzed. Using statistical methods such as kernel density estimation, Gaussian mixture models, or machine learning models such as hidden Markov models, a probability distribution model is established. This model is used to describe the occurrence probability of the disturbance term and its impact on the load, for example, outputting the magnitude and probability of increased load due to infiltration under specific conditions.

[0043] Fuse the topological network, the efficiency matrix, and the probability distribution model into a set of discrete parameters.

[0044] Specifically, integrate the aforementioned information into a multi-dimensional state vector S that can be understood by the agent. This vector includes key environmental parameters such as indoor and outdoor temperature and humidity, CO2, and radiation intensity; equipment status such as switch, mode, output percentage, and the current efficiency calculated by the efficiency model; and the current or predicted disturbance information obtained from the disturbance model. Continuous variables can be discretized as needed. The state vector S comprehensively describes the current environment and operating conditions, forming the basis for the agent's decision-making.

[0045] Exemplarily, consider the energy consumption optimization scenario of a large shopping mall building. The constructed topological network model shows that there is a significant heat exchange relationship between the atrium area and the surrounding store areas, and these two types of areas are served by two large air handling units (AHUs) together. The extracted equipment efficiency model, such as the model represented by the function Fchiller or Fboiler, further indicates that when the outdoor temperature is lower than 10 degrees Celsius, the heating efficiency of the AHU will decrease by 15%. At the same time, the probability distribution model of the environmental load disturbance term generated based on historical data shows that during the peak pedestrian flow period on weekend afternoons, there is an 80% probability that the carbon dioxide (CO2) concentration in the area near the main entrance of the shopping mall will exceed the safety or comfort threshold of 1000 ppm, which will lead to a significant increase in the required fresh air load. Therefore, the reinforcement learning state space S constructed for this shopping mall contains information in multiple dimensions, such as: the current temperature and humidity readings in the atrium and each store area, the real-time outdoor temperature, the current operating modes and output air volumes of the two AHUs, the current actual heating efficiency of the AHU queried from the efficiency model based on the real-time outdoor temperature, and the probability value of the CO2 concentration exceeding the standard in the entrance area output by the disturbance term model, etc. Based on this comprehensive and dynamic set of state information S, the reinforcement learning model can make predictive decisions. For example, before predicting the upcoming peak of pedestrian flow, it appropriately increases the introduction volume of the fresh air system in advance and intelligently selects to perform preheating or precooling operations during the period when the equipment operating efficiency is relatively high, so as to actively respond to the upcoming personnel load disturbance and the challenge of the equipment efficiency fluctuating with the external environment change.

[0046] Optionally, the generating of the probability distribution model of the environmental load disturbance term by collecting the barometric gradient data and the historical equipment operation logs includes: Collecting the barometric gradient data and the historical equipment operation logs; Specifically, the barometric gradient data is collected by deploying barometric sensors at representative measuring points inside and outside the building, recording the differential pressure data at a preset frequency to form a time series of the barometric field reflecting the effects of wind pressure, thermal pressure, etc. The historical equipment operation logs are collected by extracting the operation records of key energy-consuming equipment from the BAS, etc., including timestamps, status, set values, and measured performance parameters, providing a data basis for subsequent analysis.

[0047] Analyzing the dynamic characteristics of the building internal environment based on the barometric gradient data and performing microenvironment zoning; Specifically, in addition to monitoring the barometric gradient, micro wind speed sensors can also be deployed in key areas inside the building, or the calibrated CFD simulation in the digital twin can be used to monitor or simulate the internal air flow velocity and direction. According to the characteristics such as air flow stability, temperature uniformity, and pollutant diffusion mode, the large area is subdivided into microenvironment zones with different dynamic characteristics. For example, the open office area can be divided into the near-window area, the internal stable area, and the downstream area of the air outlet.

[0048] Calculate the equipment energy consumption fluctuation thresholds within each partition according to the historical equipment operation logs; Specifically, for each microenvironment partition, screen the operation logs of the service equipment. Analyze the temporal fluctuation statistical characteristics of the partition equipment energy consumption under similar external conditions and personnel patterns, such as calculating the standard deviation or interquartile range. Set dynamic energy consumption fluctuation thresholds based on the statistics, such as calculating by adding or subtracting several times the standard deviation from the recent average energy consumption, to define the normal energy consumption fluctuation range of the partition.

[0049] When the real-time energy consumption data exceeds the equipment energy consumption fluctuation threshold, trigger the discrete parameter set.

[0050] Specifically, monitor the energy consumption of the equipment in each microenvironment partition in real time and compare it with the preset dynamic threshold. If the energy consumption continuously exceeds the threshold, for example, for several consecutive sampling periods, it is judged that an unexpected disturbance has occurred, such as a window being opened, a sudden increase in personnel, or a precursor to equipment failure. At this time, trigger the state space dynamic reorganization: temporarily add an abnormal flag bit, increase the update frequency or accuracy of the partition state parameters, or activate a fine-grained local disturbance sub-model to replace / supplement the global model prediction, and incorporate the results into the state space. This reorganization aims to enable the reinforcement learning agent to perceive local anomalies faster and more accurately and quickly adjust the strategy to respond.

[0051] Exemplarily, consider an application scenario of an intelligent ward floor. Among them, the ward near the balcony door is identified and divided into an independent microenvironment partition. According to the analysis of historical operation data, under normal circumstances where there is no personnel activity at night and the doors and windows are confirmed to be closed, the standard deviation of the energy consumption fluctuation of the fan coil unit serving this ward is 5 watts. Based on this, the dynamic fluctuation threshold of the energy consumption of this partition is set to float up and down 10 watts from the normal average value. One night, real-time monitoring found that the energy consumption of the fan coil in this ward suddenly exceeded its night average level by 30 watts, and this high energy consumption state lasted for five minutes, significantly exceeding the preset positive 10-watt fluctuation threshold upper limit. Based on this, it is judged that an abnormal situation may have occurred in this microenvironment partition, such as the balcony door not being properly closed, thus triggering the state space dynamic reorganization mechanism. The specific reorganization operation may be to specifically mark the state of this ward as "suspected high infiltration load" in the current state vector of the reinforcement learning model, and temporarily increase the weight of the temperature reading of this ward in the calculation of the state vector or reward function. After receiving this reorganized state information containing abnormal indications, the reinforcement learning agent may preferentially choose to perform control actions such as increasing the air supply volume to this ward or adjusting the temperature set point to try to maintain the temperature stability in the ward. At the same time, it may also send a prompt message about this potential abnormal situation to the building operation and maintenance management platform. This dynamic reorganization and response mechanism enables the overall optimization to respond more sensitively to local, sudden disturbance events.

[0052] Optionally, the physical constraint conditions for embedding the building digital twin in the reinforcement learning model include: Based on the stress simulation results of the BIM, generate a constraint rule set that prohibits equipment from operating overloaded; Specifically, using the equipment location, weight, and structure data provided by the BIM, combined with finite element analysis software, simulate the static and dynamic loads of large or vibration-sensitive equipment under different working conditions on the structure. According to the simulation results, design specifications, and safety standards, determine the safety boundary conditions, such as the maximum allowable vibration amplitude of the floor slab or the maximum instantaneous stress of the pipeline. Convert these boundary conditions into clear constraint rule sets, such as specifying the maximum operating frequency or maximum starting acceleration of the equipment.

[0053] During the reinforcement learning training process, generate and evaluate action sequences representing candidate policies, and intercept the candidate policies that violate the constraint rule set; Specifically, in the reinforcement learning training session, it is usually carried out in the digital twin simulation environment. When the agent generates a candidate action sequence, it is reviewed through a "safety layer" or a "constraint check module". This module predicts the future state after the execution sequence, may call the digital twin simulation, and compare it with the constraint rule set. If the prediction result shows that a constraint will be violated, such as excessive vibration, then intercept the action sequence or policy, such as preventing its execution or giving a very large negative reward, to prevent the agent from learning dangerous policies.

[0054] When a conflict is detected, modify the weight distribution of the preset composite reward function to guide the policy convergence.

[0055] Specifically, when a candidate action is intercepted due to a predicted violation of physical constraints, it indicates that the policy tends to explore dangerous areas. To guide learning, the weight of the preset composite reward function can be temporarily adjusted. If the conflict is related to equipment overload, the energy efficiency weight Wenergy can be temporarily reduced, and at the same time, the penalty strength of the penalty term Penalty for violating this constraint can be increased. This dynamic weight adjustment aims to emphasize that avoiding specific constraints is more critical than energy conservation currently, guiding the agent to select safe actions away from the constraint boundary in the future, and making the policy converge to a safe and optimal area.

[0056] Optionally, the modification of the weight distribution of the preset composite reward function includes: Obtain data on the conflict between the historical optimized policy and physical constraints, and analyze and construct a conflict type - correction coefficient mapping table based on the data; Specifically, during the training and long-term operation of the reinforcement learning model, record every physical constraint conflict event that occurs or is attempted, including the type of conflict, such as device overclocking, pipeline overpressure, minimum temperature limit, or maximum start-stop times, etc., as well as the state at the time of occurrence, the attempted action, and the relevant environmental background. By performing offline analysis on these historical cases, it is possible to identify which types of conflicts are more frequent or are more likely to occur under certain specific conditions. Based on the severity, frequency, and potential impact on overall objectives, such as energy consumption and comfort, set a set of correction factors for each identified conflict type. These factors clearly define how each item in the preset composite reward function, such as the energy consumption weight Wenergy, the comfort weight Wcomfort, or a specific penalty term Penaltytype, should be adjusted when a conflict of this type occurs. Finally, store these conflict types and their corresponding correction factors, which can be a set of multipliers or addends, in a lookup table or rule base to form a mapping table.

[0057] When a new policy conflict occurs, call the corresponding correction factor in the conflict type - correction factor mapping table to adjust the weight distribution related to comfort and energy consumption in the preset composite reward function.

[0058] Specifically, during the online decision-making or offline training of the reinforcement learning model, when the security layer detects an impending policy conflict, that is, when the agent attempts to execute an action that violates physical constraints, first identify the specific type of this conflict. Then, query the pre-constructed mapping table according to the identified conflict type to obtain the corresponding correction factor. These factors are used to dynamically and temporarily adjust the weights of each item in the preset composite reward function currently used to calculate the reward, such as the weight ratio of the comfort factor Wcomfort and the energy consumption factor Wenergy, as well as the intensity of a possible specific penalty term. For example, if the mapping table indicates that for a "pipeline overpressure" conflict, Wenergy should be multiplied by 0.9, that is, reduce the energy consumption weight, and the penalty term Penaltypressure related to pressure should be multiplied by 1.5, that is, increase the penalty intensity, then these adjusted weights will be used when calculating the reward generated by this attempted violation action. This adjustment is usually temporary and aims to give a more accurate feedback signal to the current violation attempt to guide the agent to correct its behavior.

[0059] Optionally, after feeding back the adjustment result to the building digital twin for real-time physical property simulation, it includes: Collect the actual operating parameters of the energy equipment and calculate the deviation rate from the simulation results of the building digital twin; Specifically, after the device executes the control instruction, collect the actual operating parameters of the device through deployed sensors, such as power meters, flow meters, or thermometers, etc., such as the actual power consumption , the actual flow rate , or the actual temperature of the area . Meanwhile, in the building digital twin, simulation calculations are performed using the same control instructions to obtain corresponding simulation results, such as simulation power , simulation flow rate , or simulation temperature . Then, calculate the deviation between the actual value and the simulation value, such as calculating the relative deviation rate of power and the absolute deviation of temperature . The example formulas are as follows: , , In these formulas refers to the actually measured power, refers to the power predicted by simulation, is the calculated relative deviation rate of power; refers to the actually measured temperature, refers to the temperature predicted by simulation, is the calculated absolute deviation of temperature. By calculating such deviations for key parameters, the current prediction accuracy of the digital twin model can be effectively quantified, or the gap between the physical actual response and the simulation expectation can be reflected.

[0060] When the deviation rate exceeds the preset threshold, trigger the online fine-tuning signal of the reinforcement learning model and perform the fine-tuning of the reinforcement learning; Specifically, a reasonable deviation threshold needs to be set for each key parameter to be monitored. The setting of this threshold should comprehensively consider the measurement accuracy of the sensors used, the inherent errors that may be brought about by the simplification of the digital twin model, and the random fluctuation range under normal operating conditions. For example, for the relative deviation rate related to energy consumption, its threshold can be set to ten percent, while for indicators such as the area temperature, its absolute deviation threshold can be set to 0.5 degrees Celsius. These deviation values will be continuously monitored. When it is detected that the deviation of one or several key parameters continuously exceeds their corresponding preset thresholds, for example, this over-threshold state lasts for a set time window, such as thirty minutes, then it is judged that there is a significant and continuous deviation between the digital twin model and the actual operating state of the physical building, as Figure 4 shown. This deviation may be due to the inaccuracy of the model parameters themselves, the existence of dynamic influencing factors not fully considered by the model, or the change of the performance of physical devices over time. In this case, an online fine-tuning signal will be generated and sent to the reinforcement learning model, indicating that model adjustment is required.

[0061] Re-inject the fine-tuned optimization strategy into the device control instruction set until the deviation rate returns to the tolerance interval.

[0062] Specifically, after receiving an online fine-tuning signal, the reinforcement learning model fine-tunes its network parameters using recently accumulated real interaction data. This adjustment typically employs an optimization algorithm, but with a small learning rate and a limited number of iterations, aiming to absorb real information to correct the model's understanding without disrupting the learned strategy. The newly generated optimized strategy after fine-tuning is parsed into instructions and executed. Continuously monitor the deviation rate: if the deviation returns to the tolerance interval, fine-tuning pauses; if the deviation continues to exceed the standard, further fine-tuning may be carried out or a deeper model calibration may be triggered. This process forms a high-level adaptive closed-loop, continuously correcting the model to adapt to the dynamics of the real world.

[0063] Optionally, the online fine-tuning signal for triggering the reinforcement learning model includes: Switch the dynamic threshold of the deviation rate according to the building operation mode switching signal; Specifically, receive the building operation mode signal, such as working hours, vacant mode, pre-treatment, high-density activities, holidays, etc. Different deviation rate thresholds are preset for each mode. For example, in the "working hours mode", to ensure stability and comfort, the tolerance for energy consumption deviation can be set relatively high to avoid frequent fine-tuning, while in the "vacant mode", to finely calibrate the model during low-disturbance periods, the threshold can be set lower. When the operation mode switches, the preset threshold standard of the new mode is automatically adopted to determine whether the deviation exceeds the standard and decide whether to trigger online fine-tuning.

[0064] During the low-energy consumption period at night, lower the dynamic threshold of the deviation rate; Specifically, it is first necessary to be able to identify that the building has entered the low-energy consumption operation mode at night, which can be achieved by, for example, querying a preset schedule or determining that the current overall energy load level of the building has dropped below a certain predetermined value. Once it is confirmed that the building has entered this mode, the deviation rate thresholds of the key parameters used to trigger online fine-tuning will be automatically lowered to a more stringent level. For example, if the absolute deviation threshold for temperature during working hours is set at plus or minus 0.5 degrees Celsius, then in the low-energy consumption mode at night, this threshold can be reduced to plus or minus 0.3 degrees Celsius. Similarly, the deviation thresholds for other parameters such as energy consumption may also be lowered accordingly. The purpose of adopting this strategy is to make full use of the favorable conditions of relatively stable internal loads in the building at night and fewer disturbing factors from the external environment to more precisely calibrate the reinforcement learning model and its dependent digital twin model. By capturing and correcting more subtle prediction deviations, the prediction accuracy of the model under low-load conditions can be effectively improved, which also lays a more solid foundation for optimized control under complex working conditions during the day.

[0065] During the peak electricity consumption period, expand the dynamic threshold of the deviation rate.

[0066] Specifically, it is necessary to be able to identify whether the current period belongs to the peak power consumption period of the power grid. For example, it can be judged based on a preset time period or a peak time electricity price signal received from the power company. During this peak period, in order to prioritize the stable operation of the power, avoid triggering a high demand charge mechanism, and reduce the disturbances that frequent control adjustments may cause to energy equipment and indoor users, the deviation rate thresholds for triggering online fine-tuning will be adjusted to a relatively loose level. For example, if the energy consumption deviation threshold during normal working hours is set at 15%, it may be temporarily increased to 25% during the peak power consumption period. This means that unless there is a very significant deviation between the prediction of the digital twin model and the physical reality, the triggering of online fine-tuning will tend to be suppressed, allowing the currently executed optimization strategy to operate stably. At this time, the priority is to ensure "no errors" rather than pursuing the ultimate "optimal" performance.

[0067] Optionally, the parsing of the multi-objective optimization strategy into a device control instruction set includes: Identifying the communication interface types of different device protocols and constructing an instruction format conversion template library; Specifically, first, it is necessary to establish a device asset information database to record the detailed information of all controllable energy equipment in the building, such as chillers, pumps, fans, valves, variable air volume (VAV) terminals, and intelligent lighting fixtures. The content of the information database should include the model, manufacturer, and supported communication protocols of each device, such as common building automation protocols like BACnet, Modbus, KNX, or DALI. It also needs to include the network address information of the device, such as IP address, device ID, or port number, as well as the specific point information required for control and status monitoring, such as the object identifier of BACnet or the register address of Modbus. Based on this information database, for each supported communication protocol and common device control actions, such as reading and writing binary or analog values, changing setpoints, or switching operation modes, standardized instruction format templates are created in advance. For example, a template for writing an analog value for the BACnet protocol may contain placeholders for fields such as protocol data unit (PDU) type, service type, object identifier, attribute identifier, priority, and the value to be written. This pre-constructed instruction format conversion template library is the basis for achieving unified control and interoperability of various heterogeneous devices in the building.

[0068] Matching the control protocol characteristics of the target device in the instruction format conversion template library to generate an adapted binary instruction block; Specifically, when the reinforcement learning model outputs a high-level optimization policy, such as the instruction "Set the VAV box damper opening in area Z01 to 60%", the parsing module first identifies the controlled target device, i.e., the VAV box in area Z01, its communication protocol, assumed to be BACnet / IP, and the point to be written, such as the analog output object AO-1 corresponding to the damper opening. Then, a template applicable to "writing analog value via BACnet / IP" is selected from the template library. Information such as the specific device address, object identifier such as AO-1, attribute identifier such as PresentValue, priority such as 8, and the target value 60.0 to be written is filled into the corresponding fields of the template. Finally, according to the selected protocol specification such as BACnet, the filled template is encoded into a binary data packet that conforms to the protocol standard, i.e., the application protocol data unit APDU, for network transmission preparation.

[0069] Dynamically queue-sort the sending order of instruction blocks according to the device response latency requirements.

[0070] Specifically, considering that not all energy devices respond to control instructions at the same speed, and some complex control logics may have strict timing requirements, such as the need to open the water valve first and then start the associated fan. An information library should be maintained, which contains typical response latency information of different device types or specific device instances, or set priorities for them. When there are multiple control instructions to be sent to different devices simultaneously or within a very short time, it is not necessarily simply sent in the order of instruction generation. Instead, these generated binary instruction blocks are first put into a queue to be sent. Then, the queue manager dynamically sorts the instruction blocks in the queue according to a series of factors. These factors include, for example, the expected response latency of the target devices of each instruction. The sorting strategy may give priority to sending to devices with fast responses to observe the effects as soon as possible, or give priority to sending to devices with slow responses to ensure their early start, depending on the control objective at that time, the urgency of the instruction itself, such as safety-related instructions having the highest priority, and the possible control logic dependencies between instructions. Finally, the binary instruction blocks are sent to the target devices for execution one by one through the corresponding network interfaces in the order of dynamic sorting.

[0071] Exemplarily, assume that at a certain moment, the optimization strategy determines that it is necessary to simultaneously adjust the chilled water outlet temperature setpoint of the central chiller and the damper opening degrees of the VAV boxes distributed on multiple floors. The former is controlled, for example, through the Modbus TCP protocol, and its response is usually slow, so an instruction needs to be sent in advance; the latter is controlled, for example, through the BACnet / IP protocol, and the response is relatively fast. Based on this, the parsing module generates a Modbus instruction block and ten BACnet instruction blocks. Subsequently, according to the preset sorting rules, for example, the rules stipulate that the instructions for controlling the central system-level devices should be sent first to guide the overall operation trend, the instruction queue manager places that Modbus instruction block at the head of the sending queue, and arranges the ten BACnet instruction blocks behind it. Finally, the instructions are sent in this sorted order. This can ensure that the central chiller starts to adjust its outlet temperature first, and then the VAV boxes on each floor can accurately adjust the air volume according to the changed water temperature and the real-time demands of their respective areas, thus realizing a more stable and coordinated control state transition process.

[0072] Based on the same inventive concept, the present invention also provides a building energy consumption dynamic optimization system based on BIM and reinforcement learning, as Figure 5 shown, the system includes: A digital twin construction module, used to establish a BIM containing the physical attribute parameters of building components, and fuse real-time monitoring data to generate a building digital twin reflecting the evolution of dynamic thermal properties; A state space construction module, used to extract the spatial topological relationship and component physical property parameters based on the building digital twin, and construct a multi-dimensional state space of a preset reinforcement learning model; A policy generation module, used to embed physical constraint conditions from the building digital twin into the reinforcement learning model, and train the reinforcement learning model based on the physical constraint conditions and a preset composite reward function to generate a multi-objective optimization policy for energy equipment; A control and feedback module, used to parse the multi-objective optimization policy into a device control instruction set, control the energy equipment to execute parameter adjustment, and feedback the adjustment result of the energy equipment to the building digital twin for real-time physical property simulation.

[0073] It should be noted that the functional division and information interaction among the above-mentioned various modules are logical. Physically, they can be integrated on the same software platform or distributedly deployed. The connections between them represent data flow and control flow, aiming to jointly achieve the building energy consumption dynamic optimization goal of the present invention. The above are only exemplary embodiments of the present invention, and the protection scope of the present invention cannot be limited thereby.

Claims

1. A dynamic optimization method for building energy consumption based on BIM and reinforcement learning, characterized in that: The method comprises: Establish BIM that includes physical property parameters of building components and integrate real-time monitoring data to generate a building digital twin that reflects the evolution of dynamic thermal properties; Extracting spatial topological relationships and component physical property parameters based on the building digital twin, and constructing a multidimensional state space of a preset reinforcement learning model; Embedding physical constraints derived from the building digital twin in the reinforcement learning model, and training the reinforcement learning model based on the physical constraints and a preset composite reward function to generate a multi-objective optimization strategy for energy equipment; The multi-objective optimization strategy is parsed into a device control instruction set to control the energy equipment to perform parameter adjustment, and the adjustment results of the energy equipment are fed back to the building digital twin for real-time physical property simulation.

2. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 1 is characterized in that: The generation of a building digital twin reflecting the evolution of dynamic thermal properties includes: Obtain the aging coefficient data set of building materials and establish a dynamic heat conduction model associated with seasonal changes; Identify the solar radiation exposure factor of outdoor facade components and generate light reflectivity attenuation curves based on historical dust monitoring data; The thermal property parameters of the components in the BIM are updated according to the dynamic thermal conductivity model and the light reflectivity attenuation curve.

3. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 1 is characterized in that: The multidimensional state space of reinforcement learning is constructed as follows: Dividing the building structure topology network from the building digital twin and extracting the energy conversion efficiency matrix of HVAC equipment; Collect air pressure gradient data and historical equipment operation logs to generate a probability distribution model for environmental load disturbance terms; The topological network, the efficiency matrix and the probability distribution model are integrated into a discrete parameter set.

4. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 3 is characterized in that: The method of collecting air pressure gradient data and historical equipment operation logs to generate a probability distribution model for environmental load disturbance terms includes: Collect pressure gradient data and historical equipment operation logs; Analyze the dynamic characteristics of the building's internal environment according to the air pressure gradient data and perform microenvironment zoning; Calculate the energy consumption fluctuation threshold of the equipment in each partition based on the historical equipment operation log; When the real-time energy consumption data exceeds the device energy consumption fluctuation threshold, the discrete parameter set is triggered.

5. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 1 is characterized in that: The physical constraints embedded in the building digital twin in the reinforcement learning model include: Based on the stress simulation results of the BIM, a set of constraint rules for prohibiting overload operation of equipment is generated; During the reinforcement learning training process, generating and evaluating action sequences representing candidate strategies, and intercepting the candidate strategies that violate the constraint rule set; When a conflict is detected, the weight distribution of the preset composite reward function is modified to guide the policy to converge.

6. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 5 is characterized in that: The modified preset compound reward function weight allocation includes: Obtain data on conflicts between historical optimization strategies and physical constraints, and analyze and construct a conflict type-correction coefficient mapping table based on the data; When a new policy conflict occurs, the corresponding correction coefficient in the conflict type-correction coefficient mapping table is called to adjust the weight distribution related to comfort and energy consumption in the preset composite reward function.

7. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 1 is characterized in that: Feeding back the adjustment result to the building digital twin for real-time physical property simulation includes: Collecting actual operating parameters of the energy equipment and calculating the deviation rate from the simulation results of the building digital twin; When the deviation rate exceeds a preset threshold, an online fine-tuning signal of the reinforcement learning model is triggered, and the reinforcement learning is fine-tuned; The fine-tuned optimization strategy is reinjected into the device control instruction set until the deviation rate returns to the tolerance interval.

8. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 7 is characterized in that: The online fine-tuning signal triggering the reinforcement learning model includes: According to the building operation mode switching signal, the dynamic threshold of the switching deviation rate is switched; During the nighttime low energy consumption period, lowering the dynamic threshold of the deviation rate; During peak electricity consumption periods, the dynamic threshold of the deviation rate is expanded.

9. The method for dynamic optimization of building energy consumption based on BIM and reinforcement learning according to claim 1 is characterized in that: The parsing of the multi-objective optimization strategy into a device control instruction set comprises: Identify the communication interface types of different device protocols and build a template library for instruction format conversion; Matching the control protocol characteristics of the target device in the instruction format conversion template library to generate an adapted binary instruction block; According to the device response latency requirements, the sending order of the instruction blocks is dynamically queued.

10. A building energy consumption dynamic optimization system based on BIM and reinforcement learning, applied to execute a building energy consumption dynamic optimization method based on BIM and reinforcement learning as claimed in any one of claims 1 to 9, characterized in that: The system comprises: The digital twin construction module is used to build BIM containing the physical property parameters of building components and integrate real-time monitoring data to generate a building digital twin that reflects the evolution of dynamic thermal properties; A state space construction module, used to extract spatial topological relationships and component physical property parameters based on the building digital twin, and construct a multidimensional state space of a preset reinforcement learning model; A strategy generation module, used to embed physical constraints derived from the building digital twin in the reinforcement learning model, and train the reinforcement learning model based on the physical constraints and a preset composite reward function to generate a multi-objective optimization strategy for energy equipment; The control and feedback module is used to parse the multi-objective optimization strategy into a device control instruction set, control the energy equipment to perform parameter adjustment, and feed back the adjustment results of the energy equipment to the building digital twin for real-time physical property simulation.

Citation Information

Patent Citations

  • Air conditioner energy-saving control method and system based on reinforcement learning and digital twinborn model

    CN118031385A

  • Building thermal performance intelligent energy-saving analysis method based on BIM

    CN118981816A

  • Building environment dynamic regulation and control method and system based on Internet of Things driving

    CN119356083A

  • Intelligent supervision method for digital power system based on cloud computing

    CN119727137A

  • BIM-based building energy consumption simulation analysis optimization method and system

    CN119849297A

Cited By

  • Production decision adaptive optimization method and system

    CN120406165A

  • Project quality visual supervision system and method applied to municipal construction

    CN120410338A

  • Efficient machine room multi-subsystem cooperative control method based on intelligent simulation

    CN120428596A

  • Low-carbon building monitoring system and method based on digital twinning

    CN120542676A

  • Wallboard digital optimization method for arc-shaped clear water wall surface

    CN120562209A