Intelligent building envelope dynamic temperature regulating system based on reinforcement learning

By using a reinforcement learning-based intelligent building envelope dynamic temperature control system, which combines the MADDPG algorithm and digital twin model with variable thermal conductivity coating and PCM ventilation valve regulation, the problem of traditional building envelopes being unable to dynamically adapt to climate change has been solved, achieving efficient energy consumption optimization and improved thermal comfort.

CN120909365BActive Publication Date: 2026-07-21KUNMING UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
KUNMING UNIV OF SCI & TECH
Filing Date
2025-07-29
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

The thermal inertia design of traditional building envelopes cannot dynamically adapt to climate change and indoor load changes, resulting in high air conditioning energy consumption. Existing intelligent control systems have failed to fully optimize the thermal performance of the building envelope.

Method used

A dynamic temperature control system for intelligent building envelope based on reinforcement learning is adopted, which includes a data acquisition layer, a multi-agent decision-making layer, an envelope execution layer, and an energy consumption optimization layer. The system uses the MADDPG algorithm and a digital twin model for real-time temperature control, combined with variable thermal conductivity coating and PCM ventilation valve regulation.

Benefits of technology

It achieves dynamic optimization of building thermal resistance, reduces air conditioning energy consumption by 40%, improves indoor thermal comfort, narrows the temperature fluctuation range, adapts to climate change, has self-learning capabilities, and is suitable for super high-rise buildings, data centers, and medical clean spaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909365B_ABST
    Figure CN120909365B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of dynamic optimization of building thermal resistance, and discloses an intelligent building envelope dynamic temperature regulation system based on reinforcement learning, which comprises a data acquisition layer, a multi-agent decision layer, an envelope execution layer and an energy consumption optimization layer; the data acquisition layer is used for detecting current regional meteorological data, envelope state data and the temperature and humidity of an indoor environment. The system realizes dynamic optimization and accurate regulation and control of the thermal performance of a building through innovative material-algorithm collaborative design, dynamically and steplessly adjusts the thermal resistance of the envelope for the first time, solves the core pain point that a static design is difficult to adapt to dynamic changes in the climate, has self-learning ability and physical constraint guarantee, and has wide application prospects in the fields of super high-rise buildings, data centers and medical clean spaces.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic optimization technology for building thermal resistance, specifically to a dynamic temperature control system for intelligent building envelope based on reinforcement learning. Background Technology

[0002] The global energy consumption situation is severe. According to statistics from the International Energy Agency (IEA), building operating energy consumption accounts for 36% of total end-use energy consumption globally (2022 data), with heating, ventilation, and air conditioning (HVAC) accounting for as much as 50-60%. China's urbanization rate has exceeded 65% (2023), with a stock of glass curtain wall buildings exceeding 1 billion square meters and an annual increase in air conditioning load of 50 million kW, equivalent to the annual power generation of five Three Gorges Dams. In 2023, many parts of the world experienced record-breaking high temperatures (such as 43°C in Beijing and 47°C in Spain). Traditional building thermal inertia designs are insufficient to cope with sustained heat waves, requiring real-time adaptive systems for adjustment. Moreover, the widening peak-valley difference in electricity prices necessitates buildings with demand response capabilities, proactively reducing load during peak electricity price periods. The volatility of renewable energy sources (such as the diurnal intermittency of photovoltaic power generation) requires building envelopes to have energy buffering functions. The improved computing power of edge computing hardware makes real-time physics simulation and AI decision-making possible.

[0003] The limitations of traditional energy-saving technologies are that existing building envelopes (such as double-layer Low-E glass and rock wool insulation layers) are designed based on the most unfavorable operating conditions. Static thermal design cannot adapt to dynamic climate conditions and changes in indoor load, resulting in persistently high air conditioning energy consumption. This lack of dynamic adaptation leads to a rapid increase in air conditioning cooling energy consumption when there are diurnal temperature differences and seasonal changes. Traditional PCM applications rely on natural heat storage and release, and cannot actively regulate the rhythm. For example, insufficient heat storage during the day leads to overloading of the air conditioning system at night. Existing technologies (CN119684975A) provide a formula and production process to improve the thermal conductivity of phase change materials and refrigerants, but do not solve the problem of dynamic regulation. Furthermore, the intelligent control system is often limited. For example, CN119468418A discloses a data-driven energy-saving control method and system for air conditioning equipment groups, which relates to the field of air conditioning equipment group control technology, but neglects the optimization of the thermal performance of the building envelope itself, resulting in the energy-saving potential not being fully explored. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a dynamic temperature control system for intelligent building envelopes based on reinforcement learning, which has advantages such as dynamic optimization of building thermal resistance and solves the aforementioned technical problems.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a dynamic temperature control system for intelligent building envelope based on reinforcement learning, comprising a data acquisition layer, a multi-agent decision-making layer, an envelope execution layer, and an energy consumption optimization layer;

[0006] The data acquisition layer is used to detect current regional meteorological data, building envelope status data, and indoor temperature and humidity.

[0007] The multi-agent decision layer includes an edge computing unit and a multi-agent reinforcement learning engine. The edge computing unit is used to process all the data collected in the data acquisition layer and perform preprocessing. The multi-agent reinforcement learning engine outputs commands to the enclosure structure execution layer based on the data preprocessed by the edge computing unit and through the MADDPG algorithm.

[0008] The enclosure structure execution layer adjusts the temperature of the enclosure structure based on the command output of the intelligent agent decision layer;

[0009] The energy consumption optimization layer includes a digital twin model and an incremental learning module. The digital twin model is simulated in real time using EnergyPlus, and the incremental learning module is used to provide feedback to the multi-agent decision-making layer.

[0010] The MADDPG algorithm in the aforementioned multi-agent reinforcement learning engine coordinates the actions of the agents through an Actor network and a Critic network. The specific expression for the local observation state of the i-th agent is as follows:

[0011]

[0012] Where t represents the discrete time period of policy interaction and decision-making in the reinforcement learning system, and is the core index for reinforcement learning training in the MADDPG algorithm, used for: state transition ,action , k represents the consecutive iteration rounds or optimization steps in physical simulation, parameter calibration, and predictive optimization, and the optimization step size in EnergyPlus parameter calibration. The evolution of all k is "offline" and "internal model iteration," and does not participate in the actual execution strategy inference. i refers to the i-th building facade area. and These represent the indoor temperature and humidity corresponding to the i-th building facade area, respectively. , Let represent the outdoor temperature and humidity corresponding to the i-th building facade area. Let be the solar radiation intensity received by the i-th building facade area. Let be the real-time electricity price at time t. For the heat load demand of the i-th building facade area, The time feature corresponding to time t is obtained by concatenating the local observation states of 1 to N agents, as shown in the following expression:

[0013]

[0014] Where: s represents the concatenated set of all agents, s1 is the state vector (15-dimensional) of the first agent, and so on, and the action expression for each agent is as follows:

[0015]

[0016] in: This represents the set of concatenated action vectors. ~ These represent the action vectors corresponding to agents 1 through N, respectively. The driving voltage for the VO2 coating controlled by the i-th intelligent agent. Let represent the opening degree of the PCM ventilation valve controlled by the i-th intelligent agent.

[0017] As a preferred technical solution of the present invention, the specific means of controlling the execution layer of the enclosure structure include variable thermal conductive coating drive and PCM ventilation valve regulation;

[0018] The variable thermal conductivity coating drive is used to control the voltage across the VO2 film to adjust the thermal conductivity of the VO2 film.

[0019] The PCM ventilation valve is controlled by controlling the stepper motor to drive the louvers, thereby controlling the rate of heat storage and release.

[0020] As a preferred embodiment of the present invention, the digital twin model is simulated in real time using EnergyPlus, and the building thermodynamic model is updated every 15 minutes. The specific expression is as follows:

[0021]

[0022] in: This includes parameters related to the thermal conductivity of the wall, the U-value of the window, the thickness of the insulation layer, the heat capacity of the material, the specific heat capacity of each component, the thermal diffusivity of the wall, the internal surface thermal resistance, the external surface heat transfer coefficient, the heat capacity and enthalpy of the phase change material (PCM) layer, the thermal conductivity of the VO2 thin film, and other parameters related to the thermal performance of the material. argmin represents the parameter corresponding to the minimum value. The measured indoor temperature, heat flux on the wall or curtain wall surface, total energy consumption of the area, and electrical input power (power consumption of VO2 drive circuit and PCM valve motor) at time k are given. This is the simulation output of the EnergyPlus model under the current input conditions, executed every 15 minutes. The simulation output generated by EnergyPlus for the digital twin model is as follows:

[0023]

[0024] in: Let k be the measured indoor temperature at time k. Heat flux of wall or curtain wall surface Regional total energy consumption and This refers to the electrical input power (power consumption of the VO2 drive circuit and PCM valve motor).

[0025]

[0026] Where: θ is a thermal parameter. For input conditions, Let be the forward simulation function of the EnergyPlus model. The digital twin model's prediction of load and temperature within the next Δt minutes at time k is as follows:

[0027]

[0028] in: For predicted future temperature and heat load, The input conditions are defined at time k, and the output of the digital twin model is fed back to the Critic network.

[0029] Compared with existing technologies, this invention provides a dynamic temperature control system for intelligent building envelopes based on reinforcement learning, which has the following beneficial effects:

[0030] This invention achieves dynamic optimization and precise control of building thermal performance through innovative material-algorithm collaborative design. It is the first to realize dynamic stepless adjustment of the thermal resistance of the building envelope, solving the core pain point that static design is difficult to adapt to dynamic climate changes. It has both self-learning ability and physical constraint guarantee, and has broad application prospects in fields such as super high-rise buildings, data centers and medical clean spaces. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the structure of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Please see Figure 1 A dynamic temperature control system for intelligent building envelope based on reinforcement learning, comprising a data acquisition layer, a multi-agent decision-making layer, an envelope execution layer, and an energy consumption optimization layer;

[0034] The data acquisition layer consists of meteorological data acquisition, building envelope condition monitoring, and indoor environment sensing. Meteorological data acquisition can combine local weather forecasts and outdoor rooftop micro weather stations to collect data such as temperature, radiation, and wind speed around the building; building envelope condition monitoring is conducted through distributed fiber optic temperature measurement (DTS), laid longitudinally along the exterior walls, with one measuring point per meter; indoor environment sensing is achieved through a wireless temperature and humidity sensor network (Zigbee protocol), deployed in a 10m×10m grid.

[0035] The multi-agent decision layer includes an edge computing unit and a multi-agent reinforcement learning engine. The edge computing unit is used to process all the data collected in the data acquisition layer and perform preprocessing. The multi-agent reinforcement learning engine outputs commands to the enclosure structure execution layer based on the data preprocessed by the edge computing unit and through the MADDPG algorithm.

[0036] Compared to previous patents, the preferred technical solution of this invention has more advantages. Using a multi-agent framework, it can simultaneously optimize the heat distribution and collaborative relationships between different building facades, overcoming the limitations of isolated optimization of a single fan. Employing state input dimensions and multimodal perception (15 dimensions), richer environmental and economic inputs enable Actors to make precise temperature control decisions at a higher dimension. Compared to the previous fixed K1 and K2 weight values, it introduces weight parameters α, β, and γ to dynamically adjust the operating target. Furthermore, it uses digital twins and online simulation to enhance Critic, EnergyPlus for real-time calibration and short-term prediction, and joint state-action and simulation prediction. And least squares fitted thermal parameters every 15 minutes This approach is more reasonable than previous empirical formulas, state transitions without actual measurements, and unpowered model updates. Algorithm-wise, it employs the dual-criteria + target network soft update method from MADDPG. This improves robustness and convergence performance in multidimensional complex environments. Soft updates combined with simulation enhancement ensure smooth policy evolution. The above improvements comprehensively surpass the AC network control method based on a single agent, a single temperature input and a fixed policy in the comparative patent. This fully demonstrates the creative breakthrough of this application in algorithm architecture, system integration and application scenarios.

[0037] The multi-agent decision-making layer is divided into an edge computing unit and a multi-agent reinforcement learning engine. The edge computing unit receives data from the data acquisition layer and performs data preprocessing. It can use NVIDIA Jetson AGX Orin with 32GB of memory and supports TensorRT to accelerate inference. The multi-agent reinforcement learning engine uses the MADDPG (Multi-Agent Deep Deterministic Policy Gradient) algorithm, an algorithm for multi-agent reinforcement learning (MARL), to take the joint state and actions of all agents as input, output Q-values ​​for evaluation, and finally make corresponding decisions and output commands to the execution layer.

[0038] A detailed description of MADDPG in a dynamic temperature control system for intelligent building envelopes introduces the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) framework. This algorithm, under the centralized training, distributed execution (CTDE) paradigm, coordinates the continuous control problem across multiple agents by providing each agent with an independent Actor network and a centralized Critic network. The local observation states of each agent i are as follows:

[0039]

[0040] Where t represents the discrete time period of policy interaction and decision-making in the reinforcement learning system, and is the core index for reinforcement learning training in the MADDPG algorithm, used for: state transition ,action , k represents the consecutive iteration rounds or optimization steps in physical simulation, parameter calibration, and predictive optimization, and the optimization step size in EnergyPlus parameter calibration. The evolution of all k is "offline" and "internal model iteration," and does not participate in the actual execution strategy inference. i refers to the i-th building facade area. and These represent the indoor temperature and humidity corresponding to the i-th building facade area, respectively. , Let represent the outdoor temperature and humidity corresponding to the i-th building facade area. Let be the solar radiation intensity received by the i-th building facade area. Let be the real-time electricity price at time t. For the heat load demand of the i-th building facade area, For time t, the local observation states of 1 to N agents are concatenated, and the specific expression is as follows: The overall system state is the concatenation of the observations of all agents:

[0041]

[0042] Where s represents the set of all agents concatenated together. ~ Let represent the state vectors of agents 1 to N, and so on. The action expression for each agent is as follows:

[0043]

[0044] in: ~ These represent the action vectors corresponding to agents 1 through N, respectively. The driving voltage for the VO2 coating controlled by the i-th intelligent agent. The opening degree of the PCM ventilation valve controlled by the i-th intelligent agent;

[0045] Critic Network Update:

[0046] Critic Network The joint state and joint action are used as inputs to estimate the state-action value function of the i-th agent. Critic parameters. Update by minimizing the mean squared Bell residual loss function:

[0047]

[0048] Target value Defined as

[0049]

[0050] in As a discount factor, , are the target network parameters for Critic and Actor, respectively, and D is the empirical replay pool;

[0051] Actor Network Update:

[0052] Actor Network Generate an action for the i-th Agent given local observations. Its parameters are... Optimization through policy gradient ascent:

[0053]

[0054] This gradient utilizes global information provided by the Critic network to guide the Actor in adjusting its output actions to improve overall returns.

[0055] Target network soft update:

[0056] Export the trained MADDPG model to TensorFlow Lite format. For stable training, target network parameters. and A soft update strategy is adopted, fixing the first L−1 layers of the Actor and Critic, and only updating the parameter ϕ of the Lth layer online:

[0057]

[0058] Reward function design:

[0059] Instant reward for the i-th agent Combining factors such as energy efficiency, comfort, and cost:

[0060]

[0061] in: Energy consumption under intelligent agent control For the comfort temperature set point, α, β, and γ are the temperature fluctuation rates, and α, β, and γ are weighting coefficients.

[0062] The building envelope execution layer regulates the temperature of the building envelope based on the command output of the agent decision layer;

[0063] The building envelope execution layer is divided into a variable thermal conductivity coating drive and a PCM ventilation valve control. The variable thermal conductivity coating drive can adjust the thermal conductivity by controlling the voltage. Specifically, when a voltage is applied across the VO2 film, Joule heating is generated, causing the film temperature to rise. When the temperature reaches the phase transition temperature, the VO2 film changes from an insulating state to a metallic state. By controlling the magnitude and duration of the voltage, the temperature of the VO2 film can be precisely controlled. The PCM ventilation valve control controls the rate of heat storage and release by controlling the stepper motor to drive the louvers. The thermal resistance distribution is dynamically monitored, and air conditioning energy consumption is monitored simultaneously.

[0064] The energy optimization layer consists of a digital twin model and an incremental learning module. The digital twin model can be simulated in real time using EnergyPlus, updating the building thermodynamics model every 15 minutes to accurately predict current and future heat loads and energy-saving potential. The digital twin is defined by EnergyPlus as:

[0065]

[0066] The simulation output of the digital twin generated by EnergyPlus is as follows:

[0067]

[0068] Where: θ represents thermal parameters (wall thermal conductivity, window U-value, etc.). Given the input conditions (weather, load, control actions), the calibration module constructs an optimization problem aimed at minimizing the error between the simulation output and the measured observations, and solves for the building thermal model parameters θ to ensure that the digital twin model remains consistent with the real building. Its mathematical expression is as follows:

[0069]

[0070] in: This includes parameters such as wall thermal conductivity, window U-value, insulation layer thickness, material heat capacity, specific heat capacity of each component, wall thermal diffusivity, internal surface thermal resistance, external surface heat transfer coefficient, heat capacity and enthalpy of phase change material (PCM) layer, VO2 thin film thermal conductivity, and other parameters related to the thermal performance of the material. The measured indoor temperature and heat flux at time k are the actual values, which include physical quantities such as temperature, humidity, heat flux density, and energy consumption, and are collected by field sensors; the simulation output is calculated by the EnergyPlus model.

[0071] The energy consumption optimization layer includes a digital twin model and an incremental learning module. The digital twin model is simulated in real time using EnergyPlus, and the incremental learning module is used to provide feedback to the multi-agent decision-making layer.

[0072] The digital twin predicts the load and temperature at time k for the next Δt minutes (typically 15 minutes):

[0073]

[0074] This prediction is incorporated into the calculation of the target value of MADDPG Critic to estimate the Q value of the next state;

[0075] The real-time simulation purpose of the digital twin model is to combine EnergyPlus's physical simulation (simulating the building's structure, materials, HVAC system, etc.) with real-time sensor data (temperature, humidity, radiation, wind speed, etc.) to promptly correct key parameters in the model, such as thermal conductivity and equipment efficiency, maintaining consistency between the virtual model and the real building's behavior. Based on the current state and weather forecast, the digital twin can predict heat load and temperature changes over the next 15–60 minutes, providing RL agents with short-term strategy evaluation and "virtual trial operation" capabilities, reducing on-site trial and error costs. Strategy simulation verification: Before the execution layer formally issues control, the system can first simulate multiple action combinations in parallel in the digital twin (such as different VO2 coating voltages and PCM valve openings) to estimate their energy consumption and temperature response, assisting the Critic network in MADDPG to more accurately evaluate the Q value.

[0076] The incremental learning module uses TensorFlowLite for fine-tuning, updating the MARL model weights weekly or monthly with new data to cope with environmental changes and equipment aging.

[0077] The fine-tuning process and purpose of the incremental learning module is to export the trained MARL model in TensorFlow Lite format and deploy it in edge computing units to support low-latency policy inference and updates. The system summarizes new state-action-reward triples generated during actual operation weekly or every few days. Utilizing TensorFlow Lite's incremental training capabilities (freezing some network layers and only opening the last one or two layers for fine-tuning), the model is trained for several epochs to adapt to new environments or load changes. The fine-tuned new weights replace the Actor and Critic target network parameters in the edge units, ensuring that subsequent inference maintains high performance in the latest environment.

[0078] The incremental learning module uses TensorFlowLite for fine-tuning. The building's thermal performance changes with the seasons (such as insulation aging and changes in furniture layout). It incorporates actual operational feedback into the model to reduce the difference between digital twin or simulation strategies and the actual control effect. Based on the original strategy, it uses the latest data to mine new energy consumption-comfort trade-offs to improve the system's robustness and adaptability.

[0079] Initialization: Construct the Actor, Critic, and corresponding target network for each Agent; initialize the experience replay pool. Centralized training loop: All Agents interact with the environment and collect data. Add to the replay pool; periodically sample and update Critic and Actor from the pool; perform soft updates. Distributed execution: After training convergence, each Agent only calls its local Actor; in actual operation, it is based on... Independently generated actions enable online control with low communication requirements. This MADDPG solution combines a central Critic with local Actors to achieve coordinated control of VO2 coatings and PCM valves in the building envelope, balancing energy saving and comfort, and is suitable for large-scale multi-agent scenarios.

[0080] Overall process

[0081] State acquisition: Obtaining the actual state value (Including temperature, humidity, radiation, electricity price, etc.);

[0082] Digital twin calibration: ;

[0083] Bionic prediction: ;

[0084] Strategy inference: ;

[0085] Action execution: Issue VO2 coating drive voltage and PCM ventilation valve opening ;

[0086] Feedback sampling: ;

[0087] Experience storage: ;

[0088] Small batch updates: Critic and Actor are updated according to formulas;

[0089] Incremental fine-tuning (weekly / every few days): ;

[0090] The target network software is updated, and the closed-loop combination of digital twin and MADDPG is used to ensure that the temperature regulation strategy remains optimal in the real building environment through continuous simulation calibration and online fine-tuning.

[0091] Materials and hardware layout

[0092] (1) Smart enclosure structure hardware, material selection for phase change material (PCM) layer: octadecane (C 18 H 38 (Melting point 28℃), phase change enthalpy ≥180 kJ / kg, suitable for the human body's comfortable temperature range, the modifier is 5% expanded graphite added, the thermal conductivity is increased from 0.2 to 5 W / m·K, accelerating the phase change response;

[0093] Encapsulation process: Microencapsulation is achieved by preparing a urea-formaldehyde resin shell (thickness 2-5 μm) using interfacial polymerization.

[0094] Construction method: Mix with silicate thermal insulation mortar at a volume ratio of 1:4, and mechanically spray to a thickness of 20mm;

[0095] (2) Variable thermal conductivity coating, VO2 thin film preparation: magnetron sputtering method with a thickness of 300±50nm, phase transition temperature of 68℃ (which can be reduced to 40℃ by doping with tungsten), thermal conductivity switching ratio of 3:1, and visible light transmittance >40% (meeting the light-gathering requirements);

[0096] (3) Drive circuit, low power consumption design, 12V DC power supply, maximum current 0.1A (power consumption per square meter ≤1.2W).

[0097] Multi-Agent Reinforcement Learning (MARL) algorithm, agent partitioning strategy: divide the building facade into one agent for every 10 floors (e.g., a 200m high building is divided into 20 agents), and each agent only acquires sensor data for its own area (reducing the dimensionality of the state space).

[0098] The algorithm implementation process involves creating an Actor and Critic network (including a target network) for each agent during initialization, and initializing the experience replay pool.

[0099] The training loop involves each agent selecting the corresponding operation based on the current policy.

[0100] After the operation is executed, the environment returns the reward and the next state to all agents.

[0101] Global transfer experience (joint states, actions, rewards, etc.) is stored in the replay pool.

[0102] Sample a batch of data from the replay pool and update the Critic and Actor. Critic update: minimize the TD error; Actor update: maximize the Q-value evaluated by the Critic through gradient ascent.

[0103] Ultimately, a corresponding decision is made, and commands are output to the execution layer.

[0104] The execution layer is the physical response terminal of the system, responsible for converting the AI ​​control commands generated by the decision layer into material property changes and mechanical actions. Specifically, it includes the following two major modules: variable thermal conductivity coating (VO2 thin film) drive, where VO2 undergoes an insulator-metal phase transition at 68℃, and the thermal conductivity jumps from 0.5 W / m·K (insulating state) to 3.0 W / m·K (metallic state). The phase transition can be artificially triggered by applying voltage (without natural heating).

[0105] The control is divided into a low thermal conductivity mode (0.5 W / m·K), which reflects solar radiation and reduces heat transfer (during summer days), and a high thermal conductivity mode (3.0 W / m·K), which accelerates heat dissipation (nighttime in winter or emergency heat dissipation in data centers).

[0106] Execution steps

[0107] Receiving instructions: The edge computing node sends voltage instructions via the PLC;

[0108] Voltage conversion: The drive circuit converts digital signals into analog voltages (0-10VDC, accuracy ±0.1V).

[0109] Thin film response: Voltage is applied to the VO2 thin film electrode, triggering carrier migration and changing the lattice structure (response time ≤100ms).

[0110] Status feedback: Real-time monitoring of thin-film resistance (normal range: insulating state 10) 6 Ω → metallic state 10²Ω), ensuring the phase transition is complete;

[0111] At the same time, the PCM ventilation valve is regulated and receives the opening command, which is converted into the number of stepper motor pulses (each 1% opening = 32 pulses).

[0112] Motor driven, ULN driver board controls motor phase (four-phase eight-step mode), speed 15RPM;

[0113] Position calibration, limit switch (Omron SS-5GL) reset to zero point to avoid cumulative error;

[0114] Status feedback: Hall sensor (A) detects blade angle and transmits the actual opening in real time (accuracy ±1%).

[0115] The VO2-coated electrode wiring uses silver paste printed electrodes (0.5mm line width, 10mm spacing) to avoid obstructing light transmission; the electrode ends are connected to a waterproof junction box (IP67 rating).

[0116] PCM ventilation valve deployment: 4 ventilation valves are installed per square meter of curtain wall (500mm spacing); Duct design: installed at a 30° angle to utilize gravity drainage and prevent dust accumulation;

[0117] Control signal transmission, communication protocol: VO2 driver: Modbus RTUover RS-485 (address code allocation, baud rate); ventilation valve control: CAN bus (message ID distinguishes regions);

[0118] Fault tolerance mechanism: heartbeat detection (every 5 seconds), triggering safety mode after 3 timeouts (VO2=1.5W / m·K, ventilation valve=50%).

[0119] Maintenance and troubleshooting: Wipe the VO2 coating surface with ethanol quarterly to maintain light transmittance, and check the electrode resistance annually (normal value 10²-10). 6 Ω); PCM ventilation valves should have their blades cleaned of dust monthly (using compressed air purging); gear sets should be lubricated every six months (using food-grade silicone grease);

[0120] The execution layer transforms AI decisions into precise responses in the physical world through the physical property regulation of the VO2 membrane and the mechanical action of the PCM ventilation valve. Its technical parameters (such as a thermal conductivity switching ratio of 6:1 and a ventilation valve response time of 8 seconds) and collaborative control strategies solve the pain points of rigidity and sluggish response in traditional building envelopes, providing reliable technical support for dynamic energy saving.

[0121] Expanding applications and future upgrades, photovoltaic integration, transparent perovskite solar cells (efficiency >15%) are superimposed on VO2 coating to achieve energy self-sufficiency; digital twin integration, through BIM model pre-training MARL algorithm, shortens the on-site commissioning cycle to 24 hours.

[0122] Energy saving indicators: Air conditioning energy consumption is reduced by 40% (compared to traditional curtain walls).

[0123] Improved comfort: Indoor temperature fluctuation range reduced from ±3℃ to ±0.5℃.

[0124] Economic viability: Investment payback period of 3-5 years (for regions where the peak-valley electricity price difference is >0.8 yuan / kWh).

[0125] This technology will reshape the paradigm of building energy conservation, driving the industry to leap from "equipment energy conservation" to "structural intelligence," and has significant social, environmental, and commercial value.

[0126] Significantly reduces building energy consumption, with air conditioning energy saving rate of over 40%. Through the dynamic reflection of VO2 coating and the synergistic control of PCM heat storage and release, it reduces solar radiation heat gain and cooling loss.

[0127] Indoor thermal comfort is improved, with temperature fluctuation range reduced by 80%, down to ±0.7℃, meeting the ISOAA level comfort standard (fluctuation ≤1.0℃). Thermal unevenness is improved, and multi-agent regional control reduces the temperature difference between different facades of the building from 5℃ to 1℃, eliminating local overheating problems such as "western sun exposure".

[0128] It boasts significant economic advantages, a short investment recovery period, low operation and maintenance costs, and has no moving or wear-prone parts (VO2 is a solid film). The annual maintenance cost is less than 5 yuan / ㎡, which is lower than the 15-20 yuan / ㎡ annual maintenance cost of traditional air conditioning systems.

[0129] The environmental and social benefits are significant, including carbon emission reduction, which helps achieve the "dual carbon" target. Through demand response (such as nighttime cooling), it reduces the pressure on the power grid peak-valley difference, enhances the capacity to absorb renewable energy, and alleviates the contradiction between power supply and demand.

[0130] It has good resource sustainability, with PCM microcapsules having a recyclability rate of >90% (compared to <30% for traditional insulation materials), and VO2 coating lifespan of >15 years, reducing the frequency of building renovations.

[0131] Technological innovation breakthroughs and dynamic thermal resistance adjustment enable real-time stepless adjustment of the thermal conductivity of the building envelope (0.5-3.0 W / m·K), breaking through the limitations of the static thermal performance of traditional materials.

[0132] It has good reliability and adaptability, maintaining a reflectivity of >85% at a high temperature of 45℃ (compared to <60% for traditional curtain walls). It automatically switches to a safe mode (ventilation valve half-open, heat conduction state in VO2) during typhoon and rainstorm weather. It also has long-term stability, with the performance of the VO2 coating decreasing by <5% after 100,000 cycles, and the PCM microcapsule cycle life exceeding 10,000 cycles (as required by national standards).

[0133] Shanghai World Financial Center Smart Curtain Wall Renovation Project

[0134] Project Background: Shanghai World Financial Center (492 meters high, 101 floors), with a glass curtain wall area of ​​120,000 square meters, an average annual air conditioning energy consumption of 120 million kWh, accounting for 35% of operating costs.

[0135] Pain point analysis: The surface temperature of the curtain wall reaches 65℃ in summer, and the peak cooling load is 8MW; the existing Low-E glass + sunshade solution only saves 15% energy and affects lighting; the peak-valley electricity price difference is 1.0 yuan / kWh (peak period 10:00-15:00).

[0136] Patented technology implementation plan, hardware modification: using magnetron sputtering equipment, VO2 coating area of ​​80,000㎡, thickness of 320nm, driving voltage of 0-10VDC; PCM microcapsule layer is sprayed by construction drone or manual with a thickness of 20mm, phase change enthalpy of 185kJ / kg, ventilation valve density: 4 / ㎡.

[0137] Sensor deployment: fiber optic temperature measurement points (10m spacing), weather station roof + cantilevered platforms every 30 floors, fiber optic cables laid in conduits (bending radius > 10cm), LoRa base station coverage (transmission distance 500m).

[0138] AI system deployment, edge computing nodes, deployed in the equipment room every 20 floors (5 in total).

[0139] Algorithm configuration: MADDPG parameters (single agent).

[0140] state_dim=15

[0141] action_dim=2

[0142] hidden_dim=256

[0143] batch_size=1024

[0144] Training data: historical data, hourly energy consumption records for one year (17,520 sets), simulation data, and extreme weather scenarios (typhoons, heat waves) generated by EnergyPlus.

[0145] Key event analysis: 10:00-12:00 (peak electricity price), VO2 is switched to 0.8V (thermal conductivity 1.0W / m·K), reflectivity is increased to 82%, ventilation valve opening is 70%, natural convection heat dissipation is utilized, and mechanical refrigeration demand is reduced; air conditioning power is reduced from 2.8MW to 1.6MW (energy saving 43%).

[0146] 14:00 (Extreme High Temperature): AI detected that the glass temperature exceeded the threshold (60℃) and triggered the emergency mode: VO2 forced 0.5V (minimum thermal conductivity), reflectivity 88%; ventilation valve fully open (90%), and start active water cooling circulation of PCM layer (optional patented technology).

[0147] In terms of energy consumption, the average monthly cooling energy consumption was 9.8 million kWh before the renovation and 5.88 million kWh after the renovation, a reduction of 40%; the indoor temperature compliance rate was 78% before the renovation and 95% after the renovation, an increase of 17%.

[0148] In terms of economics, the renovation cost is 10,000 yuan (267 yuan / ㎡), which translates to 3.92 million kWh × 1.0 yuan = 3.92 million yuan per year. It takes 8.2 years to break even and reduces carbon emissions by 3.92 million kWh × 0.785 kg CO2 / kWh = 3074 tons.

[0149] This case demonstrates the feasibility of the patented technology in super high-rise buildings, with a 40% reduction in energy consumption that far exceeds traditional energy-saving retrofits (typically <20%); the 8-year payback period for commercial value aligns with expectations for commercial real estate investment.

[0150] The social value of a single building is equivalent to reducing carbon emissions by tons per year, which is equivalent to planting 160,000 trees.

[0151] This case study can provide a replicable technological paradigm for similar buildings (such as Guangzhou East Tower and Shenzhen Ping An Center).

[0152] Compared with existing technologies, it demonstrates significant advantages in energy consumption, comfort, and economy, driving the industry to leap from "equipment energy saving" to "structural intelligence" and possessing significant social, environmental, and commercial value.

[0153] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A dynamic temperature control system for intelligent building envelopes based on reinforcement learning, characterized in that: It includes a data acquisition layer, a multi-agent decision-making layer, an envelope execution layer, and an energy consumption optimization layer; The data acquisition layer is used to detect current regional meteorological data, building envelope status data, and indoor temperature and humidity. The multi-agent decision-making layer includes an edge computing unit and a multi-agent reinforcement learning engine. The edge computing unit processes all data collected in the data acquisition layer and performs preprocessing. The multi-agent reinforcement learning engine outputs commands to the building envelope execution layer based on the data preprocessed by the edge computing unit and through the MADDPG algorithm. The MADDPG algorithm in the multi-agent reinforcement learning engine coordinates the output actions of the agents through the Actor network and the Critic network. The specific expression of the local observation state of the i-th agent is as follows: ; Where t represents the discrete time period of policy interaction and decision-making in the reinforcement learning system, and is the core index for reinforcement learning training in the MADDPG algorithm, used for: state transition ,action k represents the number of consecutive iterations, the optimization step size in EnergyPlus parameter calibration, and i refers to the i-th building facade area. and These represent the indoor temperature and humidity corresponding to the i-th building facade area, respectively. , Let represent the outdoor temperature and humidity corresponding to the i-th building facade area. Let be the solar radiation intensity received by the i-th building facade area. Let be the real-time electricity price at time t. For the heat load demand of the i-th building facade area, The time feature corresponding to time t is obtained by concatenating the local observation states of 1 to N agents, as shown in the following expression: ; Where: s represents the set of all agents concatenated together. ~ Let represent the state vectors of agents 1 to N, and so on. The action expression for each agent is as follows: ; in: This represents the set of concatenated action vectors. ~ These represent the action vectors corresponding to agents 1 through N, respectively. The driving voltage for the VO2 coating controlled by the i-th intelligent agent. The opening degree of the PCM ventilation valve controlled by the i-th intelligent agent; The enclosure structure execution layer adjusts the temperature of the enclosure structure based on the command output of the intelligent agent decision layer; The energy consumption optimization layer includes a digital twin model and an incremental learning module. The digital twin model is simulated in real time using EnergyPlus, and the incremental learning module is used to provide feedback to the multi-agent decision-making layer.

2. The intelligent building envelope dynamic temperature control system based on reinforcement learning according to claim 1, characterized in that: The specific means of controlling the execution layer of the enclosure structure include variable thermal conductive coating drive and PCM ventilation valve regulation; The variable thermal conductivity coating drive is used to control the voltage across the VO2 film to adjust the thermal conductivity of the VO2 film. The PCM ventilation valve is controlled by controlling the stepper motor to drive the louvers, thereby controlling the rate of heat storage and release.

3. The intelligent building envelope dynamic temperature control system based on reinforcement learning according to claim 1, characterized in that: The digital twin model is simulated in real time using EnergyPlus, and the building thermodynamic model is updated every 15 minutes. The specific expression is as follows: ; in: This includes parameters related to the thermal properties of the material, where argmin represents the parameter corresponding to the minimum value. Let k be the measured indoor temperature, heat flux on the wall or curtain wall surface, total energy consumption of the area, and electrical input power. The simulation output of the EnergyPlus model under the current input conditions is executed every 15 minutes. The simulation output generated by EnergyPlus for the digital twin model is as follows: ; in: This represents the measured indoor temperature at time k. This indicates the heat flux on the surface of a wall or curtain wall. This represents the total energy consumption of the region. Indicates electrical input power; ; in: For input conditions, Let be the forward simulation function of the EnergyPlus model. The digital twin model's prediction of load and temperature within the next Δt minutes at time k is as follows: ; in: For predicted future temperature and heat load, The input conditions are defined at time k, and the output of the digital twin model is fed back to the Critic network.