Building air conditioning system control optimization method based on mechanism-reinforcement learning and intelligent agent

By building a multi-region thermodynamic mechanism model and a digital twin simulation environment, combined with reinforcement learning agents, the problems of low energy efficiency and fluctuations in traditional air conditioning systems are solved, efficient and accurate air conditioning system control is achieved, and training efficiency and physical interpretability are improved.

CN120491485AInactive Publication Date: 2025-08-15ANHUI YUANLI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510830180.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional building air conditioning system control methods rely on static mechanism models and are difficult to adapt to dynamic environmental changes and complex working conditions, resulting in low energy efficiency and fluctuations in comfort; although purely data-driven reinforcement learning methods have adaptability, they lack physical interpretability, and there are problems of low training efficiency and difficulty in convergence.

Method used

A multi-region thermodynamic mechanism model is built, combined with the digital twin simulation environment and real system transfer learning, a reinforcement learning agent is designed, a preliminary strategy model is generated through a hybrid training strategy, and the control parameters are dynamically adjusted in real time at the edge end, and a lightweight agent is used for online control.

Benefits of technology

It realizes efficient and accurate control of building air conditioning systems, improves energy efficiency and comfort, overcomes the uninterpretationary defects of pure data-driven methods, and improves training efficiency and physical interpretability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005459117210000021
    Figure BDA0005459117210000021
  • Figure BDA0005459117210000032
    Figure BDA0005459117210000032
  • Figure BDA0005459117210000061
    Figure BDA0005459117210000061
Patent Text Reader

Abstract

The invention discloses a building air conditioning system control optimization method based on mechanism-reinforcement learning and an intelligent agent, and relates to the technical field of building energy saving and intelligent control. Comprising the steps of constructing a multi-region thermodynamic mechanism model, carrying out mixed training on a strategy, carrying out online dynamic control and deploying a lightweight agent, and generating a preliminary strategy model by constructing the multi-region thermodynamic mechanism model and combining a digital twinborn simulation environment and real system transfer learning. The model can dynamically control a building air conditioning system on line, and environment data are collected in real time and control parameters are dynamically adjusted by designing a reinforcement learning agent. The problems that a traditional building air conditioning system control method is low in energy efficiency and fluctuates in comfort degree are solved, and meanwhile the physical interpretability and the training efficiency of a pure data driven reinforcement learning method are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of building energy conservation and intelligent control, and in particular to a building air-conditioning system control optimization method and an intelligent agent based on mechanism-reinforcement learning. Background Art

[0002] Building energy consumption accounts for a significant portion of China's total energy consumption, with heating, ventilation, and air conditioning (HVAC) systems being the primary energy consumer. With the acceleration of urbanization and improvements in living standards, the demand for improved comfort and energy conservation within buildings continues to rise. Due to individual differences, users in different regions may have varying thermal preferences, which impacts their actual comfort levels. Furthermore, the significant differences in winter and summer climates lead to significant peaks in seasonal energy demand, further exacerbating energy consumption. To address these challenges, improving the energy efficiency of HVAC systems and optimizing their control strategies are urgent issues.

[0003] Traditional building air conditioning system control methods (such as PID control and rule-based control) rely on static mechanism models and struggle to adapt to dynamic environmental changes and complex operating conditions, resulting in low energy efficiency and fluctuating comfort levels. While purely data-driven reinforcement learning methods offer adaptive capabilities, they lack physical interpretability and suffer from low training efficiency and convergence difficulties. Existing technologies have yet to effectively address the dynamic balance between energy efficiency optimization and indoor environmental comfort. To address this, we propose a control optimization method and intelligent agent for building air conditioning systems based on mechanism-based reinforcement learning. Summary of the Invention

[0004] The purpose of this invention is to address the problem that traditional building air conditioning system control methods (such as PID control and rule-based control) rely on static mechanism models and are difficult to adapt to dynamic environmental changes and complex operating conditions, resulting in low energy efficiency and fluctuating comfort levels. While purely data-driven reinforcement learning methods have adaptive capabilities, they lack physical interpretability and suffer from low training efficiency and convergence difficulties. This invention provides a building air conditioning system control optimization method and intelligent agent based on mechanism-based reinforcement learning.

[0005] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0006] A control optimization method for building air conditioning systems based on mechanism-reinforcement learning includes the following steps:

[0007] Step S1: Construct a multi-region thermodynamic mechanism model, obtain building CAD data, building material thermal parameters (λ, c, ρ) and historical meteorological data, establish a regional heat balance equation, calibrate the parameters to a prediction error of ≤ 0.5°C, describe the indoor and outdoor heat exchange dynamics, and output a digital twin simulation environment (FMU format);

[0008] Step S2: Hybrid training strategy: Combine the digital twin simulation environment, equipment performance curves, and comfort constraints to conduct digital twin pre-training and real system transfer learning to generate a preliminary strategy model. Through the hybrid simulation and real data training strategy, energy efficiency and comfort are jointly optimized.

[0009] Step S3: Online dynamic control. A reinforcement learning agent is designed. Its state space contains environmental parameters and its action space is the air conditioning control instructions. The edge generates control instructions every 30 seconds, obtains real-time sensor data, including temperature, humidity, and CO2 concentration, inputs the preliminary strategy model, calculates the air conditioning setting parameters, including temperature, wind speed, and mode, and outputs the air conditioning control instructions.

[0010] Step S4: deploy lightweight intelligent agents, collect environmental data in real time and dynamically adjust control parameters.

[0011] Furthermore, the multi-region thermodynamic mechanism model in step S1 divides the building into n thermodynamic sub-regions (n≥3), and establishes an independent heat balance equation for each sub-region:

[0012]

[0013] Among them, C i is the heat capacity of zone i, C i =∑(ρ m c m V m ), building material density ρ×specific heat capacity c×volume V,k ij is the heat transfer coefficient between zone i and zone j, is the cooling capacity of the air conditioner, is solar radiation, calculated by the radiation model α is the window-to-wall ratio, A is the surface area, I is the radiation intensity, Produces heat for personnel.

[0014] Furthermore, the thermodynamic model parameter calibration in step S1 uses high-precision model parameter identification, and inputs the 72-hour measured temperature sequence T real and air conditioning operation log, and the L-BFGS algorithm is used to solve it, iterating 50 times. The calculation formula is:

[0015]

[0016] After calibration, k ij and C i , model error ≤ 0.5℃.

[0017] Furthermore, in step S2, the digital twin simulation environment is constructed using TMY3 meteorological data generated by EnergyPlus, using the double-delay DDPG algorithm, the Actor network parameter is 1.2M, the acceleration strategy convergence is accelerated, and high td-error samples are sampled first, with probability weights:

[0018]

[0019] Among them, α=0.6, ε=1e-5.

[0020] Furthermore, the real system transfer learning adopts feature distribution alignment technology, and the loss function is:

[0021]

[0022] The MMD calculation uses a Gaussian kernel function with a bandwidth of σ = 1.0, freezes the weights of the feature extraction layer, fine-tunes the LSTM layer and the policy head, and achieves a performance loss of <5% after policy migration.

[0023] Furthermore, in step S3, the intelligent agent inputs 12-dimensional sensor data every 30 seconds, and the action space constraints include hard constraints and soft constraints. The hard constraint is the temperature setting: 18°C ≤ T_set ≤ 28°C, and the soft constraint is the fan speed: gradient change rate ≤ 10% / min.

[0024] An intelligent agent for controlling a building air conditioning system based on mechanism-reinforcement learning, characterized by comprising:

[0025] Input module, supports Modbus / BACnet protocol, collects 12-dimensional environmental parameters;

[0026] The decision-making core module uses a policy network with an LSTM-based Actor-Critic architecture, with 1.2M parameters, real-time calculation of the barrier function h(s), and dynamic adjustment of the action mask.

[0027] Output module, generates 7-dimensional control instructions, including 4-dimensional continuous actions and 3-dimensional discrete actions;

[0028] Energy efficiency evaluation module calculates energy saving rate and comfort index in real time, helping operation and maintenance personnel to quickly identify abnormal periods;

[0029] Edge-cloud collaboration module: The edge performs real-time inference through Jetson TX2 with a latency of ≤150ms, and the cloud incrementally updates the model every 24 hours.

[0030] Furthermore, the 12-dimensional environmental parameters collected by the input module include

[0031] Environmental parameters: indoor temperature and humidity (2), outdoor temperature and humidity (2), CO2 concentration (1);

[0032] Equipment status: air conditioner outlet water temperature (1), fan speed (1), valve opening (1);

[0033] Spatiotemporal features: time period coding (sin / cos timestamp, 2), regional population density (1);

[0034] Energy consumption indicator: cumulative power consumption in the past hour (1).

[0035] Furthermore, the policy network includes

[0036] Feature extraction layer, which maps the 12-dimensional original state to a 64-dimensional feature space to capture nonlinear relationships;

[0037] 128-unit LSTM layer: memorizes the state of the previous 3 hours and recognizes temporal features such as day and night patterns;

[0038] Actor head: outputs mean μ and variance σ, defining the Gaussian policy distribution;

[0039] Critic head: estimates the state value V(s) and guides policy gradient updates.

[0040] Furthermore, the output module's 4-dimensional continuous actions include a chilled water flow setting value of 0.8-1.2 times the rated flow, a supply air temperature setting value of 16-20°C, a zone valve opening adjustment of 0-100% and a fresh air mixing ratio adjustment of 10%-100%. The 3-dimensional discrete actions include three operating mode selections: cooling / dehumidification / ventilation, low-speed / medium-speed / high-speed fan gear selection and an emergency operation instruction for forced frequency reduction when overload protection is triggered.

[0041] The beneficial effects of the present invention are as follows:

[0042] 1. The present invention generates a preliminary policy model by constructing a multi-region thermodynamic mechanism model, combining a digital twin simulation environment with real-world system transfer learning. This model is capable of online dynamic control of building air conditioning systems. By designing a reinforcement learning agent, environmental data is collected in real time and control parameters are dynamically adjusted. The agent includes an input module, a decision-making core module, an output module, an energy efficiency evaluation module, and an edge-cloud collaboration module, which can efficiently and accurately achieve control optimization of building air conditioning systems. By providing physical constraints through the thermodynamic model, reinforcement learning achieves dynamic optimization, overcoming the uninterpretability of pure data-driven methods. The three-stage training framework of "digital twin pre-training + online transfer + continuous learning" improves the training efficiency by 5 times compared to traditional RL. A lightweight policy network solution (parameter size <1MB) is designed to achieve 10ms real-time response. A constrained reinforcement learning method based on a safety barrier function is proposed to ensure that the control strategy always operates within the safety boundary. Through this method, the present invention solves the problems of low energy efficiency and comfort fluctuations in traditional building air conditioning system control methods, while improving the physical interpretability and training efficiency of pure data-driven reinforcement learning methods.

[0043] 2. The intelligent agent of the present invention first receives 12-dimensional environmental parameters from edge sensors in real time through the input module. These parameters cover indoor and outdoor temperature and humidity, CO2 concentration, the operating status of air conditioning equipment, as well as spatiotemporal characteristics and energy consumption indicators, providing the intelligent agent with comprehensive environmental information. The decision-making core module then processes and analyzes this input data using a policy network based on an actor-critic architecture with LSTM. The policy network uses a feature extraction layer to map the raw state into a high-dimensional feature space, capturing nonlinear relationships. It also uses an LSTM layer to memorize previous states and identify temporal features. The actor head outputs the mean and variance of a Gaussian policy distribution based on the current state, while the critic head estimates the state value to guide policy gradient updates. When generating control instructions, the output module generates 7-dimensional control instructions, consisting of 4-dimensional continuous actions and 3-dimensional discrete actions, based on the output of the decision-making core module. These instructions are sent to the air conditioning equipment within the building to guide their corresponding operation. Simultaneously, the energy efficiency evaluation module monitors and evaluates the operating efficiency of the air conditioning equipment in real time, uploading the evaluation results to the cloud management platform. The cloud-based management platform optimizes the energy efficiency model of the entire system based on this data, achieving continuous system improvement and energy efficiency enhancement. In this way, the present invention can achieve intelligent control of building air conditioning systems, improve comfort and energy efficiency standards, reduce energy waste, and achieve more intelligent and efficient building environment management. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below.

[0045] The present invention provides a building air conditioning system control optimization method based on mechanism-reinforcement learning, which includes the following steps:

[0046] Step S1: Construct a multi-region thermodynamic mechanism model, obtain building CAD data, building material thermal parameters (λ, c, ρ) and historical meteorological data, establish a regional heat balance equation, calibrate the parameters to a prediction error of ≤ 0.5°C, describe the indoor and outdoor heat exchange dynamics, and output a digital twin simulation environment (FMU format);

[0047] Step S2: Hybrid training strategy: Combine the digital twin simulation environment, equipment performance curves, and comfort constraints to conduct digital twin pre-training and real system transfer learning to generate a preliminary strategy model. Through the hybrid simulation and real data training strategy, energy efficiency and comfort are jointly optimized.

[0048] Step S3: Online dynamic control. A reinforcement learning agent is designed. Its state space contains environmental parameters and its action space is the air conditioning control instructions. The edge generates control instructions every 30 seconds, obtains real-time sensor data, including temperature, humidity, and CO2 concentration, inputs the preliminary strategy model, calculates the air conditioning setting parameters, including temperature, wind speed, and mode, and outputs the air conditioning control instructions.

[0049] Step S4: deploy lightweight intelligent agents, collect environmental data in real time and dynamically adjust control parameters.

[0050] In this embodiment, preferably, the multi-region thermodynamic mechanism model in step S1 divides the building into n thermodynamic sub-regions (n≥3), and establishes an independent heat balance equation for each sub-region:

[0051]

[0052] Among them, C i is the heat capacity of zone i, C i =∑(ρ m c m V m ), building material density ρ×specific heat capacity c×volume V,k ij is the heat transfer coefficient between zone i and zone j, is the cooling capacity of the air conditioner, is solar radiation, calculated by the radiation model α is the window-to-wall ratio, A is the surface area, I is the radiation intensity, Produces heat for personnel.

[0053] The temperature change rate of each sub-area is determined by heat capacity, heat transfer coefficient, air conditioning cooling capacity, solar radiation, and occupant heat generation. Heat capacity reflects the building material's ability to store thermal energy, while the heat transfer coefficient describes the efficiency of heat transfer between different areas. Air conditioning cooling capacity, as an active control measure, directly affects indoor temperature. Solar radiation enters the room through windows, and its magnitude is related to the window-to-wall ratio, surface area, and radiation intensity. Occupant heat generation is estimated based on the number of people indoors, typically assuming 100W of heat generation per person. These parameters work together in the heat balance equation to achieve refined control of the building's air conditioning system.

[0054] In this embodiment, preferably, the thermodynamic model parameter calibration in step S1 uses high-precision model parameter identification, and inputs a 72-hour measured temperature sequence T real and air conditioning operation log, and the L-BFGS algorithm is used to solve it, iterating 50 times. The calculation formula is:

[0055]

[0056] After calibration, k ij and C i , model error ≤ 0.5℃.

[0057] In step S2, the calibrated thermodynamic model parameters are combined with mechanistic knowledge and reinforcement learning methods to construct an agent policy network. This policy network intelligently determines the optimal control action based on the current environmental conditions (such as indoor temperature, outdoor temperature, and solar radiation intensity) and the operating status of the air conditioning system, achieving precise control of indoor temperature and minimizing energy consumption. The trained agent achieves refined control of the building's air conditioning system, effectively improving indoor temperature comfort and energy efficiency.

[0058] In this embodiment, preferably, the digital twin simulation environment in step S2 is constructed using TMY3 meteorological data generated by EnergyPlus, using a double-delay DDPG algorithm, an Actor network parameter of 1.2M, an acceleration strategy convergence, and a priority sampling of high td-error samples, with a probability weight of:

[0059]

[0060] Among them, α=0.6, ε=1e-5.

[0061] In the digital twin simulation environment, TMY3 meteorological data generated by EnergyPlus can simulate a variety of complex meteorological conditions, providing a rich set of environmental states for agent training. The application of the double-delayed DDPG algorithm enables the Actor network parameter count to reach 1.2M, effectively accelerating policy convergence and improving training efficiency. The strategy of prioritizing samples with high td-error (TD-error) and adjusting probability weights allows the agent to focus more on states with large prediction errors, enabling targeted optimization and further improving the accuracy of the control strategy. The values of and are optimized after multiple experimental verifications, ensuring algorithm stability and convergence. This digital twin simulation environment enables agents to train and learn in scenarios close to the real world, laying a solid foundation for refined control of building air conditioning systems.

[0062] In this embodiment, preferably, the real system transfer learning adopts feature distribution alignment technology, and the loss function is:

[0063]

[0064] The MMD calculation uses a Gaussian kernel function with a bandwidth of σ = 1.0, freezes the weights of the feature extraction layer, fine-tunes the LSTM layer and the policy head, and achieves a performance loss of <5% after policy migration.

[0065] The application of feature distribution alignment technology effectively narrows the distribution difference between the simulation environment and the real system, allowing the strategy learned by the agent in the simulation environment to be more smoothly transferred to the real system. Through a carefully designed loss function, not only the matching degree of the feature distribution is taken into account, but also a penalty term for weight difference is added to ensure that the migrated strategy can still maintain high performance in the real system. The choice of Gaussian kernel function and the fine setting of bandwidth σ further enhance the stability and accuracy of MMD calculation. During the transfer learning process, the weights of the feature extraction layer are frozen, and only the LSTM layer and the strategy head are fine-tuned. This strategy not only ensures the efficiency of the transfer, but also maximizes the retention of the knowledge learned by the agent in the simulation environment. The performance loss after the strategy transfer is controlled within 5%, which fully verifies the effectiveness and practicality of the transfer learning method in this embodiment.

[0066] In this embodiment, preferably, in step S3, the intelligent agent inputs 12-dimensional sensor data every 30 seconds, and the action space constraints include hard constraints and soft constraints. The hard constraint is the temperature setting: 18°C ≤ T_set ≤ 28°C, and the soft constraint is the fan speed: gradient change rate ≤ 10% / min.

[0067] The intelligent agent collects data every 30 seconds, ensuring the real-time and responsiveness of the control strategy. The hard constraint setting, which limits the temperature setting range to 18°C to 28°C, meets human comfort requirements while avoiding unnecessary energy waste. Soft constraints limit the gradient change rate of the fan speed to no more than 10% / min, which helps reduce mechanical stress on the system and extend equipment life. This also ensures a smooth transition of the indoor environment and avoids discomfort caused by sudden changes in wind speed. Through this constraint design, the intelligent agent can balance efficiency and comfort during the optimization control process, achieving refined management of the building's air conditioning system.

[0068] By constructing a multi-region thermodynamic mechanism model and combining it with a digital twin simulation environment and real-world system transfer learning, a preliminary strategy model was generated. This model is capable of online dynamic control of building air conditioning systems. By designing a reinforcement learning agent, it collects environmental data in real time and dynamically adjusts control parameters. The agent includes an input module, a decision-making core module, an output module, an energy efficiency evaluation module, and an edge-cloud collaboration module, and is capable of efficiently and accurately optimizing the control of building air conditioning systems. Through this approach, the present invention solves the problems of low energy efficiency and fluctuating comfort levels in traditional building air conditioning system control methods, while also improving the physical interpretability and training efficiency of purely data-driven reinforcement learning methods.

[0069] An intelligent agent for controlling a building air conditioning system based on mechanism-reinforcement learning, characterized by comprising:

[0070] Input module, supports Modbus / BACnet protocol, collects 12-dimensional environmental parameters;

[0071] The decision-making core module uses a policy network with an LSTM-based Actor-Critic architecture, with 1.2M parameters, real-time calculation of the barrier function h(s), and dynamic adjustment of the action mask.

[0072] Dynamic barrier function:

[0073] h(s)=min(28-T in ,T in -18,70-RH in )

[0074] Policy gradient correction;

[0075]

[0076] When h(s)<1℃, λ increases linearly from 0.1 to 1.0.

[0077] Output module, generates 7-dimensional control instructions, including 4-dimensional continuous actions and 3-dimensional discrete actions;

[0078] Energy efficiency evaluation module calculates energy saving rate and comfort index in real time, helping operation and maintenance personnel to quickly identify abnormal periods;

[0079] The energy efficiency evaluation module inputs historical energy consumption data E base , real-time operation data E new Calculate the energy saving rate and comfort compliance rate using the following formula:

[0080]

[0081] Converting real-time energy consumption, comfort, and other data into charts (such as time-sharing energy consumption curves and PMV distribution heat maps) helps operation and maintenance personnel quickly identify abnormal periods. Archiving 30 days of operating data meets the requirements of energy efficiency management standards such as ISO 50001 and provides evidence for audits. By comparing historical data vertically, equipment performance degradation trends can be identified. Energy efficiency indicators in the report (such as energy saving rate and COP improvement rate) serve as the basis for revising the reward function of reinforcement learning. The correction formula is:

[0082] R new =R+η·(actual energy saving rate-expected energy saving rate)

[0083] Among them, η is the learning rate, which dynamically adjusts the optimization direction of the policy network.

[0084] The edge-cloud collaboration module uses a Jetson TX2 to perform real-time inference at the edge with a latency of ≤150ms, and the cloud incrementally updates the model every 24 hours. The edge uploads status data every 5 minutes, and the cloud aggregates this data to generate a new training set. Model training is performed using a TPUv3 at a speed of 1000 steps / min. MQTT is used for data transmission between the edge and cloud, with bandwidth usage of ≤10KB / s.

[0085] In this embodiment, preferably, the 12-dimensional environmental parameters collected by the input module include

[0086] Environmental parameters: indoor temperature and humidity (2), outdoor temperature and humidity (2), CO2 concentration (1);

[0087] Equipment status: air conditioner outlet water temperature (1), fan speed (1), valve opening (1);

[0088] Spatiotemporal features: time period coding (sin / cos timestamp, 2), regional population density (1);

[0089] Energy consumption indicator: cumulative power consumption in the past hour (1).

[0090] Twelve environmental parameters together form the basis for the intelligent agent's decision-making. Indoor temperature and humidity reflect the building's internal comfort level, while outdoor temperature and humidity affect the load demand of the air conditioning system. CO2 concentration is a key indicator of indoor air quality; excessively high CO2 concentration can cause discomfort. Air conditioning outlet water temperature, fan speed, and valve opening are directly related to the air conditioning system's energy consumption and cooling / heating performance. Time period encoding uses sin / cosine timestamps to encode different time periods of the day as two-dimensional features, helping the intelligent agent recognize and adapt to daily temperature fluctuations and occupant activity patterns. Regional occupancy density reflects the frequency of use in different areas and has a significant impact on the air conditioning system's control strategy. The accumulated power consumption over the past hour, as an energy consumption indicator, helps the intelligent agent evaluate the effectiveness of the current control strategy and dynamically adjust it. By comprehensively considering these parameters, the intelligent agent can achieve efficient control of the building's air conditioning system.

[0091] In this embodiment, preferably, the policy network includes

[0092] Feature extraction layer, which maps the 12-dimensional original state to a 64-dimensional feature space to capture nonlinear relationships;

[0093] 128-unit LSTM layer: memorizes the state of the previous 3 hours and recognizes temporal features such as day and night patterns;

[0094] Actor head: outputs mean μ and variance σ, defining the Gaussian policy distribution;

[0095] Critic head: estimates the state value V(s) and guides policy gradient updates.

[0096] Policy Network

[0097]

[0098] Learning rate 3e-4, batch_size = 512

[0099] In this embodiment, preferably, the 4-dimensional continuous action of the output module includes a chilled water flow setting value of 0.8-1.2 times the rated flow, a supply air temperature setting value of 16-20°C, a zone valve opening adjustment of 0-100% and a fresh air mixing ratio adjustment of 10%-100%, and the 3-dimensional discrete action includes three operating mode selections of cooling / dehumidification / ventilation, low speed / medium speed / high speed fan gear selection and an emergency operation instruction of forced frequency reduction when overload protection is triggered.

[0100] The parameter range of chilled water flow setting value is 0.8Q rated -1.2Q rated , Q ratedThe rated flow rate of the water pump (unit: m3 / h) is used to adjust the chilled water flow rate to control the cooling output. The flow rate, heat exchange efficiency, and cooling capacity are positively correlated. Flow rate exceeding the safe range may cause pipeline vibration or equipment overload.

[0101] The supply air temperature setting value parameter range is: 16℃~20℃. The setting value for summer conditions is close to the lower limit (16-18℃) and the setting value for transition seasons is moderately increased (18-20℃). The deviation between the supply air temperature setting value and the indoor temperature must meet ΔT≤5℃ / h to prevent condensation.

[0102] The range of regional valve opening adjustment parameters is: 0% (closed) ~ 100% (fully open), and the cooling capacity is distributed according to the regional heat load:

[0103]

[0104] Among them, Q nece,i is the real-time cooling requirement of the i-th zone (kW).

[0105] Fresh air mixing ratio adjustment parameter range: 10% to 100% (fresh air ratio), the minimum fresh air volume meets the health standard ASHRAE 62.1, the maximum fresh air volume is used for free cooling, when T out <T in At -3℃, priority should be given to increasing the fresh air ratio.

[0106] The switching logic for the three operating modes of cooling / dehumidification / ventilation is as follows:

[0107] When Tin>Tset+1℃ and RH<65%→cooling mode;

[0108] When RH>70% → dehumidification mode;

[0109] When Tout∈[Tset-3℃,Tset]→ventilation mode.

[0110] The fan gear selection of low speed / medium speed / high speed is dynamically adjusted according to the PMV comfort index:

[0111]

[0112] The triggering conditions for the emergency operation command are: current > 110% of the rated value, communication timeout > 30 seconds, and the action content is 0: reduced frequency operation (compressor frequency drops below 50Hz), 1: switch to the backup controller (enable PID algorithm) and 2: safety shutdown (turn off the refrigerator and only maintain ventilation).

[0113] Policy network output:

[0114] Continuous action: Output normalized value (-1 to 1), which is linearly mapped to the physical range:

[0115] a cont =a norm ×(a max -a min )+a min

[0116] Discrete action: Output 3D probability distribution and take the maximum probability item.

[0117] The above actions and parameters are selected to optimize the performance of the building's air conditioning system while ensuring system stability and safety. For example, by precisely regulating the chilled water flow, the system can dynamically adjust cooling output based on actual demand, thereby avoiding energy waste and equipment overload. Properly setting the supply air temperature not only meets indoor comfort requirements but also effectively prevents condensation. Adjusting the opening of zone valves ensures that cooling capacity is distributed to each zone as needed, improving the system's energy efficiency.

[0118] When the agent is in use, it first receives 12-dimensional environmental parameters from edge sensors in real time through the input module. These parameters cover indoor and outdoor temperature and humidity, CO2 concentration, air conditioning equipment operating status, spatiotemporal characteristics, and energy consumption indicators, providing the agent with comprehensive environmental information. The decision-making core module then processes and analyzes this input data using a policy network based on an actor-critic architecture with LSTM. The policy network uses a feature extraction layer to map the raw state into a high-dimensional feature space, capturing nonlinear relationships. It also uses an LSTM layer to memorize previous states and identify temporal features. The actor head outputs the mean and variance of a Gaussian policy distribution based on the current state, while the critic head estimates the state value to guide policy gradient updates. When generating control commands, the output module generates 7-dimensional control commands, consisting of 4-dimensional continuous actions and 3-dimensional discrete actions, based on the output of the decision-making core module. These commands are sent to the air conditioning equipment within the building to guide their operation. Simultaneously, the energy efficiency evaluation module monitors and evaluates the operating efficiency of the air conditioning equipment in real time, uploading the results to a cloud-based management platform. The cloud-based management platform optimizes the energy efficiency model of the entire system based on this data, achieving continuous system improvement and energy efficiency enhancement. In this way, the present invention can achieve intelligent control of building air conditioning systems, improve comfort and energy efficiency standards, reduce energy waste, and achieve more intelligent and efficient building environment management.

[0119] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A building air conditioning system control optimization method based on mechanism-reinforcement learning, characterized in that: It includes the following steps: Step S1: Construct a multi-region thermodynamic mechanism model, obtain building CAD data, building material thermal parameters (λ, c, ρ) and historical meteorological data, establish a regional heat balance equation, calibrate the parameters to a prediction error of ≤ 0.5°C, describe the indoor and outdoor heat exchange dynamics, and output a digital twin simulation environment (FMU format); Step S2: Hybrid training strategy: Combine the digital twin simulation environment, equipment performance curves, and comfort constraints to conduct digital twin pre-training and real system transfer learning to generate a preliminary strategy model. Through the hybrid simulation and real data training strategy, energy efficiency and comfort are jointly optimized. Step S3: Online dynamic control. A reinforcement learning agent is designed. Its state space contains environmental parameters and its action space is the air conditioning control instructions. The edge generates control instructions every 30 seconds, obtains real-time sensor data, including temperature, humidity, and CO2 concentration, inputs the preliminary strategy model, calculates the air conditioning setting parameters, including temperature, wind speed, and mode, and outputs the air conditioning control instructions. Step S4: deploy lightweight intelligent agents, collect environmental data in real time and dynamically adjust control parameters.

2. The building air conditioning system control optimization method based on mechanism-reinforcement learning according to claim 1 is characterized by: In step S1, the multi-region thermodynamic mechanism model divides the building into n thermodynamic sub-regions (n≥3), and establishes an independent heat balance equation for each sub-region: Among them, C i is the heat capacity of zone i, C i =∑(ρ m c m V m ), building material density ρ×specific heat capacity c×volume V,k ij is the heat transfer coefficient between zone i and zone j, is the cooling capacity of the air conditioner, is solar radiation, calculated by the radiation model α is the window-to-wall ratio, A is the surface area, I is the radiation intensity, Produces heat for personnel.

3. The control optimization method for building air conditioning system based on mechanism-reinforcement learning according to claim 2 is characterized in that: In step S1, the thermodynamic model parameter calibration uses high-precision model parameter identification, and inputs the 72-hour measured temperature sequence T real and air conditioning operation log, and the L-BFGS algorithm is used to solve it, iterating 50 times. The calculation formula is: After calibration, k ij and C i , model error ≤ 0.5℃.

4. The building air conditioning system control optimization method based on mechanism-reinforcement learning according to claim 1, characterized in that: In step S2, the digital twin simulation environment is constructed using TMY3 meteorological data generated by EnergyPlus, using the double-delayed DDPG algorithm, with an Actor network parameter of 1.2M, accelerating strategy convergence, and prioritizing high td-error samples. The probability weight is: Among them, α=0.6, ε=1e-5.

5. The building air conditioning system control optimization method based on mechanism-reinforcement learning according to claim 1 is characterized in that: The real system transfer learning adopts feature distribution alignment technology, and the loss function is: The MMD calculation uses a Gaussian kernel function with a bandwidth of σ = 1.0, freezes the weights of the feature extraction layer, fine-tunes the LSTM layer and the policy head, and achieves a performance loss of <5% after policy migration.

6. The building air conditioning system control optimization method based on mechanism-reinforcement learning according to claim 1 is characterized by: In step S3, the agent inputs 12-dimensional sensor data every 30 seconds. The action space constraints include hard constraints and soft constraints. The hard constraint is the temperature setting: 18°C ≤ T_set ≤ 28°C, and the soft constraint is the fan speed: gradient change rate ≤ 10% / min.

7. An intelligent agent implementing the method according to any one of claims 1 to 5, characterized in that: include: Input module, supports Modbus / BACnet protocol, collects 12-dimensional environmental parameters; The decision-making core module uses a policy network with an LSTM-based Actor-Critic architecture, with 1.2M parameters, real-time calculation of the barrier function h(s), and dynamic adjustment of the action mask. Output module, generates 7-dimensional control instructions, including 4-dimensional continuous actions and 3-dimensional discrete actions; Energy efficiency evaluation module calculates energy saving rate and comfort index in real time, helping operation and maintenance personnel to quickly identify abnormal periods; Edge-cloud collaboration module: The edge performs real-time inference through Jetson TX2 with a latency of ≤150ms, and the cloud incrementally updates the model every 24 hours.

8. The intelligent agent according to claim 7, characterized in that: The 12-dimensional environmental parameters collected by the input module include Environmental parameters: indoor temperature and humidity (2), outdoor temperature and humidity (2), CO2 concentration (1); Equipment status: air conditioner outlet water temperature (1), fan speed (1), valve opening (1); Spatiotemporal features: time period coding (sin / cos timestamp, 2), regional population density (1); Energy consumption indicator: cumulative power consumption in the past hour (1).

9. The intelligent agent according to claim 8, characterized in that: The policy network includes Feature extraction layer, which maps the 12-dimensional original state to a 64-dimensional feature space to capture nonlinear relationships; 128-unit LSTM layer: memorizes the state of the previous 3 hours and recognizes temporal features such as day and night patterns; Actor head: outputs mean μ and variance σ, defining the Gaussian policy distribution; Critic head: estimates the state value V(s) and guides policy gradient updates.

10. The intelligent agent according to claim 9, characterized in that: The output module's 4-dimensional continuous actions include a chilled water flow setting value of 0.8-1.2 times the rated flow, a supply air temperature setting value of 16-20°C, a zone valve opening adjustment of 0-100% and a fresh air mixing ratio adjustment of 10%-100%. The 3-dimensional discrete actions include three operating mode selections: cooling / dehumidification / ventilation, low-speed / medium-speed / high-speed fan gear selection and an emergency operation instruction for forced frequency reduction when overload protection is triggered.

Citation Information

Cited By

  • Intelligent analysis and management method and system for integrated controller

    CN120972700A

  • Multi-parameter coupling optimization control method based on deep reinforcement learning

    CN121364640A

  • Self-adaptive energy-saving control method and system for building heating and ventilation equipment based on digital twinning

    CN121804063A