A building energy-saving optimization method based on dynamic load prediction and multi-device coupling

By constructing a digital twin environment and integrating a dynamic load prediction model with an equipment coupling graph network model, and combining it with a reinforcement learning agent optimization strategy, the problem of load changes and equipment coupling relationships in building energy conservation optimization was solved, and efficient central air conditioning system management was achieved.

CN121677112BActive Publication Date: 2026-04-24INTELLIGENT TECH CO LTD OF CHINESE CONSTR THIRD ENG BUREAU
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INTELLIGENT TECH CO LTD OF CHINESE CONSTR THIRD ENG BUREAU
Filing Date
2026-02-11
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing building energy efficiency optimization methods are unable to cope with rapid load changes, complex coupling relationships between equipment, and the low efficiency of artificial intelligence systems in learning and adapting to new environments, resulting in low energy efficiency and equipment operation risks.

Method used

A digital twin environment is constructed, integrating a dynamic load forecasting model with an equipment coupled graph network model. A model-assisted reinforcement learning agent is used to optimize equipment operation strategies, and the model is dynamically updated through a closed-loop correction process to ensure the safety and efficiency of the strategies.

Benefits of technology

It enables accurate prediction of future cooling and heating loads of building central air conditioning systems and dynamic characterization of energy transfer efficiency between equipment, thereby improving system operating efficiency, reducing energy consumption, and realizing the transformation from passively waiting for response to proactive optimization management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121677112B_ABST
    Figure CN121677112B_ABST
Patent Text Reader

Abstract

The application relates to the field of energy-saving optimization, and discloses a building energy-saving optimization method based on dynamic load prediction and multi-device coupling. The method constructs a digital twin environment integrating a dynamic load prediction model and a device coupling graph network model, uses a model-assisted reinforcement learning agent to explore a device operation strategy, and performs physical feasibility verification and energy efficiency estimation before outputting an action. The agent combines future load information to perform forward-looking planning, generates an optimization strategy aiming to minimize the total energy consumption of the system under the comfort constraint, and compiles the optimization strategy into a control instruction for application to an actual central air conditioning system. Through comparison of actual operation data and simulation data, a model correction process is started to update the prediction model or the coupling network, forming a closed-loop optimization. The system can realize the transformation from passive response to active optimization, significantly improves the energy efficiency of the central air conditioning system, and reduces the building energy consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy-saving optimization, and in particular to a building energy-saving optimization method based on dynamic load forecasting and multi-device coupling. Background Technology

[0002] As a major energy-consuming device within a building, optimizing the operational efficiency of central air conditioning systems plays a crucial role in reducing overall operating costs and carbon emissions. Currently, while artificial intelligence technology demonstrates significant potential in building energy management, and intelligent control through data analysis is gradually becoming a new direction for industry development, existing systems are often limited by fixed rules or focus only on localized optimization. They struggle to flexibly respond to rapid changes in building loads and the complex interactions between various devices, which significantly restricts further improvements in energy efficiency.

[0003] Specifically, existing technologies still have some significant shortcomings in practical applications. For example, traditional control methods typically react only after actual demand arises, lacking the ability to accurately predict future load changes. This results in a consistently slow system operation and low energy efficiency. Furthermore, within central air conditioning systems, the energy conversion and interactions between devices such as chillers, pumps, and cooling towers are highly complex. Existing systems often only consider the efficiency of individual devices during optimization, without fully considering how they work together as a whole to achieve optimal results. This prevents the full realization of overall energy-saving potential. Moreover, even when artificial intelligence is introduced for decision-making, it often fails to adequately consider the physical limits and safe operating requirements of equipment, potentially leading to operational risks. Additionally, these AI systems often require large amounts of data and significant time to learn and adapt to new environments, making their application in real-world industrial scenarios less flexible and efficient, hindering the true shift from "passive waiting for response" to "proactive optimization management." Summary of the Invention

[0004] The purpose of this invention is to propose a building energy-saving optimization method based on dynamic load forecasting and multi-device coupling, which solves the shortcomings of existing building energy-saving optimization methods in dealing with rapid changes in building load, complex coupling relationships between devices and physical safety constraints, as well as the problem of low efficiency of artificial intelligence systems when learning and adapting to new environments.

[0005] The method mentioned in this invention includes the following steps:

[0006] S1. Construct a digital twin environment that integrates a dynamic load forecasting model and a device coupling graph network model. The dynamic load forecasting model is used to output the predicted values ​​of the future cooling and heating loads of the building's central air conditioning system. The device coupling graph network model uses the physical devices of the central air conditioning system as nodes and the connection relationships between devices as edges, and dynamically updates the edge weights based on real-time monitored heat / mass transfer data to characterize the energy transfer efficiency of each level of the central air conditioning system.

[0007] S2. In the digital twin environment, a model-assisted reinforcement learning agent is run to explore the equipment operation strategy of the central air conditioning system. The action space of the agent is the control command of each device in the central air conditioning system. Before outputting the action, the physical feasibility of the candidate action is verified and the energy efficiency is estimated by using the device coupling graph network model.

[0008] S3. The intelligent agent makes decision optimization with the goal of minimizing the total energy consumption of the simulated system under comfort constraints. Its decision-making process combines the future load information provided by the dynamic load prediction model to carry out forward planning.

[0009] S4. Compile the optimization strategy generated by the intelligent agent into a sequence of device control instructions and apply it to the actual central air conditioning system;

[0010] S5. Collect actual operating data after the strategy is applied, compare it with the simulated data in the digital twin environment, start the model correction process based on the deviation of the comparison results, update the dynamic load prediction model or the equipment coupling graph network model, and feed the updated environment back to step S2 to form a closed-loop optimization.

[0011] A building energy efficiency optimization system is provided to implement the aforementioned building energy efficiency optimization method based on dynamic load forecasting and multi-device coupling. The system includes:

[0012] Data acquisition and edge control units are deployed on the construction site to acquire sensor data and execute control commands;

[0013] A cloud-based intelligent optimization platform is connected to the edge control unit via a network. The cloud-based intelligent optimization platform includes:

[0014] (a) A digital twin environment management module for maintaining and running the dynamic load forecasting model and the device coupling graph network model;

[0015] (b) A model-assisted reinforcement learning decision engine for performing policy exploration, physics verification and optimization decisions;

[0016] (c) Closed-loop calibration and evolution module, used to perform the model calibration process;

[0017] Secure communication and command verification channels are used to ensure reliable and secure transmission of data and commands between the cloud and the edge.

[0018] The beneficial effects provided by this invention are: by constructing a digital twin environment, this invention integrates a dynamic load prediction model and a device coupling graph network model, thereby achieving accurate prediction of the future cooling and heating load of the building's central air conditioning system and dynamic characterization of the energy transfer efficiency between devices.

[0019] Building upon this foundation, model-assisted reinforcement learning agents can explore equipment operation strategies within a digital twin environment. Through physical feasibility verification and energy efficiency prediction, the safety and efficiency of the generated strategies are ensured. The agent optimizes decisions by minimizing the total energy consumption of the simulated system while satisfying comfort constraints, and incorporates future load information for forward-looking planning. This effectively addresses the shortcomings of existing technologies, such as a lack of accurate prediction of future loads, insufficient inter-device collaborative optimization, and system operation risks.

[0020] Furthermore, through the closed-loop calibration process, the system can dynamically update the model based on the deviation between actual operating data and simulated data, thereby improving the system's adaptability and robustness.

[0021] In summary, the method of this application can significantly improve the operating efficiency of central air conditioning systems, reduce building energy consumption, and realize the transformation from "passive waiting for response" to "proactive optimization management," demonstrating significant energy-saving effects and practical application value. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0024] Before formally describing the present invention, a general description of the solution of the present invention will be given first to facilitate understanding.

[0025] Example 1

[0026] Please refer to Figure 1 The present invention provides a building energy-saving optimization method based on dynamic load forecasting and multi-device coupling, comprising the following steps:

[0027] S1. Construct a digital twin environment that integrates a dynamic load forecasting model and a device coupling graph network model. The dynamic load forecasting model is used to output the predicted values ​​of the future cooling and heating loads of the building's central air conditioning system. The device coupling graph network model uses the physical devices of the central air conditioning system as nodes and the connection relationships between devices as edges, and dynamically updates the edge weights based on real-time monitored heat / mass transfer data to characterize the energy transfer efficiency of each level of the central air conditioning system.

[0028] Specifically, in step S1 of building a digital twin environment:

[0029] The dynamic load forecasting model can be implemented using various machine learning or statistical methods. For example, it can employ an autoregressive integral moving average (ARIMA) model based on historical time series data or a simple multilayer perceptron (MLP) network to predict heating and cooling loads over a future period by learning the relationship between historical load data and environmental parameters. This model can provide preliminary load forecasts, offering fundamental information for subsequent strategy development.

[0030] The device coupling graph network model can be constructed as a static topology graph, where nodes represent the main physical equipment of the central air conditioning system, such as chillers, cooling towers, water pumps, and terminal devices, and edges represent the connections between these devices for energy or material flow. The initial edge weights can be set based on the equipment's design parameters or historical average operating data; for example, they can be simply represented as a fixed percentage of energy transfer efficiency. This model can initially reflect the connection structure between devices, providing a system topology view for agent action exploration.

[0031] S2. In the digital twin environment, a model-assisted reinforcement learning agent is run to explore the equipment operation strategy of the central air conditioning system. The action space of the agent is the control command of each device in the central air conditioning system. Before outputting the action, the physical feasibility of the candidate action is verified and the energy efficiency is estimated by using the device coupling graph network model.

[0032] Specifically, in step S2, where the model-assisted reinforcement learning agent explores policies:

[0033] The action space of the intelligent agent can be defined as a combination of control commands such as the on / off status, set temperature, flow rate, or frequency of various devices in the central air conditioning system. For example, the intelligent agent can output discrete or continuous control commands such as "chiller unit on / off", "cooling tower fan speed adjustment", and "pump frequency setting". These commands directly act on the device model in the digital twin environment, simulating its operational response.

[0034] Before outputting an action, the physical feasibility of candidate actions is verified using the device coupling graph network model. For example, it can be simply checked whether a candidate action would cause the output parameters (such as temperature, pressure, flow rate) of a device to exceed its preset upper and lower limits, or whether it would cause the load rate of a device to exceed its rated capacity. If so, it is determined to be infeasible. This verification can prevent the agent from exploring strategies that obviously violate the device's operating specifications.

[0035] Simultaneously, the energy efficiency of candidate actions is predicted using the aforementioned device coupling graph network model. For example, by inputting candidate actions into a digital twin environment and running a simplified physical simulator, the total energy consumption change of the system over a short period of time after executing the action can be quickly calculated. This prediction can provide the agent with preliminary energy efficiency feedback, guiding it to explore energy-saving directions.

[0036] S3. The intelligent agent makes decision optimization with the goal of minimizing the total energy consumption of the simulated system under comfort constraints. Its decision-making process combines the future load information provided by the dynamic load prediction model to carry out forward planning.

[0037] Specifically, in step S3, where the agent optimizes the decision by minimizing the total energy consumption of the simulated system while satisfying comfort constraints:

[0038] The agent's decision-making process can incorporate future load information provided by the dynamic load forecasting model for forward-looking planning. For example, the agent can adjust equipment operating modes in advance based on load forecasts for the next few hours to cope with upcoming load changes. This planning can employ rule-based heuristic search or a simple greedy algorithm to select the optimal equipment combination under the predicted load.

[0039] S4. Compile the optimization strategy generated by the intelligent agent into a sequence of device control instructions and apply it to the actual central air conditioning system;

[0040] S5. Collect actual operating data after the strategy is applied, compare it with the simulated data in the digital twin environment, start the model correction process based on the deviation of the comparison results, update the dynamic load prediction model or the equipment coupling graph network model, and feed the updated environment back to step S2 to form a closed-loop optimization.

[0041] Specifically, in step S5, which involves collecting actual operational data after the strategy is applied and performing model calibration:

[0042] The actual total energy consumption can be simply compared with the simulated total energy consumption. If the deviation exceeds a preset threshold, model calibration is initiated. The calibration process can involve fine-tuning the parameters of the dynamic load forecasting model or the equipment coupling graph network model, for example, by retraining or updating the model weights based on error backpropagation. This calibration allows the digital twin environment to gradually adapt to the operating characteristics of the actual system.

[0043] In some embodiments described above in this application, the dynamic load forecasting model is used to output predicted values ​​of the future cooling and heating loads of a building's central air conditioning system. However, in practical applications, load forecasting models for a single structure often struggle to simultaneously consider the stability of long-term load trends and the real-time accuracy of short-term load fluctuations. For example, building loads are influenced by a variety of complex factors, including weather changes, occupant activity patterns, and equipment operating status, which exhibit different characteristics at different time scales. If the load forecasting model cannot effectively capture these multi-scale characteristics, its forecast results may lack foresight in long-term planning and fail to meet high-precision requirements in short-term scheduling, thus affecting the effectiveness of the overall energy-saving optimization strategy.

[0044] In response, this application further proposes that the above-mentioned dynamic load forecasting model is a two-layer structure model based on the attention mechanism. The first layer is a long-cycle trend forecasting network that outputs the hourly load baseline for the next 24 hours; the second layer is a short-cycle rolling correction network that receives the output of the first layer and the latest environmental sequence, and outputs high-precision load forecast values ​​and their confidence intervals for the next 1-4 hours.

[0045] Specifically, the dynamic load forecasting model is designed as a two-layer structure model based on an attention mechanism. The first layer is a long-cycle trend forecasting network, whose main function is to capture the long-term variation patterns and periodic characteristics of building loads, such as daily, weekly, and seasonal load patterns. This network analyzes a large amount of historical operating data to output an hourly load baseline for the next 24 hours. This baseline represents the expected load level under normal operating conditions, providing a stable reference for subsequent short-term forecasts. As a preferred implementation, the long-cycle trend forecasting network can employ sequence models such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), or gated recurrent units (GRUs), and can be combined with convolutional neural networks (CNNs) to extract local features of the time series.

[0046] Furthermore, the second layer is a short-cycle rolling correction network, whose function is to make high-precision predictions of the load over a shorter time range based on the long-cycle baseline. This network receives the output of the first layer and the latest environmental sequence, such as real-time outdoor temperature, humidity, solar radiation intensity, and indoor occupancy density data, and performs rolling corrections. The short-cycle rolling correction network outputs high-precision load predictions for the next 1-4 hours and their confidence intervals. The confidence interval quantifies the uncertainty of the prediction results, helping the model-assisted reinforcement learning agent to assess risk during decision-making. As a specific implementation, the short-cycle rolling correction network can employ a Transformer model or an attention-enhanced LSTM / GRU model to effectively capture complex and nonlinear dependencies in the short term. The attention mechanism dynamically allocates weights to different input features (such as historical load, weather data, equipment status, etc.), allowing it to focus more on the information that has the greatest impact on the current prediction. The solution in this application effectively solves the limitations of a single model in terms of the accuracy of long- and short-cycle load prediction by introducing a two-layer dynamic load prediction model based on an attention mechanism.

[0047] In some preferred embodiments, this application is implemented as follows: Assume that the central air conditioning system of a commercial building needs energy-saving optimization. The dynamic load forecasting model first uses its first-layer long-cycle trend forecasting network to train and output the hourly cooling and heating load baseline for the next 24 hours based on historical load data, weather data, and holiday information from the past year. For example, the network may predict that the peak office hours are from 9:00 AM to 5:00 PM, during which the load will increase significantly, while the load will be lower at night. Subsequently, the second-layer short-cycle rolling correction network receives this 24-hour baseline and combines it with the latest environmental sequences such as outdoor temperature, indoor CO2 concentration, and personnel entry and exit data at the current moment (e.g., every 15 minutes). If a sudden increase in outdoor temperature or a surge in personnel density in the conference room area is detected in real time, the short-cycle rolling correction network will immediately adjust the load forecast value for the next 1-4 hours upward and provide a corresponding confidence interval. For example, it may adjust the cooling load forecast value for the next hour from 100kW to 120kW and give a confidence interval of ±5kW. This high-precision, short-term forecast information with confidence intervals will be used by a model-assisted reinforcement learning agent to adjust the operating parameters of equipment such as chillers, cooling towers, and water pumps in real time. For example, it can pre-start standby units or adjust the supply and return temperatures of chilled water to ensure the system operates with minimal energy consumption while maintaining comfort levels. Simultaneously, the attention mechanism ensures that during the correction process, the model prioritizes real-time factors that have the greatest impact on the current load. For example, it focuses more on outdoor temperature during hot summer months and indoor CO2 concentration during periods of high population density, thereby improving the accuracy of the forecasts and the model's response speed.

[0048] It should be noted that in the device coupling graph network model, the energy transfer efficiency represented by the edge weights is the dynamic energy efficiency transfer coefficient η, which is obtained through... To update, where η t Let λ be the coefficient at the current time, λ be the forgetting factor, k be the system characteristic constant, and ΔT be the coefficient. t To address the real-time temperature difference of the medium between connected devices, F t For real-time quality flow, P t This represents the real-time power consumption of upstream equipment.

[0049] Specifically, the dynamic energy transfer coefficient η t This refers to the efficiency of energy transfer between connected devices within a central air conditioning system at the current time t. The dynamic update mechanism of this coefficient aims to reflect the impact of equipment operating status and environmental changes on energy transfer efficiency in real time. Wherein, η... t-1 The energy efficiency transfer coefficient from the previous time step is used to incorporate historical information and maintain the smoothness of updates. The forgetting factor λ is a weighting coefficient between 0 and 1, which serves to balance the historical data η. t-1 The impact of current real-time monitoring data on energy efficiency coefficient updates. A larger λ value indicates a greater emphasis on historical trends, while a smaller λ value makes the update focus more on current real-time data. The system characteristic constant k is a calibration parameter related to the physical characteristics of a specific central air conditioning system, used to convert real-time monitored physical quantities into the dimension of energy transfer efficiency. ΔT t This refers to the real-time temperature difference of the medium between connected devices. For example, in a heat exchanger, it reflects the temperature difference between the refrigerant or water at the inlet and outlet, and is a key indicator for measuring heat transfer efficiency. t Real-time mass flow rate represents the mass of the medium flowing through the connected equipment per unit time, directly affecting the amount of energy delivered. t This represents the real-time power consumption of upstream equipment, such as the real-time electrical energy consumption of pumps or compressors, which signifies the cost incurred to achieve energy transfer. By substituting this real-time monitoring data into the formula, the energy transfer efficiency at the current moment can be calculated, thereby enabling dynamic and accurate updates to the edge weights of the device coupling graph network model.

[0050] This application's solution introduces a dynamic energy efficiency transfer coefficient η and its specific update formula, enabling precise and dynamic adjustment of edge weights in the device coupling graph network model based on real-time monitored heat / mass transfer data. This update mechanism incorporates historical information η. t-1 Compared with the current real-time running data (ΔT) t F t P t By weighting the edges using a forgetting factor λ, it is ensured that the edge weights can respond promptly to changes in the system's operating state, while avoiding oversensitivity caused by instantaneous fluctuations.

[0051] In some of the above embodiments, when a model-assisted reinforcement learning agent explores equipment operation strategies for a central air conditioning system in a digital twin environment, it is necessary to perform physical feasibility verification on candidate actions before outputting the action. Specifically, the physical feasibility verification in S2 aims to ensure that the strategy explored by the agent is safe and executable in the actual physical system.

[0052] The physical feasibility verification in S2 specifically involves: inputting the candidate action into the device coupling graph network model for forward propagation, simulating the evolution of the system state, and determining whether any of the following situations exist in the simulation results: (a) the load rate of any node device exceeds its safe operating range; (b) the simulated flow rate in any closed-loop pipeline does not satisfy the mass conservation law; (c) the simulated temperature difference of any heat exchange node violates the basic laws of heat transfer; if any of these exist, the candidate action is determined to be physically infeasible.

[0053] The forward propagation of candidate actions into the device coupling graph network model refers to simulating how the states (such as temperature, pressure, flow rate, power consumption, etc.) of each device within the central air conditioning system evolve over time after receiving a specific control command (i.e., candidate action). This forward propagation process aims to predict the instantaneous or short-term response of the system after executing the action. Specifically, simulating system state evolution involves calculating and updating the operating parameters and media state parameters of each node device (such as chillers, pumps, cooling towers, terminal devices, etc.) in the device coupling graph network model based on the control commands from the input actions to each node device, combined with the device's own operating characteristic curves and the coupling relationships between devices (represented by edge weights).

[0054] Determining whether any of the following scenarios exist in the simulation results is to evaluate the physical rationality and safety of candidate actions from multiple dimensions. Scenario (a) The load rate of any node device exceeds its safe operating range. This means checking whether the simulated equipment operating state would lead to equipment overload, no-load, or operation under non-design conditions, thereby avoiding equipment damage, sudden efficiency drops, or system instability. For example, the cooling capacity of chillers, the head and flow rate of water pumps, and the heat dissipation capacity of cooling towers all have their specific safe operating ranges. Scenario (b) The simulated flow rate in any closed-loop pipeline does not satisfy mass conservation. This means verifying whether the inflow and outflow mass flow rates are equal in any closed loop when simulating fluid circulation systems (such as chilled water systems and cooling water systems). This ensures the physical consistency of the fluid dynamics simulation and avoids spurious flow rate increases and decreases. Scenario (c) indicates that the simulated temperature difference at any heat exchange node violates the fundamental laws of heat transfer. This refers to checking whether the temperature difference between the inlet and outlet media of heat exchange equipment (such as evaporators, condensers, coils, etc.) conforms to the second law of thermodynamics, which states that heat always flows from a higher-temperature object to a lower-temperature object, and the temperature difference cannot be negative (in the normal heat transfer direction). This ensures the physical realism of the energy transfer simulation. If any of the above scenarios exist, the candidate action is deemed physically infeasible, meaning that the action would cause the system to enter an unsafe, unstable, or physically incompatible state, and therefore the action should not be actually executed.

[0055] This application's solution effectively addresses the problem of traditional reinforcement learning potentially generating physically infeasible or unsafe control strategies during the exploration process by introducing physical feasibility verification before the reinforcement learning agent outputs actions. Specifically, the device coupling graph network model, acting as a carrier of physical constraints, can quickly simulate the impact of candidate actions on the system state. By performing multi-dimensional physical law and safety boundary checks on the simulation results, including equipment load rate, mass conservation, and fundamental laws of heat transfer, unreasonable actions that may lead to equipment damage, system failure, or low operating efficiency can be identified and filtered out in a timely manner. This pre-verification mechanism ensures that the agent always operates within the physically feasible range during the exploration process, thereby avoiding the application of dangerous or ineffective strategies to the actual central air conditioning system and guaranteeing the safe and stable operation of the system.

[0056] Specifically, in the energy efficiency prediction step of S2 above, this application proposes a more accurate calculation method to guide reinforcement learning agents in exploring equipment operation strategies.

[0057] The energy efficiency prediction in S2 specifically involves: using the device coupled graph network model, calculating the energy efficiency change gradient of the system from the current state to the new stable state after executing the candidate action. The energy efficiency change gradient is obtained by calculating the Jacobian matrix of the input layer, i.e. the action vector, of the output layer of the graph network, i.e. the total power consumption of the system, and serving as a key indicator for evaluating the merits of the candidate action.

[0058] Specifically, energy efficiency prediction refers to the forward-looking energy efficiency assessment of each candidate action generated by a model-assisted reinforcement learning agent when exploring equipment operation strategies, in order to determine the potential energy consumption changes caused by that action. This prediction is crucial for the agent to select the optimal strategy. The device coupling graph network model is used as a high-efficiency simulator capable of simulating the state evolution and energy consumption performance of a central air conditioning system after executing specific control commands (i.e., candidate actions).

[0059] The energy efficiency change gradient can be understood as the sensitivity of the system's total energy consumption to the input action vector. When an agent proposes a candidate action, this action changes the operating state of various devices in the central air conditioning system, thereby affecting the energy transfer efficiency and total power consumption of the entire system. By calculating this gradient, the immediate impact and potential trend of the candidate action on the system's energy consumption can be quantitatively assessed.

[0060] Specifically, the energy efficiency change gradient is obtained by calculating the Jacobian matrix of the output layer to the input layer of the graph network. The output layer of the graph network refers to the total system power consumption estimated by the device-coupled graph network model after simulating the evolution of the system state. This total system power consumption estimate is a key indicator for measuring the overall energy efficiency of the system. The input layer corresponds to the action vector generated by the agent, which contains specific control commands for each device in the central air conditioning system. The calculation of the Jacobian matrix reveals how the total system power consumption estimate changes with the changes in the control commands in the action vector, thus providing a multi-dimensional energy efficiency sensitivity analysis. Therefore, this energy efficiency change gradient is used as a key indicator to evaluate the merits of candidate actions, guiding the agent to select actions with better energy efficiency. The solution in this application, by utilizing the device-coupled graph network model for energy efficiency prediction, can provide accurate energy consumption feedback for reinforcement learning agents.

[0061] In some implementations of the above methods, S3 involves an agent performing forward planning using future load information provided by a dynamic load forecasting model. However, in the complex operating environment of building central air conditioning systems, future load information is uncertain, and the equipment's operational space is vast. If traditional planning methods are used, they may face problems such as low computational efficiency, difficulty in fully exploring the optimal strategy path, and insufficient robustness to uncertainty, thereby affecting the real-time performance of decision-making and optimization results.

[0062] In response, this application further proposes that the prospective planning in S3 above adopts a Monte Carlo tree search based on a physical model, including:

[0063] The extended steps of tree search are as follows: using the device coupled graph network model as a fast simulator, the subsequent states of multiple candidate actions are simulated in parallel;

[0064] The simulation steps of tree search are as follows: A lightweight value network is used to evaluate the long-term energy efficiency potential of the leaf node state, which takes the future load sequence output by the dynamic load forecasting model as input features.

[0065] The backtracking steps of tree search are as follows: update the node statistics based on the simulation results, and guide the search to delve deeper into branches that are more energy efficient and physically feasible.

[0066] Specifically, Monte Carlo Tree Search (MCTS) based on a physical model is a heuristic search algorithm that evaluates the long-term value of actions by performing multiple simulations in the search tree. It is particularly suitable for decision-making problems with large state and action spaces. The expansion step of the tree search involves selecting one or more underexplored candidate actions at the current search node (representing the system state) according to a preset strategy to expand the search and generate new child nodes. During this process, the device coupling graph network model is used as a fast simulator, capable of efficiently simulating the evolution of the physical state of the central air conditioning system after executing these candidate actions, and simultaneously extrapolating the subsequent states of multiple candidate actions, thereby quickly assessing their short-term impact.

[0067] Furthermore, the tree search simulation step involves starting from the expanded leaf nodes and performing a series of random or policy-guided simulations until a preset simulation depth or termination condition is reached. During this process, a lightweight value network is used to evaluate the long-term energy efficiency potential of these leaf node states. This value network is designed to quickly estimate the future energy consumption performance of the system and uses the future load sequence output by the dynamic load forecasting model as input features, enabling the value assessment to fully consider future load change trends and thus provide a more forward-looking long-term energy efficiency assessment.

[0068] Furthermore, the backtracking step in tree search refers to the process of propagating simulation results (e.g., cumulative energy consumption, comfort satisfaction, etc.) backward from the leaf nodes along the search path to the root node after the simulation ends, and updating the statistics of all nodes along the way, such as the number of visits and average reward. By continuously updating these statistics, Monte Carlo tree search can gradually learn which action sequences are more likely to lead to better long-term energy efficiency, and guide the search to delve deeper into branches with higher energy efficiency and physical feasibility, avoiding getting trapped in local optima or exploring physically infeasible paths.

[0069] This application's solution effectively addresses the limitations of traditional planning methods in complex dynamic environments by introducing a physical model-based Monte Carlo tree search. Specifically, in the extended step of the tree search, a device-coupled graph network model is used as a fast simulator, which can efficiently and accurately simulate the physical consequences of multiple candidate actions. This not only accelerates the exploration of potential strategies but also ensures the physical feasibility of the inferred states. Through parallel inference, search efficiency is significantly improved, enabling the agent to evaluate more possibilities within a limited time.

[0070] In the tree search simulation step, a lightweight value network, combined with the future load sequence output by a dynamic load forecasting model, evaluates the long-term energy efficiency potential of leaf node states. This combination allows the agent to consider not only immediate energy consumption when making decisions, but also to proactively assess energy consumption performance over a future period, thus avoiding short-sighted decisions and ensuring the long-term optimization effect of the strategy. The design of the value network enables it to provide evaluations quickly without the need for time-consuming full simulations.

[0071] In the backtracking step of tree search, node statistics are updated based on simulation results, guiding the search towards more energy-efficient and physically feasible branches. This mechanism allows the agent to learn from each simulation, gradually focusing on the most promising decision paths, thereby efficiently finding the minimum total energy consumption strategy that satisfies comfort constraints within a vast action space. Through continuous iteration and optimization of the search process, the agent can generate more robust and efficient operating strategies.

[0072] In some preferred embodiments, assuming the central air conditioning system faces a predicted rapid increase in load during a certain period, the agent needs to decide how to adjust the operating parameters of the chillers, cooling towers, and pumps to cope. At this point, a Monte Carlo tree search based on a physics model will be initiated.

[0073] First, in the extended step of tree search, the agent generates multiple candidate actions based on the current system state, such as "increasing the cooling capacity of chiller unit A," "starting standby chiller unit B," and "adjusting the cooling tower fan speed." The device coupling graph network model, acting as a fast simulator, simulates in parallel the real-time impact of these candidate actions on the state of each device in the system, the temperature and flow rate of the medium, and energy consumption. For example, it simulates the changes in parameters such as chilled water outlet temperature, chilled water flow rate, and chiller unit power consumption after "increasing the cooling capacity of chiller unit A."

[0074] Next, in the tree search simulation step, for each expanded leaf node (i.e., the new system state after simulation), the lightweight value network receives the future load sequence output from the dynamic load forecasting model and quickly assesses the long-term energy efficiency potential of that state over the next few hours. For example, if a leaf node state has slightly higher current energy consumption but its configuration is better suited to handle future peak loads, the value network will assign it a higher long-term energy efficiency potential assessment.

[0075] Finally, in the backtracking step of the tree search, based on these simulation results and value assessments, the Monte Carlo tree search updates the visit count and average reward of all traversed nodes. For example, if the simulation results of a certain branch show that its long-term energy consumption is low and comfort constraints are consistently met, the node statistics of that branch are updated, giving it a higher exploration priority in subsequent searches. Through this iterative process, the agent eventually selects an optimal sequence of actions, such as "gradually increasing the cooling capacity of chiller unit A to 80% over the next 15 minutes, while simultaneously increasing the cooling tower fan speed to 70%", to minimize the total energy consumption while satisfying comfort constraints.

[0076] It should be noted that existing model calibration processes, when discovering discrepancies between actual operational data and simulated data in a digital twin environment, typically only determine the need to update the dynamic load forecasting model or the equipment coupling graph network model, but struggle to accurately identify the primary root cause of the discrepancy. Blindly adjusting the model without addressing this issue can lead to inefficient calibration or even introduce new errors, thereby affecting the performance and reliability of the entire building energy efficiency optimization method. To address this, this application proposes a model calibration process that, by performing quantitative attribution of the root causes of the discrepancy, intelligently determines whether the discrepancy primarily stems from inaccurate load forecasting or changes in equipment coupling relationships, and then prioritizes updating the corresponding models to achieve more accurate and efficient model calibration.

[0077] The model calibration process in S5 performs quantitative attribution of the root causes of bias, including:

[0078] (a) Calculate the actual total energy consumption E a With simulated total energy consumption E s The deviation ΔE;

[0079] (b) Calculate the gradient g of ΔE with respect to the parameters of the dynamic load prediction model. load And the gradient g of the edge weight parameters of the device coupled graph network model. coupling ;

[0080] (c) Comparing gradient norms ||g load ||and||g coupling ||, Where α is the proportionality coefficient, if the deviation is determined to be due to inaccurate load forecasting, the dynamic load forecasting model will be updated first; otherwise, the equipment coupling graph network model will be updated first.

[0081] Specifically, after the strategy is applied, the system will collect actual operating data and calculate the actual total energy consumption E. a Simultaneously, the digital twin environment will simulate the system's operating state and calculate the simulated total energy consumption E based on the same inputs and strategies. s The deviation ΔE is the actual total energy consumption E. a With simulated total energy consumption E s The difference between the two reflects the inconsistency between the digital twin model and the actual system. To quantify the source of the attribution bias ΔE, it is necessary to calculate the sensitivity of this bias to the parameters of the two core models. Specifically, the gradient g... load This represents the rate of change of the deviation ΔE relative to the parameters of the dynamic load prediction model, while the gradient g coupling This represents the rate of change of the deviation ΔE relative to the edge weight parameters of the device coupling graph network model. These gradients can be obtained through backpropagation algorithms or adjoint sensitivity analysis, and their magnitude reflects the degree of influence of the corresponding model parameters on the total energy consumption deviation. In practical applications, the proportionality coefficient... α This is an adjustable parameter used to balance the importance of the two gradient norms. When the gradient norm of the dynamic load forecasting model parameters ||g load ||Significantly larger than the gradient norm of the edge weight parameters in the device-coupled graph network model|| g coupling When ||g||, it indicates that the accuracy of load forecasting has a greater impact on the total energy consumption deviation, and in this case, the dynamic load forecasting model should be adjusted first. Conversely, if the gradient norm of the edge weight parameters of the equipment coupling graph network model is ||g||, it indicates that the accuracy of load forecasting has a greater impact on the total energy consumption deviation. coupling If the value is larger, it indicates that the change in energy transfer efficiency or coupling relationship between devices is the main reason, and the device coupling graph network model should be updated first.

[0082] This application's solution addresses the problem of accurately identifying the source of deviation in traditional model calibration by introducing a quantitative attribution mechanism for the root causes of deviations. Specifically, it calculates the actual total energy consumption E... a With simulated total energy consumption E s The deviation ΔE between them is calculated, and the gradient g of this deviation with respect to the parameters of the dynamic load prediction model is further calculated. load and the gradient g of the edge weight parameters of the coupled graph network model of the device coupling This allows for a quantitative assessment of the contribution of the two models to the total energy consumption deviation. Gradient norm ||g load ||and|| g coupling The magnitude of || directly reflects the influence of changes in the respective model parameters on the deviation of the system's total energy consumption. When ||g load||significantly greater than|| g coupling When ||g||, it means that the inaccuracy of the load forecasting model is the main factor causing system energy consumption deviation. Therefore, prioritizing the updating of the dynamic load forecasting model can more effectively reduce the deviation. Conversely, if ||g||, then the inaccuracy of the load forecasting model is the main factor causing system energy consumption deviation. coupling A larger || indicates a significant error in the representation of the energy transfer efficiency of the actual system by the device coupling graph network model. In this case, prioritizing the updating of the device coupling graph network model will be more efficient. This gradient-based quantitative attribution method makes the model calibration process no longer a blind attempt, but a targeted approach, thus significantly improving the efficiency and accuracy of model calibration.

[0083] In some preferred embodiments, it is assumed that after applying optimization strategies, the actual total energy consumption E of a building's central air conditioning system is reduced. a The total energy consumption of the digital twin environment simulation is 1000 kWh, while the total energy consumption E s If the load factor is 950 kWh, then the deviation ΔE is 50 kWh. To determine which model primarily accounts for this 50 kWh deviation, the system performs quantitative attribution. Specifically, through backpropagation, the gradient norm ||g| of the deviation ΔE with respect to the dynamic load forecasting model parameters is obtained. load The gradient norm of the edge weight parameters in the device-coupled graph network model is 0.8, while ||g coupling || is 0.2. Set the scaling factor α to 2. Now, compare ||g load ||(0.8) and α ||g coupling || (2 0.2 = 0.4). Since 0.8 > 0.4, that is, ||g load ||>α ||g coupling The system determines that the current deviation is primarily due to inaccurate load forecasting. Therefore, the model calibration process will prioritize updating the dynamic load forecasting model, for example, by retraining or fine-tuning its parameters to enable it to more accurately predict future heating and cooling loads. Conversely, if the calculation result is ||g load || is 0.3, while ||g coupling If || is 0.6, then 0.3 < 2 0.6 (1.2), i.e., ||g load ||<α ||g couplingIn this scenario, the system determines that the deviation primarily stems from inaccurate edge weight parameters in the device coupling graph network model, such as changes in energy transfer efficiency due to equipment aging or altered operating conditions. In this case, the model calibration process prioritizes updating the edge weight parameters of the device coupling graph network model to more accurately reflect the actual coupling relationships and energy transfer efficiency of the current equipment. This approach allows the system to specifically correct the root causes of the problem, avoiding ineffective adjustments to non-critical aspects of the model, thereby significantly improving the efficiency and accuracy of model calibration.

[0084] In some of the embodiments described above in this application, when the model calibration process determines that the deviation mainly originates from the device coupling graph network model, although the update direction is clear, directly adjusting all the edge weight parameters of the model may lead to problems such as high computational resource consumption, low update efficiency, and the introduction of unnecessary model oscillations. This is especially true in large and complex building central air conditioning systems, where the device coupling graph network model may contain a large number of edge weight parameters. To address this, this application further proposes an optimization scheme aimed at improving the efficiency and accuracy of model calibration: when determining whether to prioritize updating the device coupling graph network model, a sensitivity-driven sparse update is performed.

[0085] When prioritizing updating the device coupling graph network model, a sensitivity-driven sparse update is further performed: only for gradient g. coupling The edge weights of the top N absolute values ​​are adjusted in this round, while the weights of the rest remain unchanged. This focuses on quickly correcting the coupling links that have the greatest impact on the current deviation.

[0086] Specifically, the aforementioned sensitivity-driven sparse update refers to a selective model parameter adjustment strategy. Here, the gradient g... coupling It is the actual total energy consumption E a With simulated total energy consumption E s The gradient of the deviation ΔE with respect to the edge weight parameters of the device coupled graph network model, the magnitude of its absolute value reflects the degree of influence of each edge weight parameter on the total energy consumption deviation. A larger absolute value of the gradient indicates a greater contribution of the corresponding edge weight parameter to the current system energy consumption deviation. Furthermore, this round of adjustments only targets these gradients g. coupling The correction focuses on the N edge weights with the highest absolute values. Here, N is a preset integer that can be flexibly configured based on system complexity, computational resources, and the desired correction speed. For example, N can be set to 5% to 20% of the total number of edge weights, or a threshold can be dynamically determined based on the distribution of gradient values. The remaining edge weights that do not rank in the top N will retain their original values ​​in this round of correction. Therefore, this strategy concentrates the correction on the coupling links that most significantly affect the current system performance deviation, avoiding blindly adjusting all parameters, thus achieving fast and accurate model correction.

[0087] The scheme in this application uses gradient g coupling Analysis was conducted to quantify the contribution of each edge weight parameter in the device coupling graph network model to the deviation of the system's total energy consumption. It is precisely because these gradient values ​​accurately reflect the sensitivity of the parameters that the system can identify the coupling links that cause the most significant deviation between simulation and actual operation. By adjusting only these edge weight parameters with the greatest impact, rather than updating all parameters indiscriminately, the computational cost and time required for model calibration can be significantly reduced. Furthermore, this focused update strategy helps avoid unnecessary adjustments to parameters with smaller impacts on the deviation, thereby reducing the risk of model overfitting and improving the stability and convergence speed of model updates. As a result, the model can adapt to changes in the actual operating environment more quickly, maintaining the accuracy of its predictions and simulations.

[0088] Through the above technical solution, this application achieves more efficient and accurate parameter adjustment during model calibration, especially when the device coupling graph network model needs to be updated. Compared to adjusting all edge weight parameters, sensitivity-driven sparse updates significantly reduce the computational burden and shorten the model convergence time, enabling the digital twin environment to adapt to changes in actual building operation conditions more quickly. Furthermore, this correction method, which focuses on key influencing factors, effectively avoids over-adjustment of unimportant parameters, thereby improving the stability and robustness of model updates and ensuring the long-term effectiveness and reliability of the optimization strategy.

[0089] In some preferred embodiments, consider a central air conditioning system for a large commercial building whose equipment coupling graph network model contains hundreds of edge weight parameters. During the model calibration process in step S5, when the quantitative attribution results show that the deviation mainly stems from the equipment coupling graph network model, the system calculates the gradient g corresponding to each edge weight parameter. coupling For example, if N is set to 10, the system will identify the 10 gradients g with the largest absolute values. coupling The corresponding edge weight parameters are then selected. These parameters might correspond to the heat transfer efficiency of a heat exchanger, the head efficiency of a water pump, or the flow regulation coefficient of a valve. Subsequently, only these 10 key edge weight parameters are iteratively adjusted, for example, through gradient descent, to more accurately reflect the actual operating conditions. The remaining hundreds of edge weight parameters remain unchanged. In this way, the system can quickly locate and correct the core problems causing deviations, such as the performance degradation of a key piece of equipment or an anomaly in a pipeline connection, thereby rapidly improving the accuracy of the digital twin model without sacrificing the overall stability of the model.

[0090] Traditional model-assisted reinforcement learning agents, while capable of learning efficient equipment operation strategies through exploration in digital twin environments when optimizing building energy efficiency, often require significant time and computational resources for initial training or deployment in new building scenarios—exploring strategies from scratch. This "cold start" problem limits the rapid promotion and large-scale deployment of this method in practical applications, especially when facing diverse building types and dynamically changing operational requirements, where the agent's rapid adaptability is crucial. Failure to address these issues will result in long deployment cycles, high costs, and difficulty in quickly responding to changes in operating environments or system configurations.

[0091] In response, this application further proposes a policy knowledge distillation step, which aims to transform learned complex policies into knowledge forms that are easy to understand and apply, thereby accelerating the deployment and adaptation process of intelligent agents.

[0092] In some embodiments of this application, the method further includes a policy knowledge distillation step: periodically distilling the policies learned by the model-assisted reinforcement learning agent into an interpretable graph rule set, which is used to quickly initialize the agent in a new building scenario or as an auxiliary decision knowledge base. The form is: if the pattern matching of a specific subgraph of the device coupling graph satisfies the preset policy in the auxiliary decision knowledge base, then the device action combination is directly recommended.

[0093] Specifically, the policy knowledge distillation step refers to transforming the behavioral policies learned by a complex, often difficult-to-interpret, reinforcement learning agent into a simpler, more easily understood, and applicable knowledge representation through a certain technical means. In this application, this knowledge representation is an interpretable graph rule set. Periodically performing this distillation process means that during the agent's continuous learning and optimization, its latest and optimal policies are periodically extracted and transformed into rules to maintain the timeliness and effectiveness of the knowledge base.

[0094] The interpretable graph rule set can be understood as a series of condition-action pairs based on the structure and state of the device coupled graph network model. These rules take "If a specific subgraph of a particular device coupled graph matches a pattern" as the condition and "Then directly recommend a combination of device actions" as the result. For example, when the predicted cooling load value of a certain area in the device coupled graph network model is high and the inlet water temperature of the cooling tower reaches a certain threshold, the rule set may recommend starting a second chiller unit and adjusting the fan speed of the cooling tower.

[0095] In practical applications, the graph rule set is mainly used in two aspects: First, it is used to quickly initialize agents in new building scenarios. When a new building system needs to be deployed, there is no need to train the reinforcement learning agent from scratch. Instead, the existing graph rule set can be used as the agent's initial strategy or to guide its exploration direction, thereby significantly shortening the learning cycle. Second, it serves as an auxiliary decision-making knowledge base. When the agent makes decisions, this knowledge base can provide preset, validated optimization strategies. Especially when facing common or emergency situations, it can quickly provide reliable decision suggestions, improving the system's response speed and robustness.

[0096] The solution proposed in this application effectively addresses the efficiency issues that reinforcement learning agents may face during initial training or adaptation to new scenarios by introducing a policy knowledge distillation step.

[0097] Specifically, decision tree learning or rule extraction algorithms can be used to analyze the numerous decision trajectories of agents in a digital twin environment. For example, when the device coupled graph network model shows that the outdoor temperature is higher than 30°C and the predicted cooling load of indoor area A exceeds a set threshold, the agent will typically choose to start the second chiller and adjust the cooling water pump frequency to 70%. Through distillation, graph rules such as "If (outdoor temperature > 30°C and area A has a high cooling load) then (start the chiller, cooling water frequency = 70%)" can be extracted.

[0098] When this system is deployed in a new commercial complex, these distilled graph rule sets can serve as the initial strategies for its reinforcement learning agents. The agents no longer need to start exploring from random actions; instead, they can directly refer to these rules to make initial decisions and fine-tune and optimize within the digital twin environment. Furthermore, during daily operation, if the agent hesitates or malfunctions under certain conditions, the graph rule set can serve as an auxiliary decision-making knowledge base, providing validated, experience-based recommended action combinations to ensure system stability and energy efficiency. For example, when the system detects that a subgraph pattern (such as persistently high supply air temperature in a certain area, and the corresponding fan operating power reaching its limit) matches the "local overheating treatment" rule in the knowledge base, the system can directly recommend adjusting the airflow in adjacent areas or starting backup fans without waiting for the agent to recalculate.

[0099] Example 2:

[0100] This application proposes a building energy efficiency optimization system to implement the aforementioned building energy efficiency optimization method based on dynamic load forecasting and multi-device coupling. The system constructs a hierarchical distributed architecture, bringing data acquisition and edge control functions down to the building site, while centralizing complex intelligent optimization decision-making and model management functions in the cloud. Secure communication and command verification channels ensure the reliability of cloud-edge collaboration.

[0101] The system includes:

[0102] Data acquisition and edge control units are deployed on the construction site to acquire sensor data and execute control commands;

[0103] A cloud-based intelligent optimization platform is connected to the edge control unit via a network. The cloud-based intelligent optimization platform includes:

[0104] (a) A digital twin environment management module for maintaining and running the dynamic load forecasting model and the device coupling graph network model;

[0105] (b) A model-assisted reinforcement learning decision engine for performing policy exploration, physics verification and optimization decisions;

[0106] (c) Closed-loop calibration and evolution module, used to perform the model calibration process;

[0107] Secure communication and command verification channels are used to ensure reliable and secure transmission of data and commands between the cloud and the edge.

[0108] Specifically, the data acquisition and edge control unit is a collection of hardware and software deployed on-site. Its main function is to collect real-time sensor data from various sources within and outside the building, such as temperature, humidity, CO2 concentration, equipment operating status, and energy consumption. Simultaneously, this unit is responsible for receiving control commands from a cloud-based intelligent optimization platform and translating them into actions for the actual equipment, such as adjusting the start / stop of the chiller unit, pump speed, and valve opening in the central air conditioning system. This unit typically possesses a certain level of local processing capability to ensure real-time data acquisition and low latency in control command execution.

[0109] The cloud-based intelligent optimization platform serves as the core of the system, connecting to multiple data acquisition and edge control units via a network to achieve centralized and intelligent energy-saving optimization management of the entire building complex or large building. This platform integrates several key modules. The digital twin environment management module maintains and runs the dynamic load prediction model and the equipment coupling graph network model. This module continuously receives real-time data from the edge control units and inputs it into the dynamic load prediction model to generate predicted values ​​for the building's future heating and cooling loads. Simultaneously, this module dynamically updates the edge weights of the equipment coupling graph network model using real-time monitored heat / mass transfer data, ensuring that the digital twin environment accurately reflects the energy transfer efficiency and operating characteristics of the actual central air conditioning system.

[0110] Model-assisted reinforcement learning decision engines are a core component for achieving intelligent decision-making. This engine runs a reinforcement learning agent within a digital twin environment, exploring and learning optimal operating strategies for the central air conditioning system equipment through interaction with the digital twin. Before the agent outputs an action, the engine utilizes a device coupling graph network model to perform physical feasibility checks and energy efficiency predictions on candidate actions, ensuring that the generated strategy both conforms to physical constraints and achieves energy efficiency optimization.

[0111] The closed-loop calibration and evolution module is responsible for the system's continuous learning and adaptation capabilities. This module compares actual operational data with simulated data in the digital twin environment. Once a significant deviation is detected, the model calibration process is initiated. This process aims to identify the root causes of the deviation and update the dynamic load forecasting model or equipment coupling graph network model accordingly, thereby continuously improving the accuracy and predictive power of the digital twin environment. The updated environment is then fed back to the model-assisted reinforcement learning decision engine, forming a continuously optimizing closed loop.

[0112] The secure communication and command verification channel is a critical infrastructure element ensuring the stable operation of the entire system. This channel is responsible for establishing a reliable and secure communication link between the cloud-based intelligent optimization platform and the data acquisition and edge control unit, preventing data from being tampered with or stolen during transmission. Simultaneously, this channel also verifies transmitted control commands to prevent equipment malfunctions due to incorrect commands or malicious attacks, thereby ensuring the safe operation of building equipment.

[0113] This application's solution decouples the complex AI-driven building energy-saving optimization method into a cloud-based intelligent optimization platform and an edge data acquisition and control unit, enabling distributed, real-time deployment of the method. The data acquisition and edge control unit acquires data and executes commands in real time on-site, solving the problems of real-time data acquisition and low-latency control command execution. The cloud-based intelligent optimization platform centrally handles complex model calculations, strategy exploration, and optimization decisions. Utilizing its powerful computing capabilities, it ensures the accurate operation of the dynamic load forecasting model and the equipment coupling graph network model, as well as enabling the reinforcement learning agent to perform efficient strategy exploration and optimization. Secure communication and command verification channels guarantee data integrity and command security during cloud-edge collaboration, allowing the entire optimization method to operate continuously within a reliable and efficient architecture. This effectively addresses the challenges faced by traditional methods in practical deployment, such as real-time performance, security, computing resource allocation, and continuous model adaptability.

[0114] Through the above technical solution, this application provides a system architecture capable of implementing complex AI-driven building energy-saving optimization methods. This system, through a cloud-edge collaborative model, achieves real-time data acquisition and low-latency control command execution, while ensuring centralized and efficient processing of complex model calculations and intelligent decision-making. Therefore, the system ensures the accuracy and reliability of optimization strategies, significantly improves the overall energy-saving effect of buildings, and provides building managers with a stable and secure intelligent operation and management platform. Furthermore, the introduction of closed-loop correction and evolution modules enables the system to continuously learn and adapt to changes in the building environment, further enhancing the long-term effectiveness and robustness of energy-saving optimization.

[0115] In some preferred embodiments, a specific example is given below. Suppose a large commercial complex needs to implement a building energy efficiency optimization method based on dynamic load forecasting and multi-device coupling.

[0116] First, multiple data acquisition and edge control units are deployed in various areas of the commercial complex, such as office areas, shops, and public areas. These units are responsible for collecting real-time data on temperature, humidity, CO2 concentration, pedestrian traffic, and operating parameters of various devices in the central air conditioning system (such as the inlet and outlet water temperature, flow rate, and power consumption of chillers, pump speed, and valve opening). This data is then uploaded to the cloud-based intelligent optimization platform via secure communication and command verification channels. Simultaneously, these units are also responsible for receiving and executing control commands from the cloud platform, such as adjusting the start and stop of chillers, cooling tower fan speed, and pump frequency.

[0117] On the cloud-based intelligent optimization platform, the digital twin environment management module continuously receives and processes this real-time data, updates the dynamic load forecasting model to predict the heating and cooling loads for the next 24 hours or even longer, and dynamically adjusts the edge weights of the equipment coupling graph network model based on real-time heat / mass transfer data to accurately reflect the energy transfer efficiency between the various components of the central air conditioning system.

[0118] The model-assisted reinforcement learning decision engine utilizes this high-precision digital twin environment to run a reinforcement learning agent. This agent explores various equipment operation strategies in the simulated environment, such as proactively reducing the operating load of chiller units or optimizing the operation combination of pumps and fans when lower loads are predicted in the future. Before outputting actual control commands, the decision engine uses a device coupling graph network model to physically verify the feasibility of these candidate strategies, ensuring that there are no issues such as equipment overload, flow non-conservation, or abnormal temperature differences. It also performs energy efficiency predictions and selects the strategy with the lowest energy consumption that meets comfort constraints.

[0119] Subsequently, these optimized strategies are distributed to various data acquisition and edge control units in the form of a sequence of device control commands via secure communication and command verification channels. Upon receiving the commands, the edge control units immediately execute them to precisely control the relevant equipment in the central air conditioning system.

[0120] After the strategy is applied, the data acquisition and edge control unit continues to collect actual operational data and upload it to the cloud. The closed-loop correction and evolution module compares this actual operational data with the simulated data in the digital twin environment. If a significant deviation is found between the actual and simulated energy consumption, the module will initiate a model correction process. For example, it will use quantitative attribution to determine whether the load forecasting model is inaccurate or the equipment coupling model is inaccurate, and prioritize updating the corresponding model. The updated model will then be fed back to the model-assisted reinforcement learning decision engine, thus forming a continuously iterative and self-optimizing closed loop to ensure that the system can always provide the optimal energy-saving strategy.

[0121] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A building energy-saving optimization method based on dynamic load forecasting and multi-device coupling, characterized in that, include: S1. Construct a digital twin environment that integrates a dynamic load forecasting model and a device coupling graph network model. The dynamic load forecasting model is used to output the predicted values ​​of the future cooling and heating loads of the building's central air conditioning system. The device coupling graph network model uses the physical devices of the central air conditioning system as nodes and the connection relationships between devices as edges, and dynamically updates the edge weights based on real-time monitored heat / mass transfer data to characterize the energy transfer efficiency of each level of the central air conditioning system. S2. In the digital twin environment, a model-assisted reinforcement learning agent is run to explore the equipment operation strategy of the central air conditioning system. The action space of the agent is the control command of each device in the central air conditioning system. Before outputting the action, the physical feasibility of the candidate action is verified and the energy efficiency is estimated by using the device coupling graph network model. S3. The intelligent agent makes decision optimization with the goal of minimizing the total energy consumption of the simulated system under comfort constraints. Its decision-making process combines the future load information provided by the dynamic load prediction model to carry out forward planning. S4. Compile the optimization strategy generated by the intelligent agent into a sequence of device control instructions and apply it to the actual central air conditioning system; S5. Collect actual operating data after the strategy is applied, compare it with the simulated data in the digital twin environment, start the model correction process based on the deviation of the comparison results, update the dynamic load prediction model or the equipment coupling graph network model, and feed the updated environment back to step S2 to form a closed-loop optimization.

2. The building energy-saving optimization method based on dynamic load forecasting and multi-device coupling according to claim 1, characterized in that, The dynamic load forecasting model is a two-layer structure model based on the attention mechanism. Its first layer is a long-cycle trend forecasting network that outputs the hourly load baseline for the next 24 hours. The second layer is a short-cycle rolling correction network that receives the output of the first layer and the latest environmental sequence, and outputs high-precision load forecast values ​​and their confidence intervals for the next 1-4 hours.

3. The building energy-saving optimization method based on dynamic load forecasting and multi-equipment coupling according to claim 1, characterized in that, In the device coupling graph network model, the energy transfer efficiency represented by the edge weights is the dynamic energy efficiency transfer coefficient η, which is obtained through... To update, where η t Let λ be the coefficient at the current time, λ be the forgetting factor, k be the system characteristic constant, and ΔT be the coefficient. t To address the real-time temperature difference of the medium between connected devices, F t For real-time quality flow, P t This represents the real-time power consumption of upstream equipment.

4. The building energy-saving optimization method based on dynamic load forecasting and multi-device coupling according to claim 1, characterized in that, The physical feasibility verification in S2 specifically involves: inputting candidate actions into the device coupling graph network model for forward propagation, simulating the evolution of the system state, and determining whether any of the following situations exist in the simulation results: a) The load rate of any node device exceeds its safe operating range; b. Simulated flow in any closed-loop pipeline does not satisfy mass conservation; c. The simulated temperature difference at any heat exchange node violates the fundamental laws of heat transfer; if this exists, the candidate action is deemed physically infeasible.

5. The building energy-saving optimization method based on dynamic load forecasting and multi-equipment coupling according to claim 1, characterized in that, The energy efficiency prediction in S2 specifically involves: using the device coupled graph network model, calculating the energy efficiency change gradient of the system from the current state to the new stable state after executing the candidate action. The energy efficiency change gradient is obtained by calculating the Jacobian matrix of the input layer, i.e. the action vector, of the output layer of the graph network, i.e. the total power consumption of the system, and serving as a key indicator for evaluating the merits of the candidate action.

6. The building energy-saving optimization method based on dynamic load forecasting and multi-device coupling according to claim 1, characterized in that, The prospective planning in S3 employs a Monte Carlo tree search based on a physical model, which includes: The extended steps of tree search are as follows: using the device coupled graph network model as a fast simulator, the subsequent states of multiple candidate actions are simulated in parallel; The simulation steps of tree search are as follows: A lightweight value network is used to evaluate the long-term energy efficiency potential of the leaf node states, with the future load sequence output by the dynamic load forecasting model as input features. The backtracking steps of tree search are as follows: update the node statistics based on the simulation results, and guide the search to delve deeper into branches that are more energy efficient and physically feasible.

7. The building energy-saving optimization method based on dynamic load forecasting and multi-device coupling according to claim 1, characterized in that, The model calibration process in S5 performs quantitative attribution of the root causes of bias, including: a. Calculate the actual total energy consumption E a With simulated total energy consumption E s The deviation ΔE; b. Calculate the gradient g of ΔE with respect to the parameters of the dynamic load prediction model. load And the gradient g of the edge weight parameters of the device coupled graph network model. coupling ; c. Compare the gradient norm ||g load ||and||g coupling ||, if Where α is the proportionality coefficient, if the deviation is determined to be due to inaccurate load forecasting, the dynamic load forecasting model will be updated first; otherwise, the equipment coupling graph network model will be updated first.

8. The building energy-saving optimization method based on dynamic load forecasting and multi-device coupling as described in claim 7, characterized in that, When prioritizing updating the device coupling graph network model, a sensitivity-driven sparse update is further performed: only for gradient g. coupling The edge weights of the top N absolute values ​​are adjusted in this round, while the weights of the rest remain unchanged. This focuses on quickly correcting the coupling links that have the greatest impact on the current deviation.

9. A building energy-saving optimization method based on dynamic load forecasting and multi-device coupling according to claim 1, characterized in that, The method further includes a policy knowledge distillation step: periodically distilling the policies learned by the model-assisted reinforcement learning agent into an interpretable graph rule set, which is used to quickly initialize the agent in a new building scenario or as an auxiliary decision knowledge base. The form is: if the pattern matching of a specific subgraph of the device coupling graph satisfies the preset policy in the auxiliary decision knowledge base, then the device action combination is directly recommended.

10. A building energy-saving optimization system, used to implement the building energy-saving optimization method based on dynamic load forecasting and multi-device coupling as described in any one of claims 1-9, characterized in that, The system includes: Data acquisition and edge control units are deployed on the construction site to acquire sensor data and execute control commands; A cloud-based intelligent optimization platform is connected to the edge control unit via a network. The cloud-based intelligent optimization platform includes: a. A digital twin environment management module, used to maintain and run the dynamic load forecasting model and the equipment coupling graph network model; b. A model-assisted reinforcement learning decision engine for performing policy exploration, physical verification, and optimization decisions; c. Closed-loop calibration and evolution module, used to execute the model calibration process; Secure communication and command verification channels are used to ensure reliable and secure transmission of data and commands between the cloud and the edge.

Citation Information

Patent Citations

  • Air conditioning system control method and system based on reinforcement learning and twinborn model strategy

    CN121408831A

  • Air conditioning energy-saving simulation system

    WO2020124957A1