Large building group demand response regulation method, device and equipment and readable storage medium
By training and integrating multi-objective control models of HVAC and power systems, efficient and coordinated control strategies are generated, solving the problems of low accuracy and efficiency in demand response control of large building complexes, and achieving energy optimization and system efficiency improvement.
Patent Information
- Application Number
- CN202511856133.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-12-10
AI Technical Summary
Traditional technologies are difficult to effectively regulate the demand response of large building complexes, especially when faced with heterogeneous responses to different types of loads and dynamically changing system parameters, resulting in poor regulation accuracy and efficiency.
By training multiple candidate control models, the optimal multi-objective control models for the HVAC system and power system are selected based on the multi-objective demand information of large building complexes. These models are then integrated to generate efficient and coordinated target control strategies, which are then implemented in conjunction with current environmental and load information.
It has improved the demand response control effect of large building complexes, optimized energy consumption and improved system operating efficiency, and provided a stable, low-carbon and intelligent control solution.
Smart Images

Figure CN121329082B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart grid and building energy consumption optimization technology, and in particular to a method, device, computer equipment, storage medium and computer program product for demand response control of large building complexes. Background Technology
[0002] With the continuous growth of power system load and the accelerated pace of low-carbon transformation, demand response has become one of the key means to improve building energy efficiency, alleviate grid pressure, and achieve carbon emission reduction targets. As typical large-scale energy consumers, large building complexes have stable load characteristics, large response potential, and considerable flexibility in regulation, making them a key target for load-side management and intelligent regulation.
[0003] In traditional technologies, machine learning and deep learning algorithms struggle to comprehensively regulate the systemic response of large building complexes, given the heterogeneity of load responses to different types of loads and the dynamically changing system parameters. This results in poor accuracy and efficiency in regulating large building complexes. Therefore, effectively optimizing the demand response of large building complexes has become an urgent problem to be solved. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for demand response control of large building complexes, which can improve the control effect of demand response control of large building complexes, in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a demand response control method for large building complexes. The method includes:
[0006] Based on the multi-objective demand information of the large building complex, multiple first candidate control models for the HVAC system of the large building complex and multiple second candidate control models for the power system of the large building complex are trained.
[0007] Based on the first model performance evaluation results of each of the first candidate control models, a first multi-objective control model for the HVAC system is determined from the plurality of first candidate control models; and based on the second model performance evaluation results of each of the second candidate control models, a second multi-objective control model for the power system is determined from the plurality of second candidate control models.
[0008] The first multi-objective control model and the second multi-objective control model are integrated to obtain a multi-objective control integrated model for the large building complex.
[0009] The current environmental information of the HVAC system and the current load information of the power system are input into the multi-objective control integration model to obtain target control strategy information for the large building complex, so as to control the large building complex based on the target control strategy information.
[0010] Secondly, this application also provides a demand response control device for large building complexes. The device includes:
[0011] The model training module is used to train multiple first candidate control models for the HVAC system of the large building complex and multiple second candidate control models for the power system of the large building complex based on the multi-objective demand information of the large building complex.
[0012] The model screening module is used to determine a first multi-objective control model for the HVAC system from the plurality of first candidate control models based on the first model performance evaluation results of each first candidate control model, and to determine a second multi-objective control model for the power system from the plurality of second candidate control models based on the second model performance evaluation results of each second candidate control model.
[0013] The model integration module is used to integrate the first multi-objective control model and the second multi-objective control model to obtain a multi-objective control integrated model for the large building complex.
[0014] The strategy output module is used to input the current environmental information of the HVAC system and the current load information of the power system into the multi-objective control integration model to obtain target control strategy information for the large building complex, so as to control the large building complex based on the target control strategy information.
[0015] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0016] Based on the multi-objective demand information of the large building complex, multiple first candidate control models for the HVAC system of the large building complex and multiple second candidate control models for the power system of the large building complex are trained.
[0017] Based on the first model performance evaluation results of each of the first candidate control models, a first multi-objective control model for the HVAC system is determined from the plurality of first candidate control models; and based on the second model performance evaluation results of each of the second candidate control models, a second multi-objective control model for the power system is determined from the plurality of second candidate control models.
[0018] The first multi-objective control model and the second multi-objective control model are integrated to obtain a multi-objective control integrated model for the large building complex.
[0019] The current environmental information of the HVAC system and the current load information of the power system are input into the multi-objective control integration model to obtain target control strategy information for the large building complex, so as to control the large building complex based on the target control strategy information.
[0020] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0021] Based on the multi-objective demand information of the large building complex, multiple first candidate control models for the HVAC system of the large building complex and multiple second candidate control models for the power system of the large building complex are trained.
[0022] Based on the first model performance evaluation results of each of the first candidate control models, a first multi-objective control model for the HVAC system is determined from the plurality of first candidate control models; and based on the second model performance evaluation results of each of the second candidate control models, a second multi-objective control model for the power system is determined from the plurality of second candidate control models.
[0023] The first multi-objective control model and the second multi-objective control model are integrated to obtain a multi-objective control integrated model for the large building complex.
[0024] The current environmental information of the HVAC system and the current load information of the power system are input into the multi-objective control integration model to obtain target control strategy information for the large building complex, so as to control the large building complex based on the target control strategy information.
[0025] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0026] Based on the multi-objective demand information of the large building complex, multiple first candidate control models for the HVAC system of the large building complex and multiple second candidate control models for the power system of the large building complex are trained.
[0027] Based on the first model performance evaluation results of each of the first candidate control models, a first multi-objective control model for the HVAC system is determined from the plurality of first candidate control models; and based on the second model performance evaluation results of each of the second candidate control models, a second multi-objective control model for the power system is determined from the plurality of second candidate control models.
[0028] The first multi-objective control model and the second multi-objective control model are integrated to obtain a multi-objective control integrated model for the large building complex.
[0029] The current environmental information of the HVAC system and the current load information of the power system are input into the multi-objective control integration model to obtain target control strategy information for the large building complex, so as to control the large building complex based on the target control strategy information.
[0030] The aforementioned demand response control methods, devices, computer equipment, storage media, and computer program products for large building complexes train multiple candidate control models for HVAC systems and power systems based on multi-objective demand information of the large building complex. The optimal first and second multi-objective control models are selected based on model performance evaluation results, and then integrated to obtain a comprehensive multi-objective control integrated model. This model simultaneously processes environmental information from the HVAC system and load data from the power system, generating efficient, collaborative, and adaptive target control strategies. This improves the demand response control effect for large building complexes, optimizes energy consumption, enhances system operating efficiency, and enables multi-objective collaborative management of large building complexes. It provides a stable, low-carbon, and intelligent solution for large building complexes to participate in electricity market demand response. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is a diagram illustrating the application environment of a demand response control method for large building complexes in one embodiment.
[0033] Figure 2 This is a flowchart illustrating a demand response control method for large building complexes in one embodiment.
[0034] Figure 3 This is a flowchart illustrating the steps of training multiple first candidate control models for a heating, ventilation, and air conditioning system for a large building complex and training multiple second candidate control models for a power system for a large building complex in one embodiment.
[0035] Figure 4 This is a flowchart illustrating a demand response control method for large building complexes in another embodiment;
[0036] Figure 5 This is a flowchart illustrating the demand response control method for large building complexes in yet another embodiment.
[0037] Figure 6 This is a schematic diagram illustrating the comparison results of evaluation metrics in one embodiment; wherein, Figure 6 (a) in the figure is a schematic diagram of the comparison results of evaluation index 1; Figure 6 (b) in the figure is a schematic diagram of the comparison results of evaluation index 2; Figure 6 (c) in the figure is a schematic diagram of the comparison results of evaluation index 3; Figure 6 (d) in the figure is a schematic diagram of the comparison results of evaluation index 4;
[0038] Figure 7 This is a comparative schematic diagram of load changes in a large building complex over a continuous 72-hour period, as shown in one embodiment; wherein, Figure 7 (a) in the figure is a comparative diagram of the controllable load changes of a large building complex over a continuous 72-hour period; Figure 7 (b) in the figure is a comparative diagram of discrete HVAC load changes in a large building complex over a continuous 72-hour period;
[0039] Figure 8 This is a comparative diagram of the training curves of a reinforcement learning model in one embodiment; wherein, Figure 8 (a) in the figure is a comparative diagram of the training curves of the reinforcement learning model used to control discrete HVAC loads; Figure 8 (b) in the figure is a comparative diagram of the training curves of the reinforcement learning model used to control continuous controllable load;
[0040] Figure 9 This is a structural block diagram of a demand response control device for a large building complex in one embodiment.
[0041] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0043] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0044] The demand response control method for large building complexes provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, server 101 communicates with HVAC system 102 and power system 103 via a network. A data storage system can store the data that server 101 needs to process. The data storage system can be integrated onto server 101 or located on the cloud or other network servers. Server 101 trains multiple first candidate control models for the HVAC system 102 of the large building complex and multiple second candidate control models for the power system 103 of the large building complex based on the multi-objective demand information of the large building complex. Based on the first model performance evaluation results of each first candidate control model, a first multi-objective control model for the HVAC system 102 is determined from the multiple first candidate control models, and a second multi-objective control model for the power system 103 is determined from the multiple second candidate control models based on the second model performance evaluation results of each second candidate control model. The first and second multi-objective control models are integrated to obtain a multi-objective control integrated model for the large building complex. The current environmental information of the HVAC system 102 and the current load information of the power system 103 are input into the multi-objective control integrated model to obtain target control strategy information for the large building complex, and the large building complex is then controlled based on this target control strategy information. Server 101 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Heating, ventilation and air conditioning (HVAC) system 102 refers to the system or related equipment responsible for indoor heating, ventilation and air conditioning. Power system 103 refers to the overall system responsible for the generation, transmission, transformation, distribution and consumption of electrical energy.
[0045] In one embodiment, such as Figure 2 As shown, a demand response regulation method for large building complexes is provided, which can be applied to... Figure 1 Taking the server in the example, the following steps are included:
[0046] Step S201: Based on the multi-objective demand information of the large building complex, train multiple first candidate control models for the HVAC system of the large building complex and multiple second candidate control models for the power system of the large building complex.
[0047] Large building complexes refer to building groups that are large in scale and have complex functions. Large building complexes usually consist of multiple buildings.
[0048] Among them, the first candidate regulation model and the second candidate regulation model refer to multi-objective regulation models that are waiting for further screening.
[0049] Specifically, for distributed demand response scenarios of large building complexes, multiple demand objectives of the large building complex can be determined first, i.e., multi-objective demand information can be determined; then, multiple reinforcement learning models for the multi-objective demand information can be constructed. For the HVAC system and power system of the large building complex, multiple first candidate control models based on reinforcement learning for the HVAC system and multiple second candidate control models based on reinforcement learning for the power system can be trained respectively.
[0050] Step S202: Based on the first model performance evaluation results of each first candidate control model, a first multi-objective control model for the HVAC system is determined from multiple first candidate control models; and based on the second model performance evaluation results of each second candidate control model, a second multi-objective control model for the power system is determined from multiple second candidate control models.
[0051] Among them, the multi-objective control model refers to a model that performs coordinated optimization control for multiple demand objectives. The first multi-objective control model focuses on the control of HVAC system-related equipment, while the second multi-objective control model focuses on the control of power system-related equipment.
[0052] Specifically, the candidate control models with the highest performance evaluation scores (first or second) for the HVAC system and the power system are selected respectively. The first candidate control model with the highest performance evaluation score is labeled Q1, and the second candidate control model with the highest performance evaluation score is labeled Q2. Q1 and Q2 are used as the first multi-objective control model and the second multi-objective control model, respectively, to participate in the demand response of large building complexes.
[0053] In practical applications, the first multi-objective control model and / or the second multi-objective control model can be invoked to make parallel decisions on different types of loads based on the current environmental conditions, thereby achieving intelligent, collaborative, and efficient optimization of demand response for large building complexes.
[0054] Step S203: Integrate the first multi-objective control model and the second multi-objective control model to obtain an integrated multi-objective control model for large building complexes.
[0055] Specifically, the first multi-objective control model can output the optimal switching control command based on the input state, while the second multi-objective control model can output the optimal power adjustment ratio based on the input state. By integrating the first and second multi-objective control models, a reinforcement learning model for the HVAC control system and power system of a large building complex to participate in demand response is obtained, namely, the target integrated control model.
[0056] Step S204: Input the current environmental information of the HVAC system and the current load information of the power system into the multi-objective control integration model to obtain the target control strategy information for the large building complex, so as to control the large building complex based on the target control strategy information.
[0057] The current environmental information refers to information describing the environmental conditions of a large building complex. For example, current environmental information includes outdoor temperature, target indoor temperature, and other information.
[0058] Current load information refers to information describing the operating load of the power system. For example, current load information includes real-time electricity prices, controllable load, and uncontrollable load.
[0059] Specifically, the server collects real-time environmental information through the HVAC system and real-time load information through the power system. Then, it inputs the current environmental and load information into a multi-objective control integration model, which determines the optimal control strategy and outputs target control strategy information for the large building complex. Based on the target control strategy information, the large building complex is then controlled to obtain the control results.
[0060] In practical applications, the essence of a multi-objective regulation ensemble model based on reinforcement learning is to learn the optimal policy that maximizes the expected cumulative reward under a given environment and constraints. Therefore, the target regulation policy information output by the multi-objective regulation ensemble model can be represented by the following formula:
[0061]
[0062] In the formula, Indicates the optimal strategy; Indicating in strategy The following expectations; The discount factor at time t is represented. This represents the immediate reward at time t.
[0063] In the aforementioned demand response control method for large building complexes, multiple candidate control models for HVAC systems and power systems are trained based on the multi-objective demand information of the large building complex. The optimal first and second multi-objective control models are selected based on model performance evaluation results, and then integrated to obtain a comprehensive multi-objective control integrated model. This integrated model simultaneously processes environmental information from the HVAC system and load data from the power system, generating efficient, collaborative, and adaptive target control strategies. This improves the control effect of demand response for large building complexes, optimizes energy consumption, enhances system operating efficiency, and achieves multi-objective collaborative management of large building complexes. It provides a stable, low-carbon, and intelligent solution for large building complexes to participate in electricity market demand response.
[0064] In one embodiment, such as Figure 3 As shown, step S201 above, based on the multi-objective demand information of the large building complex, trains multiple first candidate control models for the HVAC system of the large building complex, and trains multiple second candidate control models for the power system of the large building complex, specifically including the following:
[0065] Step S301: Obtain initial control models based on different reinforcement learning algorithms.
[0066] Specifically, the server can build initial control models based on multiple reinforcement learning algorithms. Each initial control model is trained independently under the same physical parameters and experimental conditions. After the initial control models of each reinforcement learning algorithm are trained, their performance is compared using a unified model performance evaluation rule.
[0067] In practical applications, five reinforcement learning algorithms can be selected to construct the initial control model: PPO-D (Proximal Policy Optimization-Demonstrations), SACD (Soft Actor-Critic with Delayed Updates, where the Actor in the Actor-Critic architecture outputs deterministic actions, i.e., continuous values; and the Critic evaluates the value of state-action pairs, i.e., the Q-function), TD3 (TwinDelayed Deep Deterministic Policy Gradient), SAC (Soft Actor-Critic), and DDPG (Deep Deterministic Policy Gradient). Based on the action space type, these five reinforcement learning algorithms can be divided into discrete models (PPOD, SACD) and continuous models (TD3, SAC, DDPG).
[0068] (1) PPOD uses a policy network to directly output the probability distribution of discrete actions and uses a dominance function to evaluate the quality of actions. The core idea of PPOD is to limit the magnitude of each policy update (i.e., the pruning objective function) in policy gradient optimization to avoid excessive changes in the policy during a single update, thereby improving training stability. It features simple computation and stable convergence. The pruning objective function of PPOD can be expressed by the following formula:
[0069]
[0070] In the formula, E t θ represents the expectation over time step t, and θ represents the average of all sampled trajectories (PPO is a policy gradient algorithm based on sampling, which estimates the expectation through sampled trajectories); θ is the current policy network parameter. The strategy probability ratio, Indicates the current policy in state Select action The probability, Indicates the old strategy in state Take action below The probability, if >1 indicates that the current strategy is more inclined to select action. Otherwise, it means that the current strategy does not favor selecting an action. ; This is the estimate of the advantage function, used to measure the action. The relative value of an action How much better is it than the average performance? This indicates a pruning operation, where ε is the pruning threshold, used to limit the magnitude of policy updates.
[0071] (2) SACD mainly adopts soft Q-network and policy network structures, and uses reparameterization techniques to improve gradient calculation efficiency. The core idea of SACD is to introduce an entropy term as regularization while maximizing the reward, encouraging the policy to maintain more exploratory behavior in the early stages of training. SACD has strong exploration ability in discrete action space and fast convergence speed. The entropy regularization objective function of SACD is expressed by the following formula:
[0072]
[0073] In the formula, This represents the state-action value estimation of a pre-Q network with parameter θ. This represents the state-action value estimation of the target Q-network; This represents the immediate reward given by the environmental simulation model after action a is performed in state s; This is a discount factor, representing the degree of decay in future rewards, used to balance long-term and immediate returns. The larger the group, the more emphasis is placed on long-term rewards; Indicates the next discrete action; Indicates the policy network in state The probability distribution of actions is shown below; 'a' represents the weighting coefficient; 'D' represents the empirical replay buffer. This represents the expectation of the state s and action a sampled from the experience replay buffer.
[0074] (3) TD3 consists of two independent value networks Q and a deterministic policy network, and uses an experience replay buffer. The core idea of TD3 is to introduce a dual Q network on top of DDPG to reduce the overestimation bias of Q values, while using delayed policy updates and target policy smoothing to improve training stability. TD3 can effectively avoid overestimation and improve the accuracy and stability of continuous control tasks. The objective function of TD3 is shown below:
[0075]
[0076] In the formula, This indicates that the i-th value network (TD3 has two independent value networks) and ; Represents the target policy network In state The deterministic optimal action output is given by β, which represents smoothed noise. y is the target value of the value network, where y represents the expected cumulative reward (composed of immediate reward + discounted future value) that can be obtained after performing action a in state s.
[0077] (4) SAC comprises a soft Q-network, a value function network, and a policy network, and uses a stochastic policy to output continuous actions. Its core idea is to combine a maximum entropy reinforcement learning framework to simultaneously optimize the expected reward and entropy value of the policy, thereby obtaining a policy that is both high-performance and robust. SAC has strong convergence performance and good generalization ability, making it suitable for continuous control tasks in dynamic environments. The objective function of SAC is shown below:
[0078]
[0079] In the formula, For strategy The joint distribution of generated states and actions; Representation Strategy Generated state-action distribution lower expectations; It is in state Next action Instant rewards; It is the entropy coefficient; The policy entropy is used to represent the uncertainty of action distribution; This represents the probability distribution of actions.
[0080] (5) DDPG directly outputs deterministic actions from the Actor network, while the Critic network evaluates the value of state-action pairs. It utilizes experience replay and target network stabilization training. DDPG combines a deep neural network with the deterministic policy gradient method to achieve end-to-end learning in a high-dimensional continuous action space. It has a simple structure and low computational cost. The objective function of DDPG is shown below:
[0081]
[0082] In the formula, For policy network parameters The gradient; The deterministic action output by the Actor network in a given state; For Criti network state-action value estimation; The gradient of the value function with respect to action a is used to measure the degree to which changes in action affect the value. For the policy network, parameters The gradient is used to measure the degree to which changes in parameters affect the action output.
[0083] Step S302: Based on the multi-objective demand information, obtain the reward function of the initial regulation model.
[0084] Multi-objective demand information refers to multiple optimization objectives set for the demand of large building complexes. This multi-objective demand information includes electricity cost information, negative penalties for participating in demand response, reward information for participating in demand response, and carbon emission penalty information.
[0085] The reward function is used to calculate the immediate reward obtained by the agent in the reinforcement learning-based model (such as the initial control model) after taking a certain action in the environment. The reward function is used to guide the agent in the initial control model to learn the optimal policy.
[0086] Specifically, to address the actual needs of large building complexes, a multi-objective reward function is designed that comprehensively considers electricity costs, carbon emissions, user comfort, and peak-shaving load. Furthermore, by configuring multi-objective weight parameters, a flexible balance between different optimization objectives is achieved. The server can then use the reward function to guide the agent in the initial control model to perform the learning process based on the current state, environment, and actions.
[0087] In practical applications, the server can first calculate electricity cost information, negative penalty information for participating in demand response, reward information for participating in demand response, and carbon emission penalty information; then, it can construct a reward function based on these information. The reward function can be expressed by the following formula:
[0088]
[0089] In the formula, Let i represent the reward function for the i-th building; This represents the electricity cost information for the i-th building; It is a punishment for user dissatisfaction caused by participating in demand response, that is, negative punishment information for participating in demand response; It refers to the rewards for participating in peak shaving demand response in the electricity market, i.e., information on rewards for participating in demand response. It concerns penalties for carbon emissions, specifically information about carbon emission penalties. This indicates the weight corresponding to the electricity cost information; This indicates the weight corresponding to the negative penalty information for participating in demand response; This indicates the weight corresponding to the reward information for participating in demand response; This indicates the weight corresponding to carbon emission penalty information.
[0090] Electricity cost information can be represented by the following formula:
[0091]
[0092] In the formula, Let be the electricity price at time t; Let be the controllable load of the i-th building at time t; Let be the uncontrollable load of the i-th building at time t.
[0093] Negative penalties for participating in demand response can be represented by the following formula:
[0094]
[0095] In the formula, Let be the indoor temperature of the i-th building in the large building complex at time t; Let represent the target indoor temperature of the i-th building at time t.
[0096] The reward information for participating in demand response can be represented by the following formula:
[0097]
[0098] In the formula, P represents the peak electricity consumption period; Let be the rated power of the HVAC system of the i-th building. Different power levels reflect different speeds of regulating indoor temperature. This represents the on / off state of the HVAC system in building i at time t, and is a binary decision variable. A value of 0 indicates that the function is off. A value of 1 indicates that the application is open. This represents the total energy consumption of the baseline model at time t. It can be expressed by the following formula:
[0099]
[0100] Carbon emission penalty information can be expressed using the following formula:
[0101]
[0102] In the formula, It is a carbon emission factor.
[0103] Step S303: Based on the preset temperature constraints, reward function, and first state space information of the HVAC system of the large building complex, the initial control model is trained in the environmental simulation model to obtain multiple first candidate control models for the HVAC system; the first candidate control models are used to output switching control information for the HVAC system based on the input first state space information.
[0104] It should be noted that the temperature constraints of the initial control model are the same as those of the baseline control model.
[0105] Among them, environmental simulation models refer to virtual simulation models used to test and train the behavior of intelligent agents. Environmental simulation models can realistically reflect various dynamic characteristics of large building complexes, such as energy consumption, temperature changes, demand response strategies, and carbon emissions.
[0106] Specifically, based on preset temperature constraints, a reward function, and the first state-space information of the HVAC system of a large building complex, each initial control model is trained in an environmental simulation model. During training, the agent continuously operates within the environmental simulation model. Each time the agent selects an action, the model provides a reward or penalty based on the action's effect. The agent dynamically adjusts its control strategy based on the reward feedback from the model. The reward function comprehensively considers multiple objectives, including electricity costs, carbon emissions, and user comfort. Through continuous interaction with the environmental simulation model, the agent constantly adjusts and optimizes its strategy network parameters, enabling it to flexibly respond to electricity market demands under complex scenarios such as dynamic electricity prices and load fluctuations. This achieves intelligent, multi-objective optimal control of the overall energy consumption of the large building complex, ultimately converging to the optimal control strategy information that participates in demand response and balances economy, low carbon emissions, and comfort. It should be noted that before training the initial control models for the HVAC system and the power system respectively, all physical parameters involved (such as multi-source heterogeneous data) and weight parameters in the multi-objective reward function were uniformly set to fixed values. By standardizing the physical parameters, it was ensured that various reinforcement algorithms were compared under the same experimental conditions, which improved the reproducibility and fairness of the experimental results and provided a reliable foundation for the model performance evaluation in subsequent steps.
[0107] State space information refers to the set of all environmental information that an agent can observe at each decision step. A well-designed state space helps the agent comprehensively perceive environmental changes and make better decisions. The state space of current reinforcement learning models may include indoor temperature, outdoor temperature, target temperature, current load, electricity price, carbon emission factor, and time-series data features to comprehensively reflect the operational status and external environment of large building complexes.
[0108] Action space information refers to the set of all possible actions that an agent can choose at each decision step. The design of the action space directly affects the flexibility and precision of the control strategy.
[0109] In practical applications, the first-state spatial information of HVAC systems Including: the wholesale electricity price at the next moment The current time t and its characteristic value t*, and the load participating in the demand response at the next time step. Carbon emission rate Peak electricity consumption times (1 represents peak electricity consumption period, 0 represents off-peak electricity consumption period) and weekday indicator variables (1 represents a working day, 0 represents a non-working day). First state space information This can be expressed by the following formula:
[0110]
[0111] First action space information of HVAC system The state of the HVAC system is represented by binary variables. 0 indicates the HVAC system is off, and 1 indicates it is on. At each decision-making moment, the agent selects the optimal action based on the current state, achieving flexible start-up and shutdown control of the HVAC system. .
[0112] Step S304: Based on the preset temperature constraints, reward function, and second state space information of the power system of the large building complex, the initial control model is trained in the environmental simulation model to obtain multiple second candidate control models for the power system; the second candidate control models are used to output power regulation information for the power system based on the input second state space information.
[0113] Specifically, the training process for multiple second candidate control models of the power system is the same as the training process for multiple first candidate control models of the HVAC system, as detailed in step S303 above, and will not be repeated here.
[0114] In practical applications, the second state-space information of the power system Including: wholesale electricity prices at the next moment The current time t and its characteristic value t*, and the continuous load that can participate in demand response at the next time. Carbon emission rate Peak electricity consumption periods and working day indicator variable This can be expressed by the following formula:
[0115]
[0116] Continuously controllable load operating space : A continuous variable, representing the power adjustment ratio of the controllable load. At each decision moment, the agent selects the optimal adjustment range based on the current state, achieving flexible control over the continuously controllable load, i.e. .
[0117] In this embodiment, by introducing an integrated reinforcement learning framework, the optimal reinforcement learning algorithms are trained and integrated for the response characteristics of HVAC systems and continuously controllable loads, respectively, thereby improving the adaptability and control accuracy of the control strategy. This leads to the construction of an efficient, collaborative, and adaptive intelligent distributed demand response strategy for building multi-source heterogeneous data environments, providing a stable, low-carbon, and intelligent solution for large building complexes to participate in electricity market demand response.
[0118] In one embodiment, before determining the first multi-objective control model for the HVAC system from multiple first candidate control models based on the first model performance evaluation results of each first candidate control model in step S202, the method further includes: obtaining the first performance evaluation result of the first candidate control model based on the first multi-objective operation results of the first candidate control model participating in demand response, and the second multi-objective operation results of the benchmark control model of the HVAC system under benchmark operating conditions; obtaining the second performance evaluation result of the first candidate control model based on the convergence speed and stability of the reward function of the first candidate control model; and obtaining the first model performance evaluation result of each first candidate control model based on the first performance evaluation result and the second performance evaluation result of each first candidate control model.
[0119] The baseline operating condition refers to the normal operating conditions when not participating in demand response.
[0120] Specifically, the server calculates the difference between the first candidate control model's first multi-objective operation result in demand response and the second multi-objective operation result of the HVAC system's benchmark control model under benchmark conditions, i.e., the change of the first candidate control model relative to the benchmark control model, to obtain the first multi-objective operation result of the first candidate control model in demand response. The server can also obtain the second performance evaluation result of the first candidate control model based on the convergence speed and stability of the reward function of the first candidate control model. Based on the first weight corresponding to the first performance evaluation result and the second weight corresponding to the second performance evaluation result, the server performs a weighted summation of the first performance evaluation result and the second performance evaluation result of each first candidate control model to calculate the first model performance evaluation result of each first candidate control model.
[0121] In practical applications, the calculation process of the performance evaluation results of the first model can be represented by the following formula:
[0122]
[0123] In the formula, This indicates the performance evaluation results of the first model; This indicates the results of the first performance evaluation; This indicates the results of the second performance evaluation; Indicates the first weight; This indicates the second weight.
[0124] The first performance evaluation result can be assessed based on multiple objective requirements such as total electricity cost, user dissatisfaction, peak shaving effect, and carbon emissions. The calculation process of the first performance evaluation result can be expressed by the following formula:
[0125]
[0126] In the formula, Δcost represents the difference between the user cost information of the first candidate control model and the user cost information of the benchmark control model; Δdissatisfaction represents the difference between the negative penalty information of the first candidate control model participating in demand response and the negative penalty information of the benchmark control model not participating in demand response; Δpeak shaved represents the difference between the reward information of the first candidate control model participating in demand response and the reward information of the benchmark control model not participating in demand response; and Δcarbon represents the difference between the carbon emission penalty information of the first candidate control model and the carbon emission penalty information of the benchmark control model. The weighting coefficients for each sub-item in the calculation formula for the first performance evaluation result.
[0127] The calculation process for the second performance evaluation result can be expressed by the following formula:
[0128]
[0129] In the formula, Speed represents the convergence speed; Stability represents the stability; η1 and η2 are the weight coefficients of each sub-item in the calculation formula of the second performance evaluation result.
[0130] In this embodiment, a first performance evaluation result is obtained by comprehensively comparing the multi-objective operation results of the first candidate control model after participating in demand response with the performance of the benchmark control model under benchmark operating conditions. Simultaneously, a second performance evaluation result is obtained by combining the convergence speed and stability analysis of the reward function. These results are then integrated to form a comprehensive and objective first model performance evaluation result. This method effectively balances actual operational benefits with the stability of the training process, and can scientifically select control models that perform well under multiple objective constraints such as energy efficiency, comfort, and response speed, significantly improving the reliability and overall performance of intelligent control of HVAC systems in large building complexes.
[0131] In one embodiment, before determining the second multi-objective control model for the power system from multiple second candidate control models based on the second model performance evaluation results of each second candidate control model in step S202, the method further includes: obtaining a third performance evaluation result of the second candidate control model based on the change between the third multi-objective operation result of the second candidate control model participating in demand response and the second multi-objective operation result; obtaining a fourth performance evaluation result of the second candidate control model based on the convergence speed and stability of the reward function of the second candidate control model; and obtaining a second model performance evaluation result of each second candidate control model based on the third and fourth performance evaluation results of each second candidate control model.
[0132] Specifically, the server calculates the deviation between the third multi-objective operating result of the second candidate control model participating in demand response and the second multi-objective operating result of the benchmark control model of the HVAC system under the benchmark operating condition, i.e., the change of the second candidate control model relative to the benchmark control model, to obtain the third multi-objective operating result of the second candidate control model participating in demand response; the server can also obtain the fourth performance evaluation result of the second candidate control model based on the convergence speed and stability of the reward function of the second candidate control model; based on the third weight corresponding to the third performance evaluation result and the fourth weight corresponding to the fourth performance evaluation result, the third performance evaluation result and the fourth performance evaluation result of each second candidate control model are weighted and summed to calculate the second model performance evaluation result of each second candidate control model.
[0133] It should be noted that the evaluation method for the second model performance evaluation results of each second candidate regulation model is the same as that for the first model performance evaluation results of each first candidate regulation model, as detailed in the above embodiments, and will not be repeated here.
[0134] In this embodiment, by analyzing the changes in the multi-objective operating results of the second candidate control model after participating in demand response compared to the baseline operating condition, a third performance evaluation result is obtained. At the same time, by combining the convergence speed and stability of the reward function during model training, a fourth performance evaluation result is obtained. Thus, a comprehensive and objective second model performance evaluation result is generated. This can effectively quantify the actual benefits of the demand response strategy and the stability of model training, providing a scientific basis for selecting power system control models with excellent performance under multi-objective constraints such as energy efficiency, load regulation, and operational reliability. This significantly improves the accuracy and comprehensive performance of intelligent control of power systems in large building complexes.
[0135] In one embodiment, before determining the first multi-objective control model for the HVAC system from multiple first candidate control models based on the first model performance evaluation results of each first candidate control model in step S202, the method further includes: constructing a benchmark control model for the demand response of the HVAC system that does not participate in multi-objective demand information based on the multi-source heterogeneous data of the large building complex; the benchmark control model is used to make the indoor temperature of the large building complex close to the target indoor temperature.
[0136] The target indoor temperature refers to the desired indoor temperature, i.e., the ideal indoor temperature.
[0137] Specifically, the server first collects and processes high-quality target multi-source heterogeneous data and target multi-source heterogeneous features, including: (1) Data collection and integration: The server can collect multi-dimensional raw data such as outdoor temperature, target indoor temperature, real-time electricity price, HVAC load, controllable load, and uncontrollable load of large building groups through multi-source systems such as HVAC system and power system, and then the server obtains multi-source heterogeneous data of large building groups. Since the data types of multi-source heterogeneous data are quite different, the multi-source heterogeneous data can be time-aligned and integrated to form aligned multi-source heterogeneous data in a unified format. (2) Data cleaning and missing value processing: The outliers, duplicates and missing values in the aligned multi-source heterogeneous data are processed. For example, interpolation, forward filling and other methods can be used to fill in the missing data in the aligned multi-source heterogeneous data to ensure the continuity and integrity of the data, and then the server obtains cleaned multi-source heterogeneous data. (3) Resampling and Feature Engineering: The cleaned multi-source heterogeneous data is resampled to unify the cleaned multi-source heterogeneous data with different sampling frequencies into a 5-minute interval, thus obtaining time-series multi-source heterogeneous data. Then, feature engineering is performed on the time-series multi-source heterogeneous data to obtain multi-source heterogeneous features. (4) Target Variable Generation: Based on the operating rules of the large building complex, the daily target indoor temperature curve (such as the temperature settings during working hours and non-working hours) is generated and used as the target variable for subsequent model training. (5) The time-series multi-source heterogeneous data and multi-source heterogeneous features are normalized to obtain target multi-source heterogeneous data and multi-target source heterogeneous features, which are saved in a standard format to provide a data foundation for subsequent benchmark model construction and reinforcement learning algorithm training.
[0138] Furthermore, based on the target multi-source heterogeneous data of large building complexes, a baseline control model for the HVAC system that does not participate in demand response is constructed. This baseline control model aims to keep the indoor temperature as close as possible to the target indoor temperature while ensuring user comfort. Specifically, it includes: (1) Parameter setting: setting parameters such as the building's heat capacity C, thermal resistance R, and the rated power h of the HVAC system, and setting the simulation step size T according to the actual operating cycle (for example, when T=5min, the number of operating steps per day is 288). (2) Using a first-order thermodynamic model to describe the change process of indoor temperature, the temperature dynamic equation of the baseline control model is constructed. The temperature dynamic equation is shown below:
[0139]
[0140] In the formula, Let be the outdoor temperature of the i-th building at time t; Δt is the time step. Let be the heat capacity of the i-th building, which represents the building's heat storage capacity; Let be the thermal resistance of the i-th building, a parameter that reflects the building's insulation performance.
[0141] (3) Determining the objective function: The benchmark control model optimizes the operation of the HVAC system to maintain the indoor temperature as close as possible to the target indoor temperature while ensuring user comfort. The benchmark control model has minimizing the temperature difference between the indoor temperature and the target indoor temperature as its sole optimization objective. Therefore, the objective function of the benchmark control model can be expressed by the following formula:
[0142]
[0143] (4) Constraints: Constraints are divided into initial temperature constraints, dynamic temperature constraints and other constraints.
[0144] 1) Initial temperature constraint:
[0145] 2) Dynamic temperature constraints:
[0146] In the formula, This indicates the maximum permissible temperature difference between the interior and exterior temperatures of a large building complex.
[0147] 3) Other constraints: All types of loads within the large building complex operate at optimal power; temperature and load power are both greater than zero.
[0148] In this embodiment, by collecting and processing high-quality target multi-source heterogeneous data, a benchmark control model for demand response of the HVAC system that does not participate in multi-objective demand information is constructed. This benchmark control model can effectively utilize multi-source heterogeneous data to achieve precise control of the indoor temperature of large building complexes under benchmark operating conditions, making it stably close to the target temperature. This ensures indoor environmental comfort while providing a reliable control benchmark and comparison basis for subsequent multi-objective collaborative optimization, thereby improving the overall energy efficiency and control stability of large building complexes.
[0149] In one embodiment, step S204 above, which involves regulating a large building complex based on target regulation strategy information, specifically includes the following: generating regulation instructions corresponding to the target regulation strategy information; regulating the large building complex according to the regulation instructions to obtain the regulation result of the large building complex; and evaluating the regulation effect of the target regulation strategy information based on the regulation result and the benchmark regulation result of the benchmark regulation model to obtain a visualized regulation effect of the target regulation strategy information.
[0150] Specifically, in actual operation, the system collects environmental status and load information in real time, calls the multi-objective control integration model to make optimal decisions on the controllable loads of the HVAC system and the power system, and thus outputs target control strategy information for large building groups; dynamically generates control instructions corresponding to the target control strategy information, and uses the control instructions to control the large building groups to obtain the control results of the large building groups.
[0151] By analyzing the multi-objective operation results of large building clusters after information regulation based on target regulation strategies, including indoor temperature changes, load response, electricity costs, and carbon emissions, and comparing them with the multi-objective operation results of the benchmark regulation model under benchmark operating conditions, the performance of the regulation strategy is comprehensively evaluated, and the regulation effect is intuitively displayed using visualization methods.
[0152] In this embodiment, specific control instructions corresponding to the target control strategy information are generated, and precise control is implemented on the large building complex based on these instructions. Then, the actual control results are compared and analyzed with the expected results of the benchmark control model to complete a comprehensive evaluation of the target control strategy. Finally, the control effect is presented intuitively in a visual form, realizing closed-loop management from strategy generation to effect verification. This not only effectively optimizes the system operation performance of the large building complex and improves energy utilization efficiency and response accuracy, but also provides reliable data support and intuitive analysis basis for the continuous improvement of system control strategies and intelligent decision-making.
[0153] In one embodiment, such as Figure 4 As shown, another demand response regulation method for large building complexes is provided, which can be applied to... Figure 1Taking the server in the example, the following steps are included:
[0154] Step S401: Based on the multi-objective demand information of the large building complex, train multiple first candidate control models for the HVAC system of the large building complex, and train multiple second candidate control models for the power system of the large building complex.
[0155] Step S402: Based on the multi-source heterogeneous data of the large building complex, a benchmark control model for demand response of the HVAC system that does not participate in multi-objective demand information is constructed; the benchmark control model is used to make the indoor temperature of the large building complex close to the target indoor temperature.
[0156] Step S403: Based on the first multi-objective operation results of the first candidate control model participating in demand response and the second multi-objective operation results of the benchmark control model of the HVAC system under benchmark operating conditions, the first performance evaluation result of the first candidate control model is obtained.
[0157] Step S404: Based on the convergence speed and stability of the reward function of the first candidate regulation model, obtain the second performance evaluation result of the first candidate regulation model.
[0158] Step S405: Based on the first performance evaluation results and the second performance evaluation results of each first candidate control model, obtain the first model performance evaluation results of each first candidate control model.
[0159] Step S406: Based on the change between the third multi-objective operation result of the second candidate control model participating in demand response and the second multi-objective operation result, obtain the third performance evaluation result of the second candidate control model.
[0160] Step S407: Based on the convergence speed and stability of the reward function of the second candidate regulation model, obtain the fourth performance evaluation result of the second candidate regulation model.
[0161] Step S408: Based on the third and fourth performance evaluation results of each second candidate regulation model, obtain the second model performance evaluation results of each second candidate regulation model.
[0162] Step S409: Based on the first model performance evaluation results of each first candidate control model, determine a first multi-objective control model for the HVAC system from multiple first candidate control models; and based on the second model performance evaluation results of each second candidate control model, determine a second multi-objective control model for the power system from multiple second candidate control models.
[0163] Step S410: Integrate the first multi-objective control model and the second multi-objective control model to obtain an integrated multi-objective control model for large building complexes.
[0164] Step S411: Input the current environmental information of the HVAC system and the current load information of the power system into the multi-objective control integration model to obtain the target control strategy information for the large building complex, so as to control the large building complex based on the target control strategy information.
[0165] The aforementioned demand response control method for large building complexes can achieve the following beneficial effects: By training multiple candidate control models for HVAC systems and power systems based on multi-objective demand information of large building complexes, and selecting the optimal first and second multi-objective control models based on model performance evaluation results, a comprehensive multi-objective control integrated model is obtained. This integrated model simultaneously processes environmental information of HVAC systems and load data of power systems, generating efficient, collaborative, and adaptive target control strategy information. This achieves optimization of energy consumption, improvement of system operating efficiency, and multi-objective collaborative management of large building complexes, providing a stable, low-carbon, and intelligent solution for large building complexes to participate in electricity market demand response.
[0166] To more clearly illustrate the demand response control method for large building complexes provided in this disclosure, a specific embodiment will be used to describe the method in detail below. For example... Figure 5 The diagram illustrates yet another demand response control method for large building complexes, which can be applied to... Figure 1 The server in the context specifically includes the following:
[0167] (1) Data preprocessing: Collect multi-source heterogeneous data of large building complexes, including outdoor temperature, target indoor temperature, real-time electricity price, controllable load and uncontrollable load. Perform data cleaning, interpolation and resampling to generate high-quality input datasets. At the same time, feature extraction is performed on time data to enhance its expressive power and provide reliable data support for subsequent model training.
[0168] (2) Baseline Model Construction: A baseline control model for the HVAC system is constructed based on the building's physical parameters to determine the building's energy consumption curve under normal operating conditions. The sole objective of this model is to maintain the indoor temperature as close to the set value as possible, ensuring user comfort without employing any demand response strategy. The generated baseline results are used for subsequent reinforcement learning optimization algorithms to compare and validate indoor temperature changes and load response.
[0169] (3) Reinforcement learning algorithm training: Based on distributed demand response, a multi-objective optimization objective function and constraints are constructed. Multi-objective weight parameters are introduced into the objective function to balance multiple objectives such as electricity cost, carbon emissions, and user comfort. Optimal single reinforcement learning models are trained for the response characteristics of HVAC systems and continuously controllable loads respectively. During the training process, each model adjusts the multi-objective weight parameters to achieve flexible trade-offs between different optimization objectives, ultimately obtaining an independent demand response strategy that takes into account multiple objectives.
[0170] (4) Optimal Algorithm Selection and Integration: The performance of different reinforcement learning algorithms is compared, and the optimal algorithms for HVAC systems and continuous controllable loads are obtained and integrated. A multi-objective optimized integrated reinforcement learning demand response control strategy is constructed to further improve the adaptability, practicality and control accuracy of the control strategy.
[0171] (5) Intelligent Regulation and Result Analysis: The integrated optimal algorithm is used to intelligently regulate the building complex and generate regulation results. The performance of the regulation strategy is verified by analyzing indicators such as indoor temperature changes, load response, electricity costs, and carbon emissions, and the regulation effect is displayed through visualization. This method provides a scientific decision-making basis for the subsequent participation of large building complexes in market-based demand response.
[0172] (6) Effect Verification: Based on real large-scale building electricity consumption data, including load type, temperature, electricity price, and other data. Load types include continuously controllable loads (such as lighting, some power equipment), discretely controllable loads (such as HVAC systems), and uncontrollable loads (such as safety and fire protection systems, information and communication systems, etc.). Loads not involved in demand response are used as baseline loads.
[0173] Table 1 Physical Parameter Settings
[0174]
[0175] Table 2 Weight Parameter Settings
[0176]
[0177] Table 3 Reinforcement Learning Model Parameter Settings
[0178]
[0179] Table 4 Model Training and Simulation Settings
[0180]
[0181] This study compares the demand response regulation effects of PPOD, SACD, TD3, SAC, DDPG, and their combined models through experiments. The evaluation indicators are: total electricity cost over 72 hours (indicator 1); total user dissatisfaction over 72 hours (indicator 2); peak shaving effect during peak electricity consumption over 72 hours (indicator 3); and total carbon emission reduction over 72 hours (indicator 4). The comparison results of each model under these four indicators are as follows: Figure 6 As shown. Figure 7 As shown, the study also compared discrete HVAC load variations and continuous controllable load variations within a large building complex over a continuous 72-hour period. Figure 8 As shown, the training curves of various reinforcement learning models used to control different types of loads (including discrete HVAC loads and continuously controllable loads) are also compared.
[0182] In this embodiment, multiple candidate control models for HVAC systems and power systems are trained based on multi-objective demand information of large building complexes. The optimal first and second multi-objective control models are selected based on model performance evaluation results, and then integrated to obtain a comprehensive multi-objective control integrated model. This model simultaneously processes environmental information of HVAC systems and load data of power systems, generating efficient, collaborative, and adaptive target control strategy information. This achieves optimization of energy consumption, improvement of system operating efficiency, and multi-objective collaborative management of large building complexes, providing a stable, low-carbon, and intelligent solution for large building complexes to participate in electricity market demand response.
[0183] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0184] Based on the same inventive concept, this application also provides a large-scale building complex demand response control device for implementing the aforementioned large-scale building complex demand response control method. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of one or more embodiments of the large-scale building complex demand response control device provided below can be found in the limitations of the large-scale building complex demand response control method described above, and will not be repeated here.
[0185] In one embodiment, such as Figure 9 As shown, a demand response control device 900 for large building complexes is provided, comprising: a model training module 901, a model screening module 902, a model integration module 903, and a strategy output module 904, wherein:
[0186] The model training module 901 is used to train multiple first candidate control models for the HVAC system of large building complexes and multiple second candidate control models for the power system of large building complexes based on the multi-objective demand information of large building complexes.
[0187] The model screening module 902 is used to determine a first multi-objective control model for the HVAC system from multiple first candidate control models based on the first model performance evaluation results of each first candidate control model, and to determine a second multi-objective control model for the power system from multiple second candidate control models based on the second model performance evaluation results of each second candidate control model.
[0188] The model integration module 903 is used to integrate the first multi-objective control model and the second multi-objective control model to obtain a multi-objective control integrated model for large building complexes.
[0189] The strategy output module 904 is used to input the current environmental information of the HVAC system and the current load information of the power system into the multi-objective control integration model to obtain the target control strategy information for the large building complex, so as to control the large building complex based on the target control strategy information.
[0190] In one embodiment, the model training module 901 is further configured to: acquire initial control models based on different reinforcement learning algorithms; obtain the reward function of the initial control models based on multi-objective demand information; perform model training processing on each initial control model in an environmental simulation model based on preset temperature constraints, the reward function, and the first state space information of the HVAC system of the large building complex, thereby obtaining multiple first candidate control models for the HVAC system; the first candidate control models are used to output switching control information for the HVAC system based on the input first state space information; and perform model training processing on each initial control model in an environmental simulation model based on preset temperature constraints, the reward function, and the second state space information of the power system of the large building complex, thereby obtaining multiple second candidate control models for the power system; the second candidate control models are used to output power regulation information for the power system based on the input second state space information.
[0191] In one embodiment, the large building complex demand response control device 900 further includes a first performance evaluation module, used to obtain a first performance evaluation result of the first candidate control model based on the first multi-objective operation result of the first candidate control model participating in demand response and the second multi-objective operation result of the benchmark control model of the HVAC system under benchmark operating conditions; to obtain a second performance evaluation result of the first candidate control model based on the convergence speed and stability of the reward function of the first candidate control model; and to obtain a first model performance evaluation result of each first candidate control model based on the first performance evaluation result and the second performance evaluation result of each first candidate control model.
[0192] In one embodiment, the large-scale building complex demand response control device 900 further includes a second performance evaluation module, used to obtain a third performance evaluation result of the second candidate control model based on the change between the third multi-objective operation result of the second candidate control model participating in demand response and the second multi-objective operation result; to obtain a fourth performance evaluation result of the second candidate control model based on the convergence speed and stability of the reward function of the second candidate control model; and to obtain a second model performance evaluation result of each second candidate control model based on the third and fourth performance evaluation results of each second candidate control model.
[0193] In one embodiment, the demand response control device 900 for large building complexes further includes a benchmark model construction module, which is used to construct a benchmark control model for demand response of the HVAC system that does not participate in multi-objective demand information based on the multi-source heterogeneous data of the large building complex; the benchmark control model is used to make the indoor temperature of the large building complex close to the target indoor temperature.
[0194] In one embodiment, the strategy output module 904 is further configured to generate control instructions corresponding to the target control strategy information; perform control processing on the large building complex according to the control instructions to obtain the control results of the large building complex; and perform control effect evaluation processing on the target control strategy information based on the control results and the benchmark control results of the benchmark control model to obtain the visualized control effect of the target control strategy information.
[0195] Each module in the aforementioned demand response control device for large building complexes can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0196] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores multi-objective demand information, current environmental information, and other data. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a demand response control method for large-scale building complexes.
[0197] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0198] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0199] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0200] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0201] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0202] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0203] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A demand response control method for large building complexes, characterized in that, The method includes: Obtain initial control models based on different reinforcement learning algorithms; Based on the multi-objective demand information, the reward function of the initial regulation model is obtained; Based on the preset temperature constraints, the reward function, and the first state space information of the HVAC system of the large building complex, each of the initial control models is trained in the environmental simulation model to obtain multiple first candidate control models for the HVAC system; the first candidate control models are used to output switching control information for the HVAC system based on the input first state space information. Based on the preset temperature constraints, the reward function, and the second state space information of the power system of the large building complex, each of the initial control models is trained in the environmental simulation model to obtain multiple second candidate control models for the power system; the second candidate control models are used to output power regulation information for the power system based on the input second state space information. Based on the first multi-objective operation result of the first candidate control model participating in demand response and the second multi-objective operation result of the benchmark control model of the HVAC system under benchmark operating conditions, the first performance evaluation result of the first candidate control model is obtained. Based on the convergence speed and stability of the reward function of the first candidate regulation model, the second performance evaluation result of the first candidate regulation model is obtained. Based on the first performance evaluation results and the second performance evaluation results of each first candidate regulation model, the first model performance evaluation results of each first candidate regulation model are obtained. Based on the first model performance evaluation results of each of the first candidate control models, a first multi-objective control model for the HVAC system is determined from the plurality of first candidate control models; and based on the second model performance evaluation results of each of the second candidate control models, a second multi-objective control model for the power system is determined from the plurality of second candidate control models. The first multi-objective control model and the second multi-objective control model are integrated to obtain a multi-objective control integrated model for the large building complex. The current environmental information of the HVAC system and the current load information of the power system are input into the multi-objective control integration model to obtain target control strategy information for the large building complex, so as to control the large building complex based on the target control strategy information.
2. The method according to claim 1, characterized in that, Before determining a second multi-objective control model for the power system from the plurality of second candidate control models based on the second model performance evaluation results of each of the second candidate control models, the method further includes: The third performance evaluation result of the second candidate control model is obtained based on the change between the third multi-objective operation result of the second candidate control model participating in demand response and the second multi-objective operation result. Based on the convergence speed and stability of the reward function of the second candidate regulation model, the fourth performance evaluation result of the second candidate regulation model is obtained. Based on the third and fourth performance evaluation results of each of the second candidate regulation models, the second model performance evaluation results of each of the second candidate regulation models are obtained.
3. The method according to claim 2, characterized in that, Before determining the first multi-objective control model for the HVAC system from the plurality of first candidate control models based on the first model performance evaluation results of each of the first candidate control models, the method further includes: Based on the multi-source heterogeneous data of the large building complex, a benchmark control model for the demand response of the HVAC system that does not participate in the multi-objective demand information is constructed. The benchmark control model is used to bring the indoor temperature of the large building complex close to the target indoor temperature.
4. The method according to claim 3, characterized in that, The process of regulating the large building complex based on the target regulation strategy information includes: Generate control instructions corresponding to the target control strategy information; According to the control instructions, the large building complex is controlled to obtain the control result of the large building complex. Based on the control results and the benchmark control results of the benchmark control model, the control effect of the target control strategy information is evaluated to obtain a visualized control effect of the target control strategy information.
5. A demand response control device for large building complexes, characterized in that, The device includes: The model training module is used to acquire initial control models based on different reinforcement learning algorithms; obtain the reward function of the initial control model according to multi-objective demand information; and perform model training processing on each of the initial control models in an environmental simulation model according to preset temperature constraints, the reward function, and the first state space information of the HVAC system of the large building complex, to obtain multiple first candidate control models for the HVAC system. The first candidate control models are used to output switching control information for the HVAC system according to the input first state space information. Based on preset temperature constraints, the reward function, and the second state space information of the power system of the large building complex, the initial control models are also trained on the environmental simulation model to obtain multiple second candidate control models for the power system. The second candidate control models are used to output power regulation information for the power system according to the input second state space information. The first performance evaluation module is used to obtain a first performance evaluation result of the first candidate control model based on the first multi-objective operation result of the first candidate control model participating in demand response and the second multi-objective operation result of the benchmark control model of the HVAC system under benchmark operating conditions; to obtain a second performance evaluation result of the first candidate control model based on the convergence speed and stability of the reward function of the first candidate control model; and to obtain a first model performance evaluation result of each first candidate control model based on the first performance evaluation result and the second performance evaluation result of each first candidate control model. The model screening module is used to determine a first multi-objective control model for the HVAC system from the plurality of first candidate control models based on the first model performance evaluation results of each first candidate control model, and to determine a second multi-objective control model for the power system from the plurality of second candidate control models based on the second model performance evaluation results of each second candidate control model. The model integration module is used to integrate the first multi-objective control model and the second multi-objective control model to obtain a multi-objective control integrated model for the large building complex. The strategy output module is used to input the current environmental information of the HVAC system and the current load information of the power system into the multi-objective control integration model to obtain target control strategy information for the large building complex, so as to control the large building complex based on the target control strategy information.
6. The apparatus according to claim 5, characterized in that, The device further includes a second performance evaluation module, used to obtain a third performance evaluation result of the second candidate control model based on the change between the third multi-objective operation result of the second candidate control model participating in demand response and the second multi-objective operation result. Based on the convergence speed and stability of the reward function of the second candidate regulation model, the fourth performance evaluation result of the second candidate regulation model is obtained. Based on the third and fourth performance evaluation results of each of the second candidate regulation models, the second model performance evaluation results of each of the second candidate regulation models are obtained.
7. The apparatus according to claim 6, characterized in that, The device also includes a benchmark model construction module, which is used to construct a benchmark control model of the HVAC system that does not participate in the demand response of the multi-objective demand information based on the multi-source heterogeneous data of the large building complex. The benchmark control model is used to bring the indoor temperature of the large building complex close to the target indoor temperature.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Air conditioning system flexible adjustment potential quantification method and device, terminal and storage medium
CN120027509A
Regional building group multi-energy complementary cooperative regulation and control method
CN120782161A