Micro-grid-heating ventilation air conditioner coordinated optimization method based on deep reinforcement learning

By employing a two-layer optimization control strategy based on deep reinforcement learning and an improved PER-DDPG algorithm, the coupling relationship between HVAC systems and microgrids under uncertain conditions was resolved, achieving optimization of energy costs and improvement of comfort, thereby enhancing the system's adaptability and intelligence.

CN120975528AInactive Publication Date: 2025-11-18NANJING NORMAL UNIVERSITY

Patent Information

Application Number
CN202511506855.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2025-11-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively address the close bidirectional coupling between HVAC systems and microgrids under uncertain conditions, especially when dealing with high-dimensional, continuous, and heterogeneous operating spaces, making it difficult to achieve dynamic characteristic optimization of the HVAC thermal inertia and microgrid power balance requirements.

Method used

A two-layer optimization control strategy based on deep reinforcement learning is adopted, which combines enhanced model predictive control and Markov decision process. The agent is trained by an improved Priority Experience Playback Deep Deterministic Policy Gradient Algorithm (PER-DDPG) to achieve coordinated optimization between the HVAC system and the microgrid.

Benefits of technology

While ensuring comfort, we aim to reduce the energy costs of building systems, improve the adaptability and intelligence of system behavior, avoid strategies getting trapped in local optima, and enhance model convergence speed and sample utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975528A_ABST
    Figure CN120975528A_ABST
Patent Text Reader

Abstract

The invention discloses a micro-grid-heating ventilation air-conditioning coordinated optimization method based on deep reinforcement learning, and relates to the field of building energy system management and scheduling, and the method comprises the following steps: S1, constructing a refined model of a heating ventilation air-conditioning system and a micro-grid system, which considers the comfort level of people and the building energy consumption cost; s2, aiming at the refined model constructed in the step S1, adopting a double-layer optimization control strategy to construct a micro-grid energy management system framework integrated with the heating ventilation air conditioning system; s3, based on the step S2, constructing a multi-energy coupling model of the micro-grid-heating ventilation air conditioning system, converting a time sequence optimization problem into a Markov decision process MDP, and performing mathematical representation; and S4, training the intelligent agent by adopting an improved priority experience playback depth deterministic policy gradient algorithm PER-DDPG. According to the method, a collaborative scheduling optimization framework of the micro-grid and the heating ventilation air conditioning system is constructed, and the energy cost of a building system can be effectively reduced under the condition that the comfort level of personnel is effectively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of building energy system management and scheduling, in particular to a micro-grid-heating ventilation air conditioning coordinated optimization method based on deep reinforcement learning. BACKGROUND

[0002] With the large-scale access of distributed renewable energy, the power supply and demand sides show deep integration characteristics, and the power grid load increases year by year, which brings challenges to the safe and stable operation of the power grid. The distributed renewable energy consumption project is based on the load center to use the distributed renewable energy immediately, which reduces the pressure on the power grid brought by uncertain renewable energy generation resources, and improves the stability of the power grid operation.

[0003] With the low-carbon and intelligent transformation of building systems, building energy consumption is increasing year by year. As a major component of building energy consumption, the heating ventilation air conditioning (HVAC) system contributes up to 70% of the total building energy consumption load, and has high adjustable potential, which has become the focus of energy-saving control research. Micro-grid driven by renewable energy has become a key technical solution to alleviate the building load pressure and reduce the power grid consumption pressure.

[0004] This feature requires the building energy management system to coordinate and maintain the optimization goal of indoor thermal comfort and HVAC power cost. However, the nonlinear and complex dynamic characteristics of the HVAC system and the micro-grid bring great challenges to the design and optimization of the control strategy.

[0005] Existing methods are often limited to single-layer optimization or limited to the application of deep reinforcement learning in a single field, and are difficult to effectively cope with the close two-way coupling relationship between fine-grained HVAC dynamic characteristics and micro-grid operating conditions under uncertain conditions. In addition, there are two key problems in practical application: (1) the high-dimensional, continuous and heterogeneous action space across the control layer needs to be handled simultaneously; (2) it is difficult to capture the strong multi-time scale coupling dynamic characteristics between HVAC thermal inertia and micro-grid power balance requirements. SUMMARY

[0006] The purpose of the present application is to provide a micro-grid-heating ventilation air conditioning coordinated optimization method based on deep reinforcement learning to solve the problems raised in the background. The method integrates the device-level driving of the HVAC system and the system-level energy scheduling decision of the micro-grid. It can effectively reduce the energy cost of the building system while ensuring personnel comfort.

[0007] To achieve the above purpose, the present application provides the following technical solutions, including the following steps: S1, a fine model of the heating ventilation air conditioning system and the micro-grid system considering the comfort of the crowd and the cost of building energy consumption is constructed; S2, for the fine model built in step S1, a double-layer optimization control strategy is adopted to build a micro-grid energy management system framework for integrated heating, ventilation and air conditioning system; S3, based on step S2, a multi-energy coupling model of micro-grid-heating, ventilation and air conditioning system is built, and the time sequence optimization problem is converted into Markov decision process and mathematically characterized; S4, an improved priority experience replay deep deterministic policy gradient algorithm is adopted to train the agent, and a day-ahead optimization scheduling strategy is obtained, which takes the minimum comprehensive power cost and the highest personnel comfort as the target.

[0008] Further, the step S1 comprises the following steps: S101, modeling the micro-grid: the micro-grid in the present application refers to an energy network composed of wind turbines WT, photovoltaic PV cells, direct current diesel generators (DG) and energy storage systems (ESS). The photovoltaic PV cell converts solar energy into electrical energy through the photoelectric effect. The output power calculation formula is: Among them, represents the power output of the photovoltaic PV cell, represents its conversion efficiency, which is the ratio of electrical energy output to incident solar irradiance. represents the photovoltaic output under standard test conditions (photovoltaic output STC), represents the solar irradiance actually irradiated to the surface of the photovoltaic PV panel in the actual environment, corresponds to the irradiance value under the photovoltaic output STC condition, which is usually set to 1000 W / m². Coefficient is used to consider the temperature sensitivity, represents the deviation between the actual ambient temperature and the photovoltaic output STC temperature.

[0009] The output power of the wind turbine (WT) is derived from the mechanical power transmitted by the rotor shaft, and the calculation formula of the mechanical power is: Among them, represents the output power of the wind turbine, represents the efficiency of the generator. Parameter refers to the mechanical energy extracted from the wind, which depends on multiple parameters: represents the air density, is the wind speed, r represents the radius of the turbine blade, is called the power coefficient, which reflects the ability of the wind turbine to convert wind kinetic energy into mechanical energy For the energy storage system, the charge and discharge state of the energy storage battery is represented as: in, Indicates the energy storage system at time step k The state of charge, and It represents the amount of power exchanged with the energy storage system—positive during discharge and negative during charging. The parameter represents the self-discharge rate of the energy storage system. and These correspond to the efficiency during the discharge and charging processes, respectively. This indicates the rated energy capacity of the energy storage system. As a signal to control the charging / discharging mode.

[0010] Regarding the dynamic characteristics of diesel generator sets, their power output follows the equation: in, It is the response time constant (s) of the generator set. It is the output power fluctuation coefficient.

[0011] S102. Modeling HVAC: In the energy consumption structure of HVAC systems, the main energy-consuming units are variable frequency fans and refrigeration units.

[0012] The power consumption of the variable frequency fan can be described by a mathematical model, and its calculation formula is as follows: in, It is the material coefficient of the variable frequency fan. It is the air mass flow rate of the variable frequency fan.

[0013] The model calculation formula for HVAC refrigeration units is as follows: in, This represents the coefficient of performance (COP) of the refrigeration unit; It is the air supply quality flow rate; This indicates the temperature of the air entering the refrigeration unit; while This corresponds to the temperature of the air discharged from the refrigeration unit.

[0014] Furthermore, step S2 employs an enhanced model predictive control (eMPC) framework to construct an environmental control system, the construction method of which is as follows: Let the prediction range be... The sampling period is Objective function The definition is as follows: wherein, represents the total energy consumption of the heating ventilation and air conditioning system in the time interval . represents the comfort coefficient of the area i , which evaluates the closeness of the actual temperature to the desired set value - the higher the value, the smaller the deviation, thus obtaining better thermal comfort.

[0015] The system adopts a double-layer optimization control strategy, aiming to minimize the operation cost of the building group power grid. At the monitoring control level, a multi-objective dynamic optimization algorithm is used to develop optimal scheduling strategies for each component of the microgrid, including the energy storage system (ESS), wind turbine (WT), photovoltaic (PV) cell, diesel generator set (DG), and heating ventilation and air conditioning system. These control decisions are based on the real-time state vector , which reflects the real-time operation state of the microgrid. Specifically, the scheduling instruction of the energy storage system corresponds to the charge and discharge power allocation strategy in the k , k + 1] time interval, and the corresponding scheduling instructions 、 and correspond to the planned power generation of the diesel generator set, wind turbine, and photovoltaic cell. For the heating ventilation and air conditioning system, the monitoring layer determines the operation mode by optimizing the energy efficiency balance coefficient ( , ) in real time. When the system model is known, the controller outputs the optimal parameter combination; in the model-unknown scenario, the control instructions of the heating ventilation and air conditioning system are directly generated . The workflow of the controller can be expressed as: Further, the step S3 uses Markov Decision Process (MDP) to mathematically represent the time series optimization model, and MDP is formally defined as a four-tuple M = ( S, A, P, R ), where S is the set of environmental states, A is the set of action spaces, P is the state transition probability function, R is the reward function.

[0016] To ensure the safe operation of the ESS, a dynamic safety coefficient is introduced. Therefore, the operation constraints of the ESS can be represented as: At time t , we assume that the ESS will not generate any loss in the charging, discharging, or idle state. represents the predicted photovoltaic power, represents the predicted wind power generation. The total power demand is defined as follows: (1) State space The real-time electricity price of the grid can be defined as where the weight coefficients (energy consumption and comfort) in the microgrid energy management system (EMS) are denoted as and respectively. Therefore, the HVAC-microgrid state can be represented as follows: (2) Action space In the action space, the power exchange between the microgrid and the distribution grid is denoted as , and the charge-discharge power of the energy storage system is denoted as . Its numerical range is defined as where and represent the maximum discharge power and the maximum charge power, respectively. To maintain power balance, the total supply must equal the total demand. Therefore, the power balance constraint condition can be expressed as follows: where it is affected by the parameter , both parameters and need to be optimized to minimize the total energy consumption. The set of tuning parameters is denoted by the action . Therefore, A is defined in terms of these control variables: (3) Reward function The electricity selling price of the microgrid is . Let be the reward function, which is defined as: where represents the penalty coefficient imposed by the excessive charge-discharge of the energy storage system (ESS), and and represent the penalty coefficients associated with the HVAC parameters regulated by the eMPC. These coefficients are constrained within the interval [0, 1].

[0017] (4) Action value function The relationship between the system parameters, total power usage, and comfort indicators is modeled through the function : where is an approximation function used to model the relationship between input data.

[0018] Further, the step S4 introduces a prioritized experience replay mechanism (PER) into the deep deterministic policy gradient algorithm (DDPG) to solve the continuous control problem in Markov decision processes.

[0019] Within the framework of the HVAC-microgrid energy management system control, the training process of the PER-DDPG algorithm contains the following key stages: (1) State representation and action definition: At each discrete time step t , the agent perceives the current environment state , which contains key parameters such as indoor temperature deviation, state of charge (SoC) of energy storage units, and current electricity price of the main grid. The corresponding control action is composed of a series of continuous decision variables, including adjustments to the temperature setpoint and adjustments to the charge and discharge power levels of the energy storage system.

[0020] (2) Action generation and exploration: At each time step, the action network generates the optimal control action based on the current state. To avoid falling into local optima, the system generates random noise using the Ornstein-Uhlenbeck (OU) noise process and superimposes it on the control action, resulting in an exploratory action: (3) Experience replay and priority calculation: The time difference (TD) error and the priority of each sample are calculated, and the priority is used to adjust the sampling probability in experience replay: (4) Network update and policy optimization: A small batch of experience samples is randomly selected from the replay buffer. For each sample, the target Q value is calculated, and the evaluation network parameters are optimized by minimizing the loss function, which is defined based on the difference between the estimated Q value and the target Q value: a squared function based on Q value estimation error: The policy network is optimized using the policy gradient method, and the update of the policy parameters is as follows: (5) Target network update: The target critic network and the target actor network are updated using a soft update mechanism: (6) Training completion and decision execution: When the training phase is completed, the parameters of the policy network will be fixed and deployed for real-time control. Based on the current system state , the trained model generates corresponding control instructions , the system state , control instructions can adjust the temperature setting of heating, ventilation and air conditioning, manage the charging and discharging operation of the energy storage system, and adjust other related variables - all of these measures aim to improve indoor comfort and maximize energy efficiency.

[0021] Compared with the prior art, the beneficial effects of the present application are: Unlike most microgrid-heating, ventilation and air conditioning system collaborative scheduling methods, which use model predictive control to build optimization framework and rely on direct environmental observation, this study combines deep reinforcement learning with enhanced model predictive control technology. The DRL agent can dynamically adjust multiple key parameters of the MPC controller, making the system behavior more adaptive and intelligent. In addition, this study improves the DDPG algorithm by introducing the priority experience replay (PER) algorithm. This mechanism ensures that important experiences with larger time difference errors (TD errors) in the experience pool can be sampled more frequently in each training iteration, significantly improving model convergence speed and optimizing sample utilization. In addition, by combining the priority sampling strategy with the noise superposition technology, PER-DDPG can more accurately focus on key state-action pairs during exploration, effectively avoiding the dilemma of falling into local optimal solutions. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is the overall model and algorithm schematic diagram of the present application; Figure 2 is the schematic diagram of the interaction between the agent and the environment of the present application; DETAILED DESCRIPTION The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0023] Please refer to Figure 1 , Figure 1An overall schematic diagram of a micro-grid-heating ventilation air conditioning coordination optimization method based on deep reinforcement learning is provided, and the method comprises the following steps: (1) constructing a refined model of the heating ventilation air conditioning system and the micro-grid system considering the comfort of the crowd and the cost of building energy consumption; (2) for the refined model built in step 1, a double-layer optimization control strategy is adopted to build a micro-grid energy management system framework integrating the heating ventilation air conditioning system; (3) based on the multi-energy coupling model of the micro-grid-heating ventilation air conditioning system built in step 2, the time sequence optimization problem is converted into a Markov decision process and is mathematically characterized; (4) an improved priority experience replay deep deterministic policy gradient algorithm is used to train the agent, and a day-ahead optimization scheduling strategy with the minimum comprehensive power cost and the highest personnel comfort as the target is obtained.

[0024] Further, the following models are included in the figure: Micro-grid refined model: the micro-grid in the present application refers to an energy network composed of a wind turbine WT, a photovoltaic PV cell, a direct current diesel generator DG and an energy storage system ESS. The photovoltaic PV cell converts solar energy into electrical energy through the photoelectric effect. The output power calculation formula is: wherein, represents the power output of the photovoltaic PV cell, represents the conversion efficiency, which is the ratio of electrical energy output to incident solar irradiance. represents the photovoltaic output under standard test conditions (photovoltaic output STC), represents the solar irradiance actually irradiated to the surface of the photovoltaic PV panel in the actual environment, corresponds to the irradiance value under the photovoltaic output STC condition, which is usually set to 1000 W / m². The coefficient is used to consider the temperature sensitivity, represents the deviation between the actual environmental temperature and the photovoltaic output STC temperature.

[0025] The output power of the wind turbine (WT) is derived from the mechanical power transmitted by the rotor shaft, and the calculation formula of the mechanical power is: wherein, represents the output power of the wind turbine, represents the efficiency of the generator. The parameter refers to the mechanical energy extracted from the wind, which depends on multiple parameters: represents the air density, is the wind speed, r represents the radius of the turbine blade, The power coefficient, which reflects the ability of a wind turbine to convert wind kinetic energy into mechanical energy For the energy storage system, the state of charge of the energy storage battery is denoted as: wherein, denotes the state of charge of the energy storage system at time step k , represents the amount of power exchange with the energy storage system, positive for discharging and negative for charging. The parameter and correspond to the efficiency during discharging and charging, respectively. denotes the rated energy capacity of the energy storage system. is the signal to control the charging / discharging mode.

[0026] Regarding the dynamic characteristics of the diesel generator set, its power output follows the equation: wherein, is the response time constant of the generator set (s), is the output power fluctuation coefficient.

[0027] Fine model of heating, ventilation and air conditioning: In the energy consumption structure of heating, ventilation and air conditioning system, the main energy consumption unit is variable frequency fan and refrigeration unit.

[0028] The power consumption of the variable frequency fan can be described by a mathematical model, and its calculation formula is: wherein, is the material coefficient of the variable frequency fan, is the air supply mass flow rate of the variable frequency fan.

[0029] The calculation formula of the model of the heating, ventilation and air conditioning refrigeration unit is: wherein, denotes the performance coefficient of the refrigeration unit; is the air supply mass flow rate; denotes the temperature of the air entering the refrigeration unit; and corresponds to the temperature of the air discharged from the refrigeration unit.

[0030] Further, the step 2 adopts an enhanced model predictive control (eMPC) framework to construct the environment regulation system, which is constructed as follows: let the prediction range be , and the sampling period be . The objective function is The definitions are as follows: wherein, represents the total energy consumption of the heating ventilation and air conditioning system in the time interval. represents the comfort coefficient of the area , which evaluates the closeness of the actual temperature to the desired set value - the higher the value, the smaller the deviation, thereby obtaining better thermal comfort. i

[0031] The system adopts a double-layer optimization control strategy, aiming to minimize the operation cost of the building group power grid. At the monitoring control level, a multi-objective dynamic optimization algorithm is used to develop optimal scheduling strategies for each component of the microgrid, including the energy storage system (ESS), wind turbine (WT), photovoltaic PV cell, diesel generator set (DG) and heating ventilation and air conditioning system. These control decisions are based on the real-time state vector , which reflects the real-time operation state of the microgrid. Specifically, the scheduling instruction of the energy storage system corresponds to the charge and discharge power allocation strategy in the time interval k , k + 1], and the corresponding scheduling instructions 、 and correspond to the planned power generation of the diesel generator set, the fan and the photovoltaic. For the heating ventilation and air conditioning system, the monitoring layer determines the operation mode by optimizing the energy efficiency balance coefficient ( , ) in real time. When the system model is known, the controller outputs the optimal parameter combination; in the model unknown scenario, the control instruction of the heating ventilation and air conditioning system is directly generated. The workflow of the controller can be expressed as: Figure 2 The present application provides an intelligent agent interacting with the environment of a microgrid-heating ventilation and air conditioning coordinated optimization method based on deep reinforcement learning, including the following steps: using Markov decision process (MDP) to mathematically represent the time series optimization model, and the MDP is formally defined as a four-tuple M = (S, A, P, R). Its design integrates multiple evaluation factors, including power cost, equipment degradation rate and comfort index, etc. S, A, P, R

[0032] In order to ensure the safe operation of the ESS, a dynamic safety coefficient is introduced. Therefore, the operation constraint of the ESS can be expressed as: ​​ At time t, we assume that the ESS does not incur any losses in charging, discharging, or idle state. Ppv(t) represents the predicted photovoltaic power, Pwind(t) represents the predicted wind power. Total load power is defined as follows: (1) State space The real-time electricity price of the grid can be defined as where the weight coefficients (energy consumption and comfort) in the microgrid energy management system (EMS) are denoted as and respectively. Therefore, the HVAC-microgrid state can be represented as follows: (2) Action space In the action space, the power exchange between the microgrid and the distribution grid is denoted as , and the charging and discharging power of the energy storage system (ESS) is denoted as . Its numerical range is defined as where and represent the maximum discharging power and the maximum charging power, respectively. To maintain power balance, the total supply must equal the total demand. Therefore, the power balance constraint condition can be expressed as follows: where it is affected by the parameter , and Both parameters need to be optimized to minimize the total energy consumption. The set of tuning parameters is denoted by the action . Therefore, A is defined in terms of these control variables: (3) Reward function The electricity selling price of the microgrid is . Let R be the reward function, which is defined as: where represents the penalty coefficient imposed by the excessive charging and discharging of the energy storage system (ESS), and and represent the penalty coefficients associated with the HVAC parameters regulated by the eMPC. These coefficients are constrained within the interval [0, 1].

[0033] (4) Action value function The relationship between system parameters, total power usage, and comfort index is modeled by a function where is an approximation function used to model the relationship between input data.

[0034] Further, Figure 1 The algorithmic framework presented introduces a prioritized experience replay mechanism (PER) into the deep deterministic policy gradient algorithm (DDPG) to solve continuous control problems in Markov decision processes.

[0035] Within the control framework of the HVAC-microgrid energy management system, the training process of the PER-DDPG algorithm contains the following key stages: (1) State representation and action definition: At each discrete time step t , the agent perceives the current environment state , which contains key parameters such as indoor temperature deviation, state of charge (SoC) of energy storage units, and current electricity price of the main grid. The corresponding control action is composed of a series of continuous decision variables, including adjustments to the temperature setpoint and adjustments to the charge-discharge power level of the energy storage system.

[0036] (2) Action generation and exploration: At each time step, the action network generates the optimal control action based on the current state. To avoid falling into local optima, the system generates random noise using the Ornstein-Uhlenbeck (OU) noise process and superimposes it on the control action, resulting in an exploratory action: (3) Experience replay and priority calculation: The time-difference (TD) error and the priority of each sample are calculated, and the priority is used to adjust the sampling probability in experience replay: (4) Network update and policy optimization: A small batch of experience samples is randomly selected from the replay buffer. For each sample, the target Q value is calculated, and the evaluation network parameters are optimized by minimizing the loss function, which is defined based on the difference between the estimated Q value and the target Q value: a squared function based on Q value estimation error: ​The policy network is optimized using the policy gradient method, and the update of the policy parameters is as follows: (5) Target network update: The target critic network and the target actor network are updated using a soft update mechanism: (6) Training completion and decision execution: When the training phase is completed, the parameters of the policy network will be fixed and deployed for real-time control. Based on the current system state , the trained model generates corresponding control instructions These instructions can adjust the temperature settings of heating, ventilation, and air conditioning, manage the charging and discharging operations of the energy storage system, and adjust other related variables—all of which aim to improve indoor comfort and maximize energy efficiency.

[0037] It is apparent to those skilled in the art that the present application is not limited to the details of the foregoing exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered in all respects as illustrative and not restrictive, and the scope of the present application should be defined by the appended claims rather than the above description, and it is intended to include all changes falling within the meaning and range of equivalents of the claims. Any reference signs in the claims should not be considered as limiting the claims involved.

[0038] Furthermore, it should be understood that, although the present specification is described in terms of embodiments, not every embodiment contains only one independent technical solution, and the description of the specification is only for the sake of clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be appropriately combined to form other embodiments that those skilled in the art can understand.

Claims

1. A microgrid-HVAC coordinated optimization method based on deep reinforcement learning, characterized in that, Includes the following steps: Step S1: Construct a refined model of the HVAC system and microgrid system that takes into account the comfort of the population and the energy cost of the building; Step S2: Based on the refined model constructed in Step S1, a two-layer optimization control strategy is adopted to construct a microgrid energy management system framework that integrates the HVAC system. Step S3: Based on step S2, construct a multi-energy coupling model of the microgrid-HVAC system, transform the time-series optimization problem into a Markov decision process and perform mathematical representation; Step S4: The agent is trained using the improved Priority Experience Replay Deep Deterministic Strategy Gradient Algorithm (PER-DDPG). The agent is defined as a microgrid control system integrating HVAC. Finally, a day-ahead optimization scheduling strategy is obtained with the goal of minimizing overall power cost and maximizing personnel comfort.

2. The microgrid-HVAC coordinated optimization method based on deep reinforcement learning according to claim 1, characterized in that, Step S1 specifically involves: (1.1) Modeling the microgrid: (1.1.1) Photovoltaic (PV) cells convert solar energy into electrical energy through the photoelectric effect. The formula for calculating the output power is: in, This indicates the power output of a photovoltaic (PV) cell. Representing its conversion efficiency, it is the ratio of electrical energy output to incident solar irradiance. This represents the photovoltaic output STC under standard test conditions. This represents the solar irradiance that hits the surface of a photovoltaic (PV) solar panel under actual conditions. This corresponds to the irradiance value under STC conditions for photovoltaic output, and the coefficient... For consideration of temperature sensitivity, This indicates the deviation between the actual ambient temperature and the STC temperature at the photovoltaic output; (1.1.2) The output power of the wind turbine generator WT is derived from the mechanical power transmitted by the rotor shaft. The formula for calculating the mechanical power is: in, This indicates the output power of the wind turbine. Represents the efficiency of the generator, parameters Mechanical energy extracted from the wind, the value of which depends on several parameters: Represents air density, It's wind speed. r Indicates the radius of the turbine blade. Known as the power factor, it reflects the ability of a wind turbine to convert wind kinetic energy into mechanical energy; (1.1.3) For energy storage systems, the charging and discharging states of the energy storage battery are expressed as follows: in, Indicates the energy storage system at time step k The state of charge, This represents the amount of power exchanged with the energy storage system—positive during discharge and negative during charging. The parameter represents the self-discharge rate of the energy storage system. and These correspond to the efficiency during the discharge and charging processes, respectively. Indicates the rated energy capacity of the energy storage system. As a signal to control the charging / discharging mode; (1.1.4) Regarding the dynamic characteristics of the diesel generator set, its power output follows the following equation: in, It is the response time constant (s) of the generator set. It is the output power fluctuation coefficient; (1.2) Modeling the HVAC system: (1.2.1) The power consumption of the variable frequency fan is described by a mathematical model, and its calculation formula is as follows: in, It is the material coefficient of the variable frequency fan. It is the air delivery mass flow rate of the variable frequency fan; (1.2.2) The model calculation formula for HVAC refrigeration units is as follows: in, This represents the coefficient of performance (COP) of the refrigeration unit; It is the air supply quality flow rate; This indicates the temperature of the air entering the refrigeration unit; This corresponds to the temperature of the air discharged from the refrigeration unit.

3. The microgrid-HVAC coordinated optimization method based on deep reinforcement learning according to claim 1, characterized in that, Step S2 specifically involves: (2.1) An environmental control system is constructed using the enhanced model predictive control (eMPC) framework, as follows: Let the prediction range be... The sampling period is objective function The definition is as follows: in, Indicates in Total energy consumption of the HVAC system during the time interval Indicates the area i The comfort factor assesses how close the actual temperature is to the desired set value; a higher value indicates a smaller deviation. (2.2) A two-layer optimization control strategy is adopted. In the monitoring and control layer, a multi-objective dynamic optimization algorithm is used to formulate the optimal scheduling strategy for each component of the microgrid, covering the energy storage system (ESS), wind turbine (WT), photovoltaic (PV) cells, diesel generator (DG), and HVAC system. The control decision is based on the real-time state vector. Dispatch instructions for energy storage systems Corresponding to [ k , k + 1] The charging and discharging power allocation strategy within the time period, and the corresponding scheduling instructions. , and This corresponds to the planned power generation of diesel generator sets (DG), wind turbine generators (WT), and photovoltaic (PV) cells; for HVAC systems, the monitoring and control layer optimizes the energy efficiency balance coefficient in real time. , The controller determines the operating mode. When the system model is known, it outputs the optimal parameter combination; in scenarios where the model is unknown, it directly generates control commands for the HVAC system. The controller's workflow is described as follows: 。 4. The microgrid-HVAC coordinated optimization method based on deep reinforcement learning according to claim 1, characterized in that, Step S3 specifically involves: (3.1) The time series optimization model is mathematically represented using a Markov Decision Process (MDP), which is defined as a quadruple. M = ( S, A, P, R ),in S A collection of environmental states. A For the collection of action spaces, P Let be the state transition probability function. R For the reward function; (3.2) A dynamic safety factor is introduced. To ensure the safe operation of the energy storage system (ESS), the operating constraints of the ESS are expressed as follows: (3.3) In time t Assuming that the energy storage system (ESS) does not incur any losses during charging, discharging, or idle states, Indicates the predicted photovoltaic power. This indicates the predicted wind power generation capacity and total load power. The definition is as follows: (3.4.1) State Space The real-time electricity price of the power grid can be defined as In the microgrid energy management system (EMS), the weighting coefficients for energy consumption and comfort are respectively expressed as: and The status of the HVAC system-microgrid is represented as follows: (3.4.2) Space of Action In the action space, the power exchange between the microgrid and the distribution network is represented as: The charging and discharging power of the energy storage system is expressed as The numerical range is defined as ,in and These represent the maximum discharge power and the maximum charging power, respectively. The power balance constraint is stated as follows: Where the parameter is Influence, and Both parameters need to be optimized to minimize total power consumption; the tuning parameter set consists of actions. express, A The definition is as follows: (3.4.3) Reward Function The electricity sold by the microgrid is , Let the reward function be defined as follows: in, This represents the penalty coefficient imposed on the overcharging and discharging of the energy storage system (ESS), while and This represents the penalty coefficient associated with the parameters of the HVAC system regulated by the enhanced model predictive control eMPC framework, and the penalty coefficient is constrained within the interval [0,1]. (3.4.4) Action Value Function The relationship between system parameters, total power usage, and comfort indicators is expressed through a function. Modeling: in, It is an approximate function used to model the relationship between input data.

5. The microgrid-HVAC coordinated optimization method based on deep reinforcement learning according to claim 1, characterized in that, Step S4 specifically involves: Within the control framework of the HVAC system-microgrid energy management system, the training process of the improved Priority Experience Replay Deep Deterministic Policy Gradient Algorithm (PER-DDPG) is as follows: (4.1) State representation and operation definition: At each discrete time step t, the agent perceives the current environmental state. This includes indoor temperature deviation, the state-of-charge (SoC) of the energy storage unit, and the current electricity price parameters of the main grid, as well as the corresponding control actions. It consists of a series of continuous decision variables, including the adjustment of the temperature setpoint and the adjustment of the charging and discharging power level of the energy storage system; (4.2) Action Generation and Exploration: At each time step t, the action network generates the optimal control action based on the current state, using an Ornstein-Uhlenbeck noise process to generate random noise. This is then superimposed on the control action to generate exploratory actions: (4.3) Experience replay and priority calculation: Calculate the time difference (TD) error and the priority of each sample Priority is used to adjust the sampling probability in experience replay: (4.4) Network updates and policy optimization: A mini-batch of empirical samples is randomly selected from the playback buffer. For each sample, the target Q-value is calculated, and the evaluation network parameters are optimized by minimizing the loss function. The loss function is defined based on the difference between the estimated Q-value and the target Q-value: a function of the squared error of the Q-value estimation. The policy network is optimized using the policy gradient method, and the policy parameters are updated as follows: (4.5) Target network update: The target critic network and the target actor network are updated using a soft update mechanism: (4.6) Training completion and decision execution: Once the training phase is complete, the parameters of the policy network are fixed and deployed for real-time control, based on the current system state. The trained model generates corresponding control commands. Control commands It can adjust the temperature settings of HVAC systems, manage the charging and discharging operations of energy storage systems, and adjust other related variables.

Citation Information

Patent Citations

  • Micro-grid energy storage optimization scheduling method based on deep reinforcement learning

    CN117833285A

  • Micro-grid energy optimization scheduling method oriented to source grid load storage

    CN118971051A

  • Metareinforcement learning-based household micro-grid energy optimization control method

    CN119382104A

  • Multi-microgrid coordination control method based on deep learning

    CN119674976A

  • Household energy low-carbon operation strategy based on block chain and hybrid evolution reinforcement learning

    CN120595572A

Cited By

  • Power cooperative scheduling evaluation method based on reinforcement learning

    CN121688843A