Energy control method, device and equipment during microgrid operation

By modeling the microgrid energy distribution process into a multi-group structure, the adaptability and computational complexity of the microgrid energy control model is solved, and the stability and adaptability of energy control are improved, meeting the energy control needs of the microgrid.

CN119628081BActive Publication Date: 2025-09-05EAST CHINA BRANCH OF STATE GRID CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411430342.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-09-05
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

In the prior art, the microgrid energy control method is difficult to update in real time due to model complexity and uncertain factors, resulting in poor adaptability, high computational complexity and unsatisfactory optimization results, making it difficult to meet energy control needs.

Method used

The Markov decision-making process of modeling multiple control factors that affect energy distribution during the operation of the microgrid is a multi-group structure. The sample data of the energy distribution decision is trained and dynamically updated and adjusted through the Markov decision-making process to obtain an energy control model, and output the energy control strategy of the current tense to maintain energy balance.

Benefits of technology

The stability and adaptability of microgrid energy control have been improved, and the energy control strategy can be continuously optimized in a changing environment to meet energy control needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119628081B_ABST
    Figure CN119628081B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, and apparatus for energy control during the operation of a microgrid, which relates to the field of energy technology and is capable of dynamically controlling energy during the operation of a microgrid, thereby improving the stability and adaptability of the energy control strategy. The method includes: determining multiple control factors that affect energy distribution during the operation of a microgrid, modeling the energy control process of the microgrid operation as a Markov decision process of a multi-tuple structure based on the multiple control factors, performing model training and dynamic update adjustment on sample data describing energy distribution decisions through the Markov decision process, and obtaining an energy control model of the microgrid. The energy control model is used to output the current energy control strategy in a given microgrid environment, and take energy distribution actions on the microgrid according to the current energy control strategy to maintain energy balance during the operation of the microgrid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of energy technology, and in particular to a method, device and equipment for energy control during the operation of a microgrid. Background Art

[0002] As a flexible and efficient energy system, microgrids can play an important role in the access and consumption of distributed energy. However, the energy control problem of microgrids is very challenging due to its high complexity.

[0003] In related technologies, model-based optimization methods typically rely on explicit mathematical models, attempting to accurately describe the various components of a microgrid and their interrelationships. These models are then used to control the microgrid's energy through these interrelationships. However, constructing an accurate model that fully captures the operational characteristics of a microgrid is extremely difficult, especially when considering multiple uncertainties. Simple microgrid models are difficult to update in real time to adapt to environmental changes in the microgrid's operation. This results in poor adaptability, high computational complexity, and unsatisfactory optimization results during microgrid operation, making it difficult to meet the microgrid's energy control needs. Summary of the Invention

[0004] In view of this, the present application provides an energy control method, device and equipment for the operation process of a microgrid. The main purpose is to solve the problems faced by the microgrid operation process in the existing technology, such as poor adaptability, high computational complexity and unsatisfactory optimization effect, which makes it difficult to meet the energy control needs of the microgrid.

[0005] According to a first aspect of the present application, a method for energy control during a microgrid operation process is provided, comprising:

[0006] Identify multiple control factors that affect energy distribution during microgrid operation;

[0007] Modeling the energy control process of the microgrid operation as a Markov decision process of a multi-tuple structure according to the multiple control factors;

[0008] The Markov decision process is used to perform model training and dynamic update adjustment on sample data describing energy allocation decisions to obtain an energy control model of the microgrid, wherein the energy control model is used to output a current energy control strategy in a given microgrid environment;

[0009] An energy distribution action is taken on the microgrid according to the energy control strategy of the current tense to maintain energy balance during the operation of the microgrid.

[0010] Furthermore, the determination of multiple control factors affecting energy distribution during the operation of the microgrid includes:

[0011] According to the changes in electricity usage during the operation of the microgrid, determine the supply control factors and demand control factors that affect energy distribution during the operation of the microgrid;

[0012] On the basis of the supply control factors and the demand control factors, according to the change of energy prices during the operation of the microgrid, the cost control factors affecting energy distribution during the operation of the microgrid are determined.

[0013] Furthermore, before performing model training and dynamic updating and adjustment on the sample data describing the energy allocation decision through the Markov decision process to obtain the energy control model of the microgrid, the method further includes:

[0014] Obtaining energy strategy data output by the microgrid during historical operation, wherein the energy strategy data includes cost data of energy flow, energy storage data, and action data of energy control;

[0015] The energy strategy data is recorded in the form of a Markov decision process to obtain sample data describing the energy allocation decision.

[0016] Furthermore, the energy control process of the microgrid operation is modeled as a Markov decision process of a multi-tuple structure according to the multiple control factors, including:

[0017] According to the energy control process of microgrid operation, the energy state data describing the microgrid operating environment at different time states are used as the system state of the Markov decision process;

[0018] According to the energy control process of microgrid operation, the energy action data of microgrid in different time states are used as the action space of Markov decision process;

[0019] According to the energy control process of the microgrid, the reward function of the Markov decision process is constructed using the reward / penalty that describes the energy allocation actions taken by the microgrid.

[0020] According to the energy control process of the microgrid, the probability of transferring the energy allocation action of the current state to the energy allocation action of the next state is used as the transition probability of the Markov decision process;

[0021] The system state, the action space, the reward function and the transition probability are used as a multi-group structure to model the energy control process of the microgrid operation as a Markov decision process with a multi-group structure.

[0022] Furthermore, the energy control model of the microgrid is obtained by performing model training and dynamic update adjustment on the sample data describing the energy allocation decision through the Markov decision process, including:

[0023] The Markov decision process is used to perform model training and dynamic update adjustment on sample data describing energy allocation decisions, so as to calculate, during the model training process, the transition probability of the microgrid taking energy allocation actions in different time states and jumping to the energy allocation actions taken by the microgrid in the next time state according to the sample data;

[0024] Calculate the cumulative reward value given to the microgrid for taking the energy allocation action according to the transition probability;

[0025] On the basis of satisfying the balance of energy supply and demand, the energy control strategy output by the energy control model in a given microgrid environment is continuously adjusted so that the energy allocation action taken by the energy control strategy on the microgrid can satisfy the maximum cumulative reward value.

[0026] Furthermore, before performing model training and dynamic updating and adjustment on the sample data describing the energy allocation decision through the Markov decision process to obtain the energy control model of the microgrid, the method further includes:

[0027] Based on the energy trajectory data output by the microgrid during its historical operation, energy trajectories that meet the set conditions are selected for cumulative rewards to construct a first resource data pool. The energy trajectory data is the energy trajectory formed by the energy strategy data continuously output by the energy dispatcher from the initial moment to the end moment.

[0028] Applying energy strategy data provided by different experience scenarios to a Markov decision process, selecting energy strategy data that meets the state transition conditions in the Markov decision process to construct a second resource data pool, wherein the different experience scenarios correspond to mixing coefficients, and the weights of the energy strategy data provided by the experience scenarios in the second resource data pool are adjusted by the mixing coefficients;

[0029] The sample data describing the energy allocation decision is updated according to the first resource data pool and the second resource data pool, so as to adjust the parameters of the energy control model through the updated sample data describing the energy allocation decision.

[0030] Furthermore, the parameter adjustment of the energy control model by using the updated sample data describing the energy allocation decision includes:

[0031] Based on the updated sample data describing the energy allocation decision, the reward function is used to calculate the cumulative reward value of the sample data;

[0032] The sample data are sorted from large to small according to the cumulative reward value, and the energy trajectory data and / or energy strategy data corresponding to the sample data whose cumulative reward value is sorted before the preset value are selected to adjust the parameters of the energy control model of the microgrid.

[0033] According to a second aspect of the present application, there is provided an energy control device for a microgrid operation process, comprising:

[0034] a determination unit, configured to determine a plurality of control factors affecting energy distribution during operation of the microgrid;

[0035] A modeling unit, configured to model the energy control process of the microgrid operation as a Markov decision process of a multi-tuple structure according to the plurality of control factors;

[0036] A training unit is configured to perform model training and dynamic update adjustment on sample data describing energy allocation decisions through the Markov decision process to obtain an energy control model of the microgrid, wherein the energy control model is configured to output a current energy control strategy in a given microgrid environment;

[0037] The control unit is used to take energy distribution actions on the microgrid according to the energy control strategy of the current time state to maintain the energy balance during the operation of the microgrid.

[0038] Furthermore, the determining unit is specifically configured to:

[0039] According to the changes in electricity usage during the operation of the microgrid, determine the supply control factors and demand control factors that affect energy distribution during the operation of the microgrid;

[0040] On the basis of the supply control factors and the demand control factors, according to the change of energy prices during the operation of the microgrid, the cost control factors affecting energy distribution during the operation of the microgrid are determined.

[0041] Furthermore, the device further comprises:

[0042] an acquisition unit, configured to acquire energy strategy data output by the microgrid during a historical operation process before performing model training and dynamic update adjustment on the sample data describing the energy allocation decision through the Markov decision process to obtain an energy control model of the microgrid, wherein the energy strategy data includes cost data of energy flow, energy storage data, and action data of energy control;

[0043] The recording unit is used to record the energy strategy data in the form of a Markov decision process to obtain sample data describing the energy allocation decision.

[0044] Furthermore, the modeling unit is specifically used to:

[0045] According to the energy control process of microgrid operation, the energy state data describing the microgrid operating environment at different time states are used as the system state of the Markov decision process;

[0046] According to the energy control process of microgrid operation, the energy action data of microgrid in different time states are used as the action space of Markov decision process;

[0047] According to the energy control process of the microgrid, the reward function of the Markov decision process is constructed using the reward / penalty that describes the energy allocation actions taken by the microgrid.

[0048] According to the energy control process of the microgrid, the probability of transferring the energy allocation action of the current state to the energy allocation action of the next state is used as the transition probability of the Markov decision process;

[0049] The system state, the action space, the reward function and the transition probability are used as a multi-group structure to model the energy control process of the microgrid operation as a Markov decision process with a multi-group structure.

[0050] Furthermore, the training unit is specifically used to:

[0051] The Markov decision process is used to perform model training and dynamic update adjustment on sample data describing energy allocation decisions, so as to calculate, during the model training process, the transition probability of the microgrid taking energy allocation actions in different time states and jumping to the energy allocation actions taken by the microgrid in the next time state according to the sample data;

[0052] Calculate the cumulative reward value given to the microgrid for taking the energy allocation action according to the transition probability;

[0053] On the basis of satisfying the balance of energy supply and demand, the energy control strategy output by the energy control model in a given microgrid environment is continuously adjusted so that the energy allocation action taken by the energy control strategy on the microgrid can satisfy the maximum cumulative reward value.

[0054] Furthermore, the device further comprises:

[0055] A first construction unit is configured to, before performing model training and dynamic updating and adjustment on the sample data describing the energy allocation decision through the Markov decision process to obtain the energy control model of the microgrid, select energy trajectories whose cumulative rewards meet set conditions based on energy trajectory data output by the microgrid during historical operation to construct a first resource data pool, wherein the energy trajectory data is an energy trajectory formed by energy strategy data continuously output by the energy dispatcher from an initial moment to an end moment;

[0056] A second construction unit is configured to apply energy strategy data provided by different experience scenarios to a Markov decision process, select energy strategy data that meets the state transition condition in the Markov decision process to construct a second resource data pool, wherein the different experience scenarios correspond to mixing coefficients, and the weights of the energy strategy data provided by the experience scenarios in the second resource data pool are adjusted by the mixing coefficients;

[0057] An updating unit is used to update the sample data describing the energy allocation decision according to the first resource data pool and the second resource data pool, so as to adjust the parameters of the energy control model through the updated sample data describing the energy allocation decision.

[0058] Furthermore, the updating unit is specifically configured to:

[0059] Based on the updated sample data describing the energy allocation decision, the reward function is used to calculate the cumulative reward value of the sample data;

[0060] The sample data are sorted from large to small according to the cumulative reward value, and the energy trajectory data and / or energy strategy data corresponding to the sample data whose cumulative reward value is sorted before the preset value are selected to adjust the parameters of the energy control model of the microgrid.

[0061] According to a third aspect of the present application, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method described in the first aspect when executing the computer program.

[0062] According to a fourth aspect of the present application, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0063] By means of the above-mentioned technical solution, the present application provides a method, apparatus, and device for energy control during the operation of a microgrid. Compared with the existing method of controlling microgrid energy through the interrelationships between the various components of the microgrid, the present application determines multiple control factors that affect energy distribution during the operation of the microgrid, models the energy control process of the microgrid operation as a Markov decision process with a multi-tuple structure based on the multiple control factors, and uses the Markov decision process to train and dynamically update and adjust the sample data describing the energy distribution decision to obtain an energy control model of the microgrid. The energy control model is used to output the current energy control strategy in a given microgrid environment, and take energy distribution actions for the microgrid based on the current energy control strategy to maintain energy balance during the operation of the microgrid. The entire process models the energy control process of the microgrid operation as a Markov decision process with a multi-tuple structure, realizes dynamic control of energy during the operation of the microgrid, improves the stability and adaptability of the energy control strategy, and enables the microgrid to continuously optimize the energy control strategy in a changing environment, greatly meeting the energy control needs of the microgrid.

[0064] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0066] Figure 1 This is a flow chart of an energy control method for a microgrid operation process according to an embodiment of the present application;

[0067] Figure 2 yes Figure 1 A schematic flow chart of a specific implementation of step 101;

[0068] Figure 3 yes Figure 1 A schematic flow chart of a specific implementation of step 102;

[0069] Figure 4 yes Figure 1 A schematic flow chart of a specific implementation of step 103;

[0070] Figure 5 This is a schematic diagram of the structure of an energy control device during the operation of a microgrid in one embodiment of the present application;

[0071] Figure 6 The figure is a schematic diagram of the device structure of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0072] The present invention will now be discussed with reference to several exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the present invention, rather than to imply any limitation on the scope of the present invention.

[0073] As used herein, the term "including" and variations thereof are to be interpreted as open-ended terms meaning "including, but not limited to." The term "based on" is to be interpreted as "based, at least in part, on." The terms "one embodiment" and "an embodiment" are to be interpreted as meaning "at least one embodiment." The term "another embodiment" is to be interpreted as meaning "at least one other embodiment."

[0074] In related technologies, data-driven offline reinforcement learning methods can be applied to the energy control process of microgrid operation. However, offline reinforcement learning methods have high requirements on the distribution of sample data. When faced with unknown scenarios and complex environments, existing offline reinforcement learning methods will produce large deviations due to changes in the distribution of sample data, thereby affecting the stability of energy control.

[0075] In order to solve this problem, this embodiment provides an energy control method for the operation process of a microgrid, such as Figure 1 As shown, the following steps are included:

[0076] 101. Identify multiple control factors that affect energy distribution during microgrid operation.

[0077] As an energy control system, a microgrid includes energy supply devices, energy consumption devices, energy storage devices, control and protection devices, and communication monitoring devices. Among them, the energy supply devices include distributed photovoltaic, wind, biomass, geothermal, wave, fuel, gas and other power generation micro-power sources and power conversion grid-connected devices. The energy consumption devices include power loads, cooling loads, heating loads, etc. The energy storage devices include a variety of distributed energy storage bodies, grid-connected devices and battery management systems. The control and protection devices include central controllers, control master stations, grid-connected switches, circuit breakers, etc. The communication monitoring devices include communication, sensing and energy management systems.

[0078] Considering that energy control is the core issue of microgrid operation, its main goal is to ensure energy supply and demand balance and maximize the use of renewable energy. Microgrids can calculate the optimal energy control strategy based on the current energy supply and demand situation and electricity price data to control the energy distribution and conversion during the microgrid operation process, ensuring the stability and reliability of the microgrid energy supply. In other words, the multiple control factors affecting energy distribution during microgrid operation include at least energy supply, energy demand, and capacity cost. The specific energy supply sources are mainly solar energy, wind energy, photovoltaic energy, etc., and the energy demand sources are mainly battery energy storage, engine load, etc. These energy supply and energy demand are characterized by uncertainty and intermittence. Considering that energy supply and energy demand have different energy costs, this energy cost is mainly composed of energy procurement cost and energy conversion cost.

[0079] The executor of this embodiment can be an energy control device or equipment in the microgrid operation process, which can be configured at the service end of the microgrid energy control. It uses multiple control factors that affect energy distribution during the microgrid operation as energy constraint conditions for the microgrid operation to determine the optimal energy control strategy, which can minimize energy costs and energy waste while ensuring the balance of microgrid energy supply.

[0080] 102. Model the energy control process of the microgrid operation as a Markov decision process with a multi-tuple structure based on the multiple control factors.

[0081] In this embodiment, the energy control process of the microgrid operation is mainly to control the energy flow and consumption of each energy resource in the microgrid system. For example, during the energy flow process of each energy resource, the charging and discharging of the energy storage system is controlled, and during the energy consumption process of each energy resource, the purchase and sale of electricity by the main grid is controlled.

[0082] Specifically, the energy control process of a microgrid is modeled as a multi-tuple Markov decision process, consisting primarily of a system state S, an action space A, a reward function R, and a state transition probability P. The goal of the Markov decision process is to find an energy control strategy that maximizes the expected cumulative reward within a given microgrid environment and outputs the energy control strategy for the current state. Here, the system state S describes the microgrid environment at a given time t, the action space A describes the energy allocation actions taken under system state S, the reward R describes the reward / penalty for taking an energy allocation action under system state S, and the state transition probability P describes the probability of a change in the energy allocation action that causes the system state to jump from the current state to the next state.

[0083] 103. Perform model training and dynamic update adjustment on sample data describing energy allocation decisions through the Markov decision process to obtain an energy control model of the microgrid.

[0084] Among them, the energy control model is used to output the current energy control strategy in a given microgrid environment. Under the influence of different control factors, the system state S, action space A, reward function R and state transition probability P of the Markov decision process usually change during the operation of the microgrid.

[0085] Specifically, different energy allocation actions are selected in the action space through changing states, and the cumulative rewards obtained by selecting energy allocation actions in different states are calculated. Based on the cumulative rewards, it is evaluated whether the energy control effect of the energy allocation action selected in the current state meets the expectations. Therefore, the goal of the Markov decision process is to find an energy control strategy so that the microgrid can achieve the expected control effect by adopting the energy allocation action corresponding to the strategy in the current state.

[0086] 104. Take energy distribution actions on the microgrid according to the current energy control strategy to maintain energy balance during the operation of the microgrid.

[0087] It can be understood that the energy control model obtained through Markov decision process training can output an energy control strategy that meets the needs for the current time state. This energy control strategy can minimize costs and maximize benefits during the operation of the microgrid through processes such as reasonable energy distribution, control of energy storage systems, and energy transactions with the main grid.

[0088] The energy control method for the microgrid operation process provided in the embodiments of the present application is compared with the existing method of controlling the energy of the microgrid through the relationship between the various components of the microgrid. The present application determines multiple control factors that affect energy distribution during the operation of the microgrid, models the energy control process of the microgrid operation as a Markov decision process with a multi-tuple structure based on the multiple control factors, and uses the Markov decision process to train the model and dynamically update and adjust the sample data describing the energy distribution decision to obtain an energy control model of the microgrid. The energy control model is used to output the current energy control strategy in a given microgrid environment, and take energy distribution actions for the microgrid according to the current energy control strategy to maintain energy balance during the operation of the microgrid. The entire process models the energy control process of the microgrid operation as a Markov decision process with a multi-tuple structure, realizes dynamic control of the energy during the operation of the microgrid, improves the stability and adaptability of the energy control strategy, enables the microgrid to continuously optimize the energy control strategy in a changing environment, and greatly meets the energy control needs of the microgrid.

[0089] In the above embodiment, considering that the operation process of the microgrid is affected by various uncertain factors such as wind energy, photovoltaic power generation, load demand and electricity price fluctuations, in order to meet the energy supply and demand balance during the operation process of the microgrid, specifically, Figure 2 As shown, step 101 includes the following steps:

[0090] 201. Based on the changes in electricity usage during the operation of the microgrid, determine the supply control factors and demand control factors that affect energy distribution during the operation of the microgrid.

[0091] 202. Based on the supply control factors and demand control factors, and according to the changes in energy prices during the operation of the microgrid, determine the cost control factors that affect energy distribution during the operation of the microgrid.

[0092] In this embodiment, the changes in electricity usage during the operation of the microgrid mainly include the following aspects: first, the microgrid utilizes distributed energy resources such as solar energy, wind energy, and hydropower to convert them into internal loads of the power supply system; on the other hand, when the distributed energy resource function is insufficient, the microgrid releases the stored energy through the energy storage device, and when the distributed energy function is excessive, the excess electricity is stored through the energy storage device.

[0093] It can be understood that the energy supply source during the operation of the microgrid is mainly distributed energy resources and energy storage equipment. The energy conversion rate of distributed energy and the energy currently stored in the energy storage equipment can be used to determine the supply control factors affecting energy distribution during the operation of the power grid. The energy demand source during the operation of the microgrid is mainly energy equipment. The demand control factors affecting energy distribution during the operation of the power grid can be determined by the load conditions of the energy equipment.

[0094] Furthermore, considering the changes in energy prices during energy operation, the cost changes will occur when the microgrid purchases and sells electricity. The cost changes caused by the microgrid purchasing and selling electricity can be used to determine the cost control factors that affect energy distribution during the operation of the microgrid. The cost control factors can be used to judge the energy benefits achieved by the microgrid using different energy control strategies.

[0095] In the above embodiment, specifically, Figure 3 As shown, step 102 includes the following steps:

[0096] 301. According to the energy control process of microgrid operation, energy state data describing the microgrid operating environment in different time states are used as the system state of the Markov decision process.

[0097] 302. According to the energy control process of microgrid operation, the energy action data describing the microgrid in different time states are used as the action space of the Markov decision process.

[0098] 303. According to the energy control process of the microgrid, a reward function of the Markov decision process is constructed using the reward / penalty description of the energy allocation action taken by the microgrid.

[0099] 304. According to the energy control process of the microgrid operation, the probability of transferring the energy allocation action of the current state to the energy allocation action of the next state is used as the transition probability of the Markov decision process.

[0100] 305. Use the system state, the action space, the reward function, and the transition probability as a multi-tuple structure to model the energy control process of the microgrid operation as a Markov decision process with a multi-tuple structure.

[0101] In this embodiment, the Markov decision process includes a system state S, an action space A, a reward function R, and a state transition probability P.

[0102] Specifically, the energy state data describing the microgrid operating environment in different time states can be used as the system state of the Markov decision process, as shown in the following expression:

[0103] S t ={t,P t,PV ,P t,Wind ,P t,load ,E t,ess ,Price buy ,Price sell}

[0104] Among them, t is the current time step (a day is divided into 24 time steps), P t,PV is the output power of photovoltaic power generation, P t,Wind is the output power of wind power generation, P t,load is the total power consumption of the load, E t,ess is the remaining energy in the energy storage system, Price buy Price is the real-time price of electricity purchased from the main power grid. sell The price of electricity sold to the main grid.

[0105] It is understandable that the energy storage system plays a key role in the operation of the microgrid. Reasonable control of the energy storage system can balance the energy supply and demand of the microgrid and enhance the stability and reliability of the microgrid operation. In addition, the energy storage system also supports the microgrid to optimize the electricity sales and / or electricity purchase strategy when energy prices fluctuate to maximize cost-benefit. In other words, the energy storage system is an important control component in the microgrid. In this embodiment, a lithium battery is used as an energy storage system model to simulate the charging and discharging process of the energy storage system through the energy storage system model. It is assumed that the charging and discharging voltage and efficiency are constants, and a linear approximation is used to limit the energy content. The battery energy state in each period is shown in the following expression:

[0106] E t =E t-1 +ΔE(t)

[0107] Among them, E t is the battery energy content at the end of the current period t, E t-1 is the battery energy content in the previous period, and ΔE(t) is the energy change in the battery during period t. ΔE(t) is calculated based on power and charge / discharge efficiency and can be defined as follows:

[0108]

[0109] Among them, P ch (t) is the power applied to the battery during time period t, η ch and η dc denote the charging and discharging efficiencies, respectively, and are assumed to be constants. ch When (t) is positive, it indicates that the energy storage device is charging, and vice versa, it indicates that the energy storage device is discharging.

[0110] Specifically, the energy action data taken by the microgrid in different time states is used as the action space of the Markov decision process, as shown in the following expression:

[0111] A t ={A 1t ,A 2t}

[0112] A 1t ={A ch ,A dc ,A buy ,A sell ,A DG})

[0113] Among them, A ch and A dc Respectively represent the charging and discharging of the energy storage system, A buy and A sell Represents electricity purchase and electricity sales, ADG Indicates the use of distributed generators for electricity generation.

[0114] -(A 2t ={B1,B2,S1,S2})

[0115] Among them, B1 and B2 represent additional electricity purchased and stored in the energy storage system, and S1 and S2 represent electricity sold to the main grid. The energy levels of B2 and S2 are higher than B1 and S1.

[0116] Specifically, the reward function of the Markov decision process is constructed by describing the rewards / penalties obtained by the microgrid in taking energy allocation actions. See the following expression:

[0117] R t =r 1t +ar 2t +br 3t

[0118] r 1t =-∑(C 1t +C 2t )

[0119] Among them, r 1t is the negative value of the total cost of the system, C 1t and C 2t They represent the cost of distributed generator power generation and the cost of purchasing and selling electricity, respectively. They are negative because the goal in the Markov decision process is to maximize the reward, so the cost is minimized. 2t It is a reward or penalty for the operating status of the energy storage system, used to maintain the energy state of the energy storage system at a desired level to increase its lifespan. 3t It is a penalty for unreasonable actions in the energy management process, such as energy waste. The reward and punishment coefficients a and b are both positive constants used to balance the effects of different rewards or penalties.

[0120] For C 1t This can be calculated using the microgrid's electricity cost C, which primarily includes the fuel cost of the distributed generator, the cost of purchasing electricity from the main grid, and the revenue from selling electricity. Specifically, the cost of a controllable generator can usually be expressed using a quadratic cost function:

[0121] C(P)=a+bP(t)+cP(t) 2

[0122] Where C(P) is the power generation cost, P(t) is the generator output power, and a, b, and c are cost function coefficients related to generator characteristics and operating conditions. It should be noted that given that distributed generators in microgrids are typically small-scale and simulation results show that distributed generators rarely start or stop, the generator startup and shutdown costs are ignored here.

[0123] For C 2t The cost of electricity purchase and sale C of microgrid net (t) is calculated, where the cost of purchasing and selling electricity of the microgrid is C net (t), can be defined as the following expression:

[0124] C net (t)=(P buy (t)×Price buy (t))-(P sell (t)×Price sell (t))

[0125] Among them, C net (t) is the electricity purchase cost or electricity sales income of the microgrid at time t, P buy (t) is the amount of electricity purchased by the microgrid from the main grid at time t, Price buy (t) is the electricity purchase price, P sell (t) is the amount of electricity sold to the main grid at time t, Price sell (t) is the electricity price. It can be understood that this expression determines the total net cost or benefit by calculating the difference between the microgrid's electricity purchase cost and electricity sales revenue at each point in time. If it is a positive value, it means that the microgrid has incurred costs overall, and if it is a negative value, it means that the microgrid has generated benefits overall.

[0126] In this embodiment, considering the constraining role of energy supply and demand factors and cost factors in the operation of the microgrid, the energy supply and demand factors and cost factors can be used to construct the conditional constraints of the energy allocation action in the Markov decision process. Specifically, Figure 4 As shown, step 103 includes the following steps:

[0127] 401. Model training and dynamic update adjustment are performed on sample data describing energy allocation decisions through the Markov decision process, so as to calculate, during the model training process, the transition probability of the microgrid taking energy allocation actions in different time states and jumping to the microgrid taking energy allocation actions in the next time state based on the sample data.

[0128] 402. Calculate, based on the transfer probability, a cumulative reward value assigned to the microgrid for taking an energy allocation action.

[0129] 403. On the basis of satisfying the energy supply and demand balance, the energy control strategy output by the energy control model in a given microgrid environment is continuously adjusted so that the energy distribution action taken by the energy control strategy on the microgrid satisfies the maximum cumulative reward value.

[0130] In this embodiment, the sample data describing the energy allocation decision can be offline data output during the historical operation process of the microgrid. Specifically, before the sample data describing the energy allocation decision is trained and dynamically updated and adjusted through the Markov decision process to obtain the energy control model of the microgrid, the energy strategy data output during the historical operation process of the microgrid is obtained. Here, the energy strategy data includes cost data of energy flow, energy storage data, and action data of energy control. The energy strategy data is recorded in the form of a Markov decision process to obtain sample data describing the energy allocation decision, so that the sample data has the system state and action space required by the Markov decision process.

[0131] In actual application scenarios, considering that the sample data describing the energy allocation decision has uncertainty in its ability to adapt to the environment, in order to improve the stability of the sample data in different microgrid environments, before the sample data describing the energy allocation decision is trained and dynamically updated through the Markov decision process to obtain the energy control model of the microgrid, a resource data pool with energy control experience is constructed to update the sample data describing the energy allocation decision through the resource data pool with energy control experience, so that the sample data is given more energy control scenarios with unknown and complex environments.

[0132] On the one hand, based on the energy trajectory data output by the microgrid during its historical operation, the first resource data pool can be constructed by selecting energy trajectories whose cumulative rewards meet the set conditions. Here, the cumulative rewards that meet the preset conditions are usually trajectory data with a high cumulative reward ranking, which can be the trajectory data of the top 10 cumulative rewards, or the trajectory data with a set cumulative reward. The energy trajectory data is the energy trajectory formed by the energy strategy data continuously output by the energy scheduling from the initial moment to the end moment.

[0133] In this embodiment, the sample data describing the energy allocation decision can be the energy trajectory formed by the energy strategy data continuously output by the energy scheduling from the initial moment to the end moment, for example, the energy trajectory formed by the energy strategy data corresponding to the hour from 8:00 to 18:00 every day. In order to make better use of the sample data describing the energy allocation decision, the resource function set by expert experience can filter the sample data describing the energy allocation decision to obtain the energy trajectory whose cumulative reward meets the set conditions.

[0134] Specifically, the energy trajectory data output by the microgrid during its historical operation can be screened by the expert function p0, thereby generating a first resource data pool with energy control experience. Here the first resource data pool can be defined as the following expression:

[0135]

[0136] Among them, Traj i It is an energy trajectory in the energy trajectory data output by the microgrid during its historical operation, specifically including the energy strategy data from the initial scheduling moment to the end moment. t (Traj i ) is the cumulative reward function, which is used to measure the quality of the energy trajectory. Here, the cumulative reward function can be defined as the following expression:

[0137]

[0138] It is understandable that each r in the energy trajectory data t Represents the immediate reward at time t, through the cumulative reward function Z t , high-quality energy trajectories in the energy trajectory data can be screened out, so as to update the parameters of the energy control model through the high-quality energy trajectories.

[0139] On the other hand, the energy strategy data provided by different experience scenarios can be applied to the Markov decision process. In the Markov decision process, the energy strategy data that meets the state transition conditions are selected to construct the second resource data pool. Here, different experience scenarios correspond to mixing coefficients, and the weights of the energy strategy data provided by the experience scenarios in the second resource data pool are adjusted by the mixing coefficients; then, according to the first resource data pool and the second resource data pool, the sample data describing the energy allocation decision is updated, so that the parameters of the energy control model are adjusted through the updated sample data describing the energy allocation decision.

[0140] Specifically, different experience scenarios provide energy strategy data R rule (s) Policy rules are typically developed by domain experts or grid operators based on their experience and expertise. These policy rules do not directly provide optimal energy allocation actions, but rather provide recommendations or constraints based on system states. For example, when energy prices are high, policy rules such as more power generation, less power purchase, or more power sales may be adopted. Policy rules can generate a candidate action set for each system state. This candidate action set can be defined as follows:

[0141]

[0142] Furthermore, the action candidate set for each system state s is Combining the constraints and state transition rules of the Markov decision process, we can get the second resource data pool Here the second resource data pool can be defined as the following expression:

[0143]

[0144] in, is the reward function set, Is the set of whether the current state is at the end of the scheduling cycle, The next state value set.

[0145] It is understandable that, considering the contribution of different experience scenarios to the update of the energy control model during the sample data training process, a mixing coefficient α can be assigned to different experience scenarios. This mixing coefficient is a fixed constant. By dynamically adjusting the mixing coefficient α, the weight of the energy policy data provided by different experience scenarios can be controlled to optimize the update process of the energy control model. Adjusting the mixing coefficient α can shorten the range of Markov decision making, thereby improving the update performance of the deep deterministic policy gradient agent. The update strategy of this mixing weight can be defined as the following expression:

[0146]

[0147] Among them, β is the learning rate, and Respectively represent data from the first resource pool and the second resource data pool The cumulative return.

[0148] It can be understood that the updated sample data is endowed with high-quality energy trajectory and high-quality energy strategy data of more experience scenarios. In the process of adjusting the parameters of the energy control model through the updated sample data that describes the energy allocation decision, the cumulative reward value of the sample data is calculated using the reward function based on the updated sample data that describes the energy allocation decision; the sample data is sorted from large to small according to the cumulative reward value, and the energy trajectory data and / or energy strategy data corresponding to the sample data whose cumulative reward value is ranked before the preset value are selected to adjust the parameters of the energy control model of the microgrid.

[0149] Specifically, the process of parameter adjustment of the energy control model of the microgrid can be realized based on the preference-guided online deep deterministic policy gradient algorithm, which includes the offline stage, the online stage, the online update and preference guidance stage, and the online update stage of the strategy and value function. This process combines offline learning and the online experience-guided deep deterministic policy gradient algorithm to realize dynamic energy control of the microgrid. By introducing experience scenarios, the adaptability and stability of the energy control strategy are effectively improved, so that the microgrid can continuously optimize energy control decisions in a changing environment, so as to adopt energy allocation actions that are more suitable for the current state, improve the operating efficiency of the microgrid, and thus reduce the operating cost of the microgrid.

[0150] Specifically, during the offline phase, the deep deterministic policy gradient algorithm performs preliminary training on the energy control model based on traditional random sampling and historical data. During this phase, the policy and value network are pre-trained using empirical data from a simulated environment. The primary goal is to initialize model parameters and lay the foundation for subsequent online updates. The offline training policy optimization can be defined as follows:

[0151]

[0152] Among them, Q(s,a||||θ Q ) is the value function, μ(s||||θ μ ) is the strategy function. It should be noted that the training of the energy control model in the offline stage does not rely on expert experience, but learns the basic strategy through autonomous exploration and accumulation of historical data.

[0153] After entering the online stage, the deep deterministic policy gradient algorithm begins to update the energy control model using different experience scenarios, and updates the energy control model from the first resource data pool through policy rules or historical data. and the second resource data pool High-quality energy trajectory and / or energy policy data is screened during the process. This screening process can rank the energy trajectory and / or energy policy data based on the cumulative reward function, and select the top-ranked energy trajectory and / or energy policy data as empirical data to guide online updates. It should be noted that the resource data pool provides rich domain knowledge and high-quality operation recommendations, which can effectively improve the decision-making quality of energy control models in real-world environments.

[0154] In the specific online update and preference guidance phase, in each state S, the deep deterministic policy gradient algorithm uses the first mixed resource data pool and the second resource data pool The strategy rules obtained in the above code generate the action candidate set It can be defined as the following expression:

[0155]

[0156] It's easy to understand that this action selection process combines the policy network of a deep deterministic policy gradient algorithm with guidance from a resource data pool. The introduction of the resource data pool helps the energy control model quickly adapt to environmental changes during online learning, enabling it to make more robust energy control decisions under uncertainty. This allows the energy control model to reduce ineffective exploration and more quickly approach the optimal energy control strategy.

[0157] Specifically, during the online update phase of the policy and value functions, at each time step t, the deep deterministic policy gradient algorithm updates the policy and value networks by combining the empirical data obtained online. The update process is expressed as follows:

[0158]

[0159] Among them, L(θ Q ) is the loss function of the value network Q, which can be defined as the following expression in the form of mean square error:

[0160] L(θ Q )=E (s,a,r,s′) [(r+γQ′(s′,μ′(s′∣θ μ′ ))-Q(s,a|θ Q )) 2 ]

[0161] Here, Q′ and μ′ represent the target value network and target policy network, respectively. By combining empirical data from a resource database for online updates, the deep deterministic policy gradient algorithm can dynamically adjust the energy control policy output by the energy control model, improving its adaptability and stability, enabling the energy control model to achieve the desired energy control effect.

[0162] In the online update phase of model training, an embodiment of the present invention uses a deep deterministic policy gradient algorithm for policy optimization, so that the energy control model can extract high-quality empirical data from the hybrid resource data pool for guidance according to the current state at each time step, and output it in combination with the policy network of the deep deterministic policy gradient algorithm to control the microgrid to take the optimal energy allocation action. Specifically, during the policy update process, the energy control model iteratively adjusts the parameters of the policy and value network through the policy gradient algorithm, so that the policy can dynamically adapt to changes in the environment. The core of this policy update process is to combine high-return empirical suggestions in the resource data pool to effectively reduce the uncertainty and instability problems in online learning, improve energy control efficiency and the accuracy of energy control decisions. After experience guidance and online policy updates, the energy control model can continuously optimize the energy control strategy of the microgrid in a changing environment, achieving cost reduction and resource optimization.

[0163] In actual application scenarios, taking the microgrid action time interval of 1 hour as an example, real data is used for testing, and four 24-hour scenarios are randomly selected as test scenarios. At each time step t, the energy control model outputs the corresponding energy control decision based on the current state of the microgrid. This process does not use prediction information.

[0164] In the aforementioned test scenarios, energy control models trained for four different scenarios were applied to solve energy control problems in microgrid processes. It's important to note that these tests don't involve any prior knowledge of the future. Therefore, the energy control model doesn't require forecast data for future scenarios during calculations. Instead, it makes real-time decisions based on the current state, which is more suitable for practical applications.

[0165] Table 1 shows the microgrid costs for the four different test scenarios described above. This table shows the cost data for model training and updating using different deep deterministic policy gradient algorithms in different microgrid scenarios, with the lowest cost data shown in bold. Specifically, it can be observed that compared to the offline deep deterministic policy gradient algorithm, the online-updating deep deterministic policy gradient algorithm does improve the performance of the energy control model in some scenarios. For example, in test scenario 2, the cost is reduced compared to that in scenario 1. However, in test scenarios 1, 3, and 4, the online-updating deep deterministic policy gradient algorithm has a higher cost, indicating that simple online updates can sometimes lead to worse results. In contrast, the preference-guided deep deterministic policy gradient algorithm maintains a steady performance improvement, achieving lower costs in all four scenarios compared to the offline deep deterministic policy gradient algorithm.

[0166] Table 1. Cost comparison of different methods in different microgrid scenarios

[0167]

[0168] Due to the nature of offline deep deterministic policy gradient algorithms, their cost remains constant. However, compared to preference-guided deep deterministic policy gradient algorithms, offline deep deterministic policy gradient algorithms exhibit significant cost differences. This is primarily because offline learning relies on historical data for training, which can introduce sample bias and affect the generalization of the energy control model in new environments. Furthermore, when equipped with real-time update capabilities, online deep deterministic policy gradient algorithms experience significant cost fluctuations. While achieving performance improvements in some cases, they also occasionally underperform. This fluctuation may stem from changes in data distribution during online learning, leading to inconsistencies between offline and online data distributions, which in turn affect the learning of the energy control model. In contrast, preference-guided deep deterministic policy gradient algorithms exhibit significant cost reductions during the online update phase while maintaining lower stability. This demonstrates that by combining empirical knowledge with online learning, preference-guided deep deterministic policy gradient algorithms ensure a more robust and efficient model update process. They demonstrate superior stability and lower cost across multiple test scenarios, validating the effectiveness of leveraging empirical knowledge to guide online updates in reinforcement learning.

[0169] It's understandable that the energy policy model based on the preference-guided deep deterministic policy gradient algorithm can implement an energy control strategy that sells electricity during periods of high electricity prices and a purchase strategy during periods of low electricity prices, effectively reducing the operating costs of the microgrid. Furthermore, this energy control strategy can rationally utilize the energy storage system by controlling its charging and discharging to meet load demand. For example, during periods of low electricity prices, the energy control strategy increases charging, while during periods of high electricity prices, it increases discharging or even sells electricity, maximizing revenue. This energy control strategy not only optimizes the economic benefits of the microgrid but also improves the overall operational efficiency of the system. Furthermore, it can rationally allocate the proportion of renewable energy based on real-time load demand, activating controllable generators when necessary to ensure energy supply and demand balance and stable system operation. This significantly reduces the operating costs of the microgrid while meeting load demand and improving its operational efficiency.

[0170] Further, as Figure 1-4 The specific implementation of the method, the embodiment of the present application provides an energy control device for the operation process of a microgrid, such as Figure 5 As shown, the device includes: a determination unit 51, a modeling unit 52, a training unit 53, and a control unit 54.

[0171] a determination unit 51, configured to determine a plurality of control factors affecting energy distribution during operation of the microgrid;

[0172] A modeling unit 52 is configured to model the energy control process of the microgrid operation as a Markov decision process of a multi-tuple structure according to the multiple control factors;

[0173] A training unit 53 is configured to perform model training and dynamic update adjustment on sample data describing energy allocation decisions through the Markov decision process to obtain an energy control model of the microgrid, wherein the energy control model is configured to output a current energy control strategy in a given microgrid environment;

[0174] The control unit 54 is configured to take energy distribution actions on the microgrid according to the energy control strategy of the current time state, so as to maintain energy balance during the operation of the microgrid.

[0175] The energy control device for the microgrid operation process provided by the embodiments of the present invention, compared to the existing method of controlling microgrid energy through the interrelationships of the various components of the microgrid, determines multiple control factors that affect energy distribution during microgrid operation, models the microgrid energy control process as a multi-tuple Markov decision process based on these multiple control factors, and uses the Markov decision process to train and dynamically update sample data describing energy distribution decisions to obtain an energy control model for the microgrid. This energy control model is used to output a current energy control strategy in a given microgrid environment and take energy distribution actions for the microgrid based on the current energy control strategy to maintain energy balance during microgrid operation. The entire process models the microgrid energy control process as a multi-tuple Markov decision process, achieves dynamic control of energy during microgrid operation, improves the stability and adaptability of the energy control strategy, and enables the microgrid to continuously optimize its energy control strategy in a changing environment, greatly meeting the energy control needs of the microgrid.

[0176] In a specific application scenario, the determining unit is specifically configured to:

[0177] According to the changes in electricity usage during the operation of the microgrid, determine the supply control factors and demand control factors that affect energy distribution during the operation of the microgrid;

[0178] On the basis of the supply control factors and the demand control factors, according to the change of energy prices during the operation of the microgrid, the cost control factors affecting energy distribution during the operation of the microgrid are determined.

[0179] In a specific application scenario, the device further includes:

[0180] an acquisition unit, configured to acquire energy strategy data output by the microgrid during a historical operation process before performing model training and dynamic update adjustment on the sample data describing the energy allocation decision through the Markov decision process to obtain an energy control model of the microgrid, wherein the energy strategy data includes cost data of energy flow, energy storage data, and action data of energy control;

[0181] The recording unit is used to record the energy strategy data in the form of a Markov decision process to obtain sample data describing the energy allocation decision.

[0182] In a specific application scenario, the modeling unit is specifically used to:

[0183] According to the energy control process of microgrid operation, the energy state data describing the microgrid operating environment at different time states are used as the system state of the Markov decision process;

[0184] According to the energy control process of microgrid operation, the energy action data of microgrid in different time states are used as the action space of Markov decision process;

[0185] According to the energy control process of the microgrid, the reward function of the Markov decision process is constructed using the reward / penalty that describes the energy allocation actions taken by the microgrid.

[0186] According to the energy control process of the microgrid, the probability of transferring the energy allocation action of the current state to the energy allocation action of the next state is used as the transition probability of the Markov decision process;

[0187] The system state, the action space, the reward function and the transition probability are used as a multi-group structure to model the energy control process of the microgrid operation as a Markov decision process with a multi-group structure.

[0188] In a specific application scenario, the training unit is specifically used to:

[0189] The Markov decision process is used to perform model training and dynamic update adjustment on sample data describing energy allocation decisions, so as to calculate, during the model training process, the transition probability of the microgrid taking energy allocation actions in different time states and jumping to the energy allocation actions taken by the microgrid in the next time state according to the sample data;

[0190] Calculate the cumulative reward value given to the microgrid for taking the energy allocation action according to the transition probability;

[0191] On the basis of satisfying the balance of energy supply and demand, the energy control strategy output by the energy control model in a given microgrid environment is continuously adjusted so that the energy allocation action taken by the energy control strategy on the microgrid can satisfy the maximum cumulative reward value.

[0192] In a specific application scenario, the device further includes:

[0193] A first construction unit is configured to, before performing model training and dynamic updating and adjustment on the sample data describing the energy allocation decision through the Markov decision process to obtain the energy control model of the microgrid, select energy trajectories whose cumulative rewards meet set conditions based on energy trajectory data output by the microgrid during historical operation to construct a first resource data pool, wherein the energy trajectory data is an energy trajectory formed by energy strategy data continuously output by the energy dispatcher from an initial moment to an end moment;

[0194] A second construction unit is configured to apply energy strategy data provided by different experience scenarios to a Markov decision process, select energy strategy data that meets the state transition condition in the Markov decision process to construct a second resource data pool, wherein the different experience scenarios correspond to mixing coefficients, and the weights of the energy strategy data provided by the experience scenarios in the second resource data pool are adjusted by the mixing coefficients;

[0195] An updating unit is used to update the sample data describing the energy allocation decision according to the first resource data pool and the second resource data pool, so as to adjust the parameters of the energy control model through the updated sample data describing the energy allocation decision.

[0196] In a specific application scenario, the updating unit is specifically used to:

[0197] Based on the updated sample data describing the energy allocation decision, the reward function is used to calculate the cumulative reward value of the sample data;

[0198] The sample data are sorted from large to small according to the cumulative reward value, and the energy trajectory data and / or energy strategy data corresponding to the sample data whose cumulative reward value is sorted before the preset value are selected to adjust the parameters of the energy control model of the microgrid.

[0199] It should be noted that for other corresponding descriptions of the functional units involved in the energy control device for the microgrid operation process provided in this embodiment, please refer to Figure 1-Figure 4 The corresponding description in will not be repeated here.

[0200] Based on the above Figure 1-Figure 4 The method shown in FIG. 1 is a method for performing the above-mentioned operation. Accordingly, the embodiment of the present application further provides a storage medium on which a computer program is stored. When the program is executed by a processor, the above-mentioned operation is performed. Figure 1-Figure 4 The energy control method of the microgrid operation process is shown.

[0201] Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each implementation scenario of the present application.

[0202] Based on the above Figure 1-Figure 4 The method shown, and Figure 5In order to achieve the above-mentioned purpose, the embodiment of the present application further provides a physical device for energy control during the operation of a microgrid, which can be a computer, a smart phone, a tablet computer, a smart watch, a server, or a network device, etc. The physical device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to achieve the above-mentioned Figures 1-6 The energy control method of the microgrid operation process is shown.

[0203] Optionally, the physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a Wi-Fi module, etc. The user interface may include a display, an input unit such as a keyboard, etc., and the optional user interface may also include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Wi-Fi interface), etc.

[0204] In an exemplary embodiment, see Figure 6 The physical device includes a communication bus, a processor, a memory, and a communication interface. It may also include an input / output interface and a display device. The various functional units can communicate with each other via the bus. The memory stores a computer program, and the processor is configured to execute the program stored in the memory and implement the energy control method for the microgrid operation process described in the above embodiment.

[0205] Those skilled in the art will understand that the physical device structure for energy control during the operation of a microgrid provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or a combination of certain components, or different component arrangements.

[0206] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the physical device responsible for energy control during the microgrid's operation, supporting the execution of information processing programs and other software and / or programs. The network communication module is used to enable communication between components within the storage medium and with other hardware and software within the physical information processing device.

[0207] Through the description of the above implementation methods, those skilled in the art can clearly understand that this application can be implemented by means of software plus the necessary general hardware platform, or by hardware. By applying the technical solution of this application, compared with the current existing methods, this application models the energy control process of the microgrid operation as a Markov decision process with a multi-tuple structure, realizes dynamic control of the energy during the operation of the microgrid, improves the stability and adaptability of the energy control strategy, enables the microgrid to continuously optimize the energy control strategy in a changing environment, and greatly meets the energy control needs of the microgrid.

[0208] Those skilled in the art will understand that the accompanying drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the accompanying drawings are not necessarily required to implement the present application. Those skilled in the art will understand that the modules in the devices in the implementation scenario can be distributed in the devices of the implementation scenario according to the implementation scenario description, or can be changed accordingly and located in one or more devices different from the implementation scenario. The modules of the above-mentioned implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0209] The serial numbers of the above application are for descriptive purposes only and do not represent the advantages or disadvantages of the implementation scenarios. The above disclosure only discloses several specific implementation scenarios of the present application, but the present application is not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the scope of protection of the present application.

Claims

1. A method for energy control during microgrid operation, characterized in that: include: Identify multiple control factors that affect energy distribution during microgrid operation; Modeling an energy control process of the microgrid operation as a Markov decision process of a multi-tuple structure according to the multiple control factors, wherein the energy control of the microgrid operation includes controlling the energy flow and consumption of each energy resource in the microgrid system; The Markov decision process is used to perform model training and dynamic update adjustment on sample data describing energy allocation decisions to obtain an energy control model of the microgrid, wherein the energy control model is used to output a current energy control strategy in a given microgrid environment; Taking energy distribution actions on the microgrid according to the energy control strategy of the current tense to maintain energy balance during the operation of the microgrid; Before performing model training and dynamic updating and adjustment on the sample data describing the energy allocation decision through the Markov decision process to obtain the energy control model of the microgrid, a first resource data pool is constructed by selecting energy trajectories whose cumulative rewards meet set conditions based on energy trajectory data output by the microgrid during historical operation, wherein the energy trajectory data is an energy trajectory formed by the energy strategy data continuously output by the energy dispatcher from the initial moment to the end moment; Applying energy strategy data provided by different experience scenarios to a Markov decision process, selecting energy strategy data that meets the state transition conditions in the Markov decision process to construct a second resource data pool, wherein the different experience scenarios correspond to mixing coefficients, and the weights of the energy strategy data provided by the experience scenarios in the second resource data pool are adjusted by the mixing coefficients; updating the sample data describing the energy allocation decision according to the first resource data pool and the second resource data pool, so as to adjust the parameters of the energy control model by using the updated sample data describing the energy allocation decision; The parameter adjustment of the energy control model by the updated sample data describing the energy allocation decision includes: calculating the cumulative reward value of the sample data using a reward function based on the updated sample data describing the energy allocation decision; The sample data are sorted from large to small according to the cumulative reward value, and the energy trajectory data and / or energy strategy data corresponding to the sample data whose cumulative reward value is sorted before the preset value are selected to adjust the parameters of the energy control model of the microgrid.

2. The method according to claim 1, characterized in that The determination of multiple control factors affecting energy distribution during the operation of the microgrid includes: According to the changes in electricity usage during the operation of the microgrid, determine the supply control factors and demand control factors that affect energy distribution during the operation of the microgrid; On the basis of the supply control factors and the demand control factors, according to the change of energy prices during the operation of the microgrid, the cost control factors affecting energy distribution during the operation of the microgrid are determined.

3. The method according to claim 1, characterized in that Before performing model training and dynamic updating and adjustment on the sample data describing the energy allocation decision through the Markov decision process to obtain the energy control model of the microgrid, the method further includes: Obtaining energy strategy data output by the microgrid during historical operation, wherein the energy strategy data includes cost data of energy flow, energy storage data, and action data of energy control; The energy strategy data is recorded in the form of a Markov decision process to obtain sample data describing the energy allocation decision.

4. The method according to claim 1, wherein The energy control process of the microgrid operation is modeled as a Markov decision process of a multi-tuple structure according to the multiple control factors, including: According to the energy control process of microgrid operation, the energy state data describing the microgrid operating environment at different time states are used as the system state of the Markov decision process; According to the energy control process of microgrid operation, the energy action data of microgrid in different time states are used as the action space of Markov decision process; According to the energy control process of the microgrid, the reward function of the Markov decision process is constructed using the reward / penalty that describes the energy allocation actions taken by the microgrid. According to the energy control process of the microgrid, the probability of transferring the energy allocation action of the current state to the energy allocation action of the next state is used as the transition probability of the Markov decision process; The system state, the action space, the reward function and the transition probability are used as a multi-group structure to model the energy control process of the microgrid operation as a Markov decision process with a multi-group structure.

5. The method according to claim 1, characterized in that The energy control model of the microgrid is obtained by performing model training and dynamic updating and adjustment on the sample data describing the energy allocation decision through the Markov decision process, including: The Markov decision process is used to perform model training and dynamic update adjustment on sample data describing energy allocation decisions, so as to calculate, during the model training process, the transition probability of the microgrid taking energy allocation actions in different time states and jumping to the energy allocation actions taken by the microgrid in the next time state according to the sample data; Calculate the cumulative reward value given to the microgrid for taking the energy allocation action according to the transition probability; On the basis of satisfying the balance of energy supply and demand, the energy control strategy output by the energy control model in a given microgrid environment is continuously adjusted so that the energy allocation action taken by the energy control strategy on the microgrid can satisfy the maximum cumulative reward value.

6. An energy control device for a microgrid operation process, characterized in that: include: a determination unit, configured to determine a plurality of control factors affecting energy distribution during operation of the microgrid; a modeling unit, configured to model an energy control process of the microgrid operation as a Markov decision process of a multi-tuple structure according to the plurality of control factors, wherein the energy control of the microgrid operation includes controlling the energy flow and consumption of each energy resource in the microgrid system; A training unit is configured to perform model training and dynamic update adjustment on sample data describing energy allocation decisions through the Markov decision process to obtain an energy control model of the microgrid, wherein the energy control model is configured to output a current energy control strategy in a given microgrid environment; A control unit, configured to take energy distribution actions on the microgrid according to the energy control strategy of the current time state, so as to maintain energy balance during operation of the microgrid; The device also includes: a first construction unit, which is used to select energy trajectories that have cumulative rewards and meet set conditions according to energy trajectory data output by the microgrid during historical operation to construct a first resource data pool before performing model training and dynamic updating and adjustment on the sample data describing the energy allocation decision through the Markov decision process to obtain the energy control model of the microgrid, wherein the energy trajectory data is an energy trajectory formed by energy strategy data continuously output by the energy dispatcher from the initial moment to the end moment; a second construction unit, which is used to apply energy strategy data provided by different experience scenarios to the Markov decision process, and select energy strategy data that meets state transition conditions in the Markov decision process to construct a second resource data pool, wherein the different experience scenarios correspond to mixing coefficients, and the weights of the energy strategy data provided by the experience scenarios in the second resource data pool are adjusted by the mixing coefficients; an updating unit, which is used to update the sample data describing the energy allocation decision according to the first resource data pool and the second resource data pool, so as to adjust the parameters of the energy control model through the updated sample data describing the energy allocation decision; The update unit is specifically used to calculate the cumulative reward value of the sample data using a reward function based on the sample data describing the energy allocation decision after the update; sort the sample data from large to small according to the cumulative reward value, and select the energy trajectory data and / or energy strategy data corresponding to the sample data whose cumulative reward value is ranked before the preset value to adjust the parameters of the energy control model of the microgrid.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the energy control method for the microgrid operation process described in any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the energy control method for the microgrid operation process described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Micro-grid energy storage optimization scheduling method based on deep reinforcement learning

    CN117833285A