Energy Internet Optimization Operation Method and Related Devices

By using generative adversarial networks and alliance game training optimized scheduling reinforcement learning models in the energy Internet, the complexity and uncertainty problems of distributed energy systems are solved, more efficient optimization scheduling and operation control are achieved, and the robustness and reliability of the system are improved.

CN118657264BActive Publication Date: 2025-07-08CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411151372.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-07-08
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

The increase in the number of autonomous control points of distributed energy in the energy Internet has led to complex system structure, and traditional mechanism models and distributed optimization methods are difficult to meet the requirements of optimization scheduling and operation control, especially when facing uncertainty in the output and load of new energy, the existing methods are difficult to adapt.

Method used

Generative adversarial networks are used to simulate the transfer probability of uncertain environment quantities, and combined with alliance games, the optimization and scheduling reinforcement learning model is trained to improve the prediction and decision-making capabilities of the microgrid, and to simulate the transfer probability of uncertain environment quantities by building generative adversarial networks to enhance the robustness of the model.

Benefits of technology

It improves the robustness of the energy Internet in the face of uncertain environments, improves the energy utilization efficiency and system security and reliability, and adapts to the high uncertainty of distributed energy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118657264B_ABST
    Figure CN118657264B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of energy Internet, and discloses an energy Internet optimal operation method and related device, including: obtaining status data of each microgrid in the energy Internet; according to the status data of each microgrid, respectively obtaining the optimal scheduling decision of each microgrid through a pre-set policy network model of each microgrid; wherein, the pre-set policy network model of each microgrid is obtained by iteratively performing the following steps: generating the transition probability of the uncertainty environmental quantity of each microgrid through a pre-trained generative adversarial network of each microgrid, and obtaining the predicted environment of each microgrid based on the transition probability, and training the optimal scheduling reinforcement learning model of each microgrid by means of coalition game according to the predicted environment of each microgrid. By constructing a generative adversarial network to simulate the transition probability of the uncertainty environmental quantity, it can better adapt to the highly uncertain environment of the microgrid, improve the robustness of the optimal scheduling decision, and ensure the optimal operation effect of the energy Internet.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of energy Internet, and relates to an energy Internet optimal operation method and related devices. Background Art

[0002] The energy Internet integrates multiple stages such as energy production, transmission, storage, and consumption. Optimizing operation involves various links of source-network-load-storage, such as primary energy, secondary energy, and energy transportation networks. The various types of energy are not simply connected and superimposed, but under the support of an information network composed of safe, reliable, advanced sensing and communication, through efficient and accurate intelligent control strategies, scientific allocation of energy is carried out. While meeting the diverse energy demand of users, the complementary and collaborative characteristics of various types of energy are utilized to improve energy utilization efficiency and renewable energy consumption capacity, and enhance the overall security, reliability, and flexibility of the system. The energy Internet will change the traditional top-down centralized decision-making mode of the traditional energy network. Each energy entity in the system will independently manage its own energy production, consumption, and transactions, realizing autonomous energy management and system decision-making.

[0003] With the gradual access of distributed energy, the number of autonomous control points has increased significantly. The energy Internet has evolved into a huge-dimensional system with complex structure, numerous devices, and complex technologies, with typical non-linear stochastic characteristics and multi-scale dynamic characteristics, greatly increasing the difficulty of optimal operation. Traditional mechanism model analysis and distributed optimization methods are no longer able to meet the requirements of optimal scheduling, analysis and evaluation, and operation control. This requires full interconnection and sharing of energy and information among various entities through advanced sensing, measurement technologies, and edge computing technologies, and realizing friendly interaction between source-network-load-storage and full consumption of a high proportion of renewable energy through artificial intelligence algorithms. To address the above problems, in recent years, research on the operation control and optimization decision of the energy Internet has received extensive attention.

[0004] For example, Chinese Patent Application CN118017492A discloses an active distribution network optimal scheduling method. By obtaining the predicted values of photovoltaic power output and load in each area of the active distribution network, according to the predicted values of photovoltaic power output and load in each area, a preset distributed optimization scheduling model based on coalition game is called to obtain the optimal output values of each controllable device in each area. Control instructions for each controllable device in each area are generated according to the optimal output values of each controllable device in each area and sent to each controllable device in each area; among them, the distributed optimization scheduling model based on coalition game includes a coalition game layer and each area agent interacting with the coalition game layer. Its distributed collaborative optimization is realized based on non-critical information sharing based on the game mechanism in each area, avoiding the centralized collection and processing of a large amount of data and greatly improving the decision-making speed. However, this method only considers the economic and safety issues of real-time decision-making and cannot well adapt to the uncertainty caused by random fluctuations of source and load. Summary of the Invention

[0005] The object of the present invention is to overcome the above-mentioned disadvantages of the prior art and provide an energy Internet optimized operation method and related devices.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In the first aspect of the present invention, an energy Internet optimized operation method is provided, including: obtaining the state data of each microgrid in the energy Internet; according to the state data of each microgrid, respectively obtaining the optimized scheduling decisions of each microgrid through the preset policy network model of each microgrid; wherein, the preset policy network model of each microgrid is obtained through the following method: obtaining the optimized scheduling reinforcement learning model of each microgrid; training step: generating the transition probability of the uncertainty environmental quantity of each microgrid through the pre-trained generative adversarial network of each microgrid, and obtaining the predicted environment of each microgrid based on the transition probability, and training the optimized scheduling reinforcement learning model of each microgrid by means of coalition game according to the predicted environment of each microgrid; iteratively training the step until the optimized scheduling reinforcement learning model of each microgrid is trained, and obtaining the policy network in the optimized scheduling reinforcement learning model of each microgrid, and obtaining the preset policy network model of each microgrid.

[0008] Optionally, the generative adversarial network is constructed based on the Wasserstein generative adversarial network with gradient penalty.

[0009] Optionally, the pre-trained generative adversarial network of each microgrid satisfies the following generation constraints: the transition probability of the uncertainty environmental quantity of each microgrid generated by the pre-trained generative adversarial network of each microgrid respectively falls within the probability distribution fuzzy set constructed with the empirical probability distribution of the uncertainty environmental quantity as the center and the preset Wasserstein distance as the radius of each microgrid.

[0010] Optionally, in the iterative training step, after each training step is completed, the pre-trained generative adversarial network of each microgrid is updated with the goal of minimizing the reward value function of the optimized scheduling reinforcement learning model of each microgrid.

[0011] Optionally, it further includes: constructing an initial generative adversarial network for each microgrid with the state of the optimized scheduling reinforcement learning model of each microgrid as the input and the transition probability of the uncertainty environmental quantity of each microgrid as the output; respectively training the initial generative adversarial network of each microgrid according to the historical data of each microgrid to obtain the pre-trained generative adversarial network of each microgrid.

[0012] Optionally, training the optimal scheduling reinforcement learning model of each microgrid by means of coalition game according to the predicted environment of each microgrid includes: the optimal scheduling reinforcement learning models of each microgrid interact with the predicted environment of each microgrid respectively to obtain the historical experience of each microgrid; based on the net power demand data in the historical experience of each microgrid, the coalition allocation benefits and unit electricity prices of each microgrid are obtained based on the market intermediate price mechanism and the Shapley value method; according to the coalition allocation benefits and unit electricity prices of each microgrid, the reward function values of the optimal scheduling reinforcement learning models of each microgrid are obtained; through the evaluation network of the optimal scheduling reinforcement learning model of each microgrid, according to the reward function values of the optimal scheduling reinforcement learning model of each microgrid and the historical data stored in the experience replay pool in the historical experience, the state value function of the optimal scheduling reinforcement learning model of each microgrid is obtained, and the generalized advantage estimation is carried out according to the state value function of the optimal scheduling reinforcement learning model of each microgrid to obtain the advantage function of the optimal scheduling reinforcement learning model of each microgrid; according to the state value function of the optimal scheduling reinforcement learning model of each microgrid, the evaluation network of the optimal scheduling reinforcement learning model of each microgrid is updated; according to the advantage function of the optimal scheduling reinforcement learning model of each microgrid, the policy network of the optimal scheduling reinforcement learning model of each microgrid is updated by using the constrained policy optimization method.

[0013] Optionally, the state data includes the active power output of the gas turbine and the energy of the energy storage system at the previous moment, as well as the new energy output and load at several previous moments; the uncertainty environmental quantity includes the new energy output and load; the optimal scheduling decision of each microgrid includes the active power output of the gas turbine of each microgrid, the reactive power output of the gas turbine, the active power output of the energy storage system and the reactive power output of the energy storage system.

[0014] In the second aspect of the present invention, an energy Internet optimal operation system is provided, including: a data acquisition module, configured to acquire the state data of each microgrid in the energy Internet; a decision module, configured to obtain the optimal scheduling decision of each microgrid respectively through the preset policy network model of each microgrid according to the state data of each microgrid; wherein, the preset policy network model of each microgrid is obtained by the following method: obtaining the optimal scheduling reinforcement learning model of each microgrid; training step: generating the transition probability of the uncertainty environmental quantity of each microgrid through the pre-trained generative adversarial network of each microgrid, obtaining the predicted environment of each microgrid based on the transition probability, and training the optimal scheduling reinforcement learning model of each microgrid by means of coalition game according to the predicted environment of each microgrid; iteratively training the step until the optimal scheduling reinforcement learning model of each microgrid is trained, and obtaining the policy network in the optimal scheduling reinforcement learning model of each microgrid to obtain the preset policy network model of each microgrid.

[0015] In the third aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method for optimizing the operation of the energy Internet are implemented.

[0016] In the fourth aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for optimizing the operation of the energy Internet are implemented.

[0017] Compared with the prior art, the present invention has the following beneficial effects:

[0018] In the method for optimizing the operation of the energy Internet of the present invention, according to the state data of each microgrid, the optimized scheduling decisions of each microgrid are obtained respectively through the pre-set policy network models of each microgrid, and then the optimized operation of the energy Internet is realized based on the optimized scheduling decisions of each microgrid. Among them, the pre-set policy network models of each microgrid adopt the policy networks in the optimized scheduling reinforcement learning models trained by each microgrid. And in the training process of the optimized scheduling reinforcement learning model, a generative adversarial network is introduced to generate the transition probabilities of the uncertain environmental quantities of each microgrid, and the predicted environment of each microgrid is obtained based on the transition probabilities. Then, according to the predicted environment of each microgrid, the optimized scheduling reinforcement learning model of each microgrid is trained by means of coalition game. Aiming at the characteristics of the uncertain environmental quantities, a generative adversarial network is constructed to simulate the transition probabilities of the uncertain environmental quantities. Compared with the current method of defaulting the state transition probability of reinforcement learning to 100%, the pre-set policy network models of each microgrid can better adapt to the highly uncertain environment of the microgrid, improve the robustness of the optimized scheduling decision, and ensure the optimization effect of the energy Internet operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of the method for optimizing the operation of the energy Internet according to an embodiment of the present invention.

[0020] Figure 2 It is a schematic diagram of a multi-agent cooperation framework according to an embodiment of the present invention.

[0021] Figure 3 It is a schematic diagram of the training framework of the optimized scheduling reinforcement learning model according to an embodiment of the present invention.

[0022] Figure 4 It is a block diagram of the system structure of the energy Internet optimized operation according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] The present invention will be further described in detail below in conjunction with the accompanying drawings:

[0026] See Figure 1 , in an embodiment of the present invention, an energy Internet optimal operation method is provided. By constructing a generative adversarial network to simulate the transition probability of uncertain environmental quantities, the robustness of the optimal scheduling decision is improved.

[0027] Specifically, the energy Internet optimal operation method of the present invention includes the following steps:

[0028] S1: Obtain the state data of each microgrid in the energy Internet.

[0029] S2: According to the state data of each microgrid, respectively obtain the optimal scheduling decision of each microgrid through the preset policy network model of each microgrid.

[0030] Among them, the preset policy network model of each microgrid is obtained through the following method:

[0031] S11: Obtain the optimal scheduling reinforcement learning model of each microgrid.

[0032] S12: Training step: Generate the transition probability of the uncertain environmental quantity of each microgrid through the pre-trained generative adversarial network of each microgrid, and obtain the predicted environment of each microgrid based on the transition probability, and train the optimal scheduling reinforcement learning model of each microgrid by means of coalition game according to the predicted environment of each microgrid.

[0033] S13: Iteratively train the steps until the training of the optimal scheduling reinforcement learning model for each microgrid is completed, and obtain the policy network in the optimal scheduling reinforcement learning model of each microgrid, thereby obtaining the preset policy network model for each microgrid.

[0034] In the energy Internet optimal operation method of the present invention, according to the state data of each microgrid, the optimal scheduling decisions of each microgrid are obtained respectively through the preset policy network model of each microgrid, and then the optimal operation of the energy Internet is realized based on the optimal scheduling decisions of each microgrid. Among them, the preset policy network model of each microgrid adopts the policy network in the optimal scheduling reinforcement learning model trained for each microgrid. And during the training process of the optimal scheduling reinforcement learning model, a generative adversarial network is introduced to generate the transition probability of the uncertain environmental quantity of each microgrid, and the predicted environment of each microgrid is obtained based on the transition probability. Then, the optimal scheduling reinforcement learning model of each microgrid is trained by means of coalition game according to the predicted environment of each microgrid. In view of the characteristics of the uncertain environmental quantity, by constructing a generative adversarial network to simulate the transition probability of the uncertain environmental quantity, compared with the current method of defaulting the state transition probability of reinforcement learning to 100%, the preset policy network model of each microgrid can better adapt to the highly uncertain environment of the microgrid, improve the robustness of the optimal scheduling decision, and ensure the optimal operation effect of the energy Internet.

[0035] Currently, the reinforcement learning methods applied to the optimal scheduling of power systems basically train the model with fixed new energy output and load, that is, the state transition probability of default reinforcement learning is 100%. However, this is inappropriate because the new energy output and load have the characteristics of high uncertainty. Once it is assumed that the new energy output and load are fixed or follow a certain fixed probability distribution, it will lead to errors in the output of the finally obtained model.

[0036] Therefore, in the energy Internet optimal operation method of the present invention, a generative adversarial network is constructed to simulate the transition probability of these uncertain environmental quantities of the microgrid, generally the new energy output and load. Then, based on the transition probability of the uncertain environmental quantity generated by the generative adversarial network, the predicted environment of the microgrid is determined, and this predicted environment is used as the training basis for the optimal scheduling reinforcement learning model of each microgrid to improve the robustness of the model.

[0037] In a possible implementation manner, the energy Internet optimal operation method further includes: taking the state of the optimal scheduling reinforcement learning model of each microgrid as the input and the transition probability of the uncertain environmental quantity of each microgrid as the output, constructing the initial generative adversarial network of each microgrid; respectively training the initial generative adversarial network of each microgrid according to the historical data of each microgrid to obtain the pre-trained generative adversarial network of each microgrid.

[0038] Specifically, the generative adversarial network consists of two parts: a generator and a discriminator. The input of the generator is random noise that conforms to a certain probability distribution. By learning the distribution pattern, new sample data is generated. The discriminator can be regarded as a binary classifier. Its inputs are real sample data and generated data respectively, which are used to distinguish the authenticity of the data. During the training and optimization process, a dynamic game process will form between the generator and the discriminator. The generator will continuously improve its ability to pass off the fake as real to deceive the discriminator, while the discriminator will continuously improve its ability to identify the true and false. The two will continuously improve the performance of their respective networks through adversarial training, and finally achieve Nash equilibrium.

[0039] The initial generative adversarial network takes the state of the optimal scheduling reinforcement learning model of each microgrid as the input to achieve cooperative training with the optimal scheduling reinforcement learning model. It takes the transition probability of the uncertainty environmental quantity of the microgrid as the output to simulate the transition probability of the uncertainty environmental quantity of the microgrid. Essentially, it predicts the probability distribution of the uncertainty environmental quantity at the next moment by fixing the state of the uncertainty environmental quantity at the previous moment, and then realizes the state prediction of the uncertainty environmental quantity at the next moment by sampling the probability distribution.

[0040] During the process of training the initial generative adversarial network of each microgrid, first calculate the empirical probability distribution of the uncertainty environmental quantity according to the historical data of the microgrid, and use the empirical probability distribution of the uncertainty environmental quantity as the training reference to make the transition probability of the uncertainty environmental quantity generated by the trained initial generative adversarial network as close as possible to the empirical probability distribution of the uncertainty environmental quantity.

[0041] In a possible implementation manner, there are several pre-trained generative adversarial networks for the microgrid, corresponding to the respective uncertainty environmental quantities of the microgrid.

[0042] Specifically, considering that the correlation between different uncertainty environmental quantities is not necessarily at a high level, different generative adversarial networks can be set for different uncertainty environmental quantities and trained separately. For example, in this implementation manner, different generative adversarial networks are set for new energy output and load.

[0043] Optionally, if there are several uncertainty environmental quantities with high correlation, the same generative adversarial network can also be set to simulate the corresponding transition probability.

[0044] In a possible implementation manner, the generative adversarial network is constructed based on the Wasserstein generative adversarial network with gradient penalty.

[0045] Specifically, WGAN (Wasserstein Generative Adversarial Networks) is a variant of GAN (Generative Adversarial Networks). By introducing the Wasserstein distance (Earth Mover's Distance) as the objective function, WGAN has significantly improved in terms of stability and the quality of generated samples compared to traditional GAN. However, the weight clipping method adopted in the training process of WGAN has some defects, such as the weakening of the model's modeling ability, vanishing or exploding gradients, and other problems.

[0046] WGAN-GP (WGAN with Gradient Penalty) is an improved version of WGAN. By introducing gradient penalty to replace the weight clipping method used in WGAN, it enables the network to converge faster. At the same time, the gradient penalty term can also help avoid the problems of vanishing or exploding gradients. The basic principle of WGAN-GP is to add a gradient penalty term to the objective function based on WGAN to limit the gradient of the discriminator. Specifically, the role of the gradient penalty term is to make the gradient norm of the discriminator at the interpolation points between real samples and generated samples as close to 1 as possible. Based on WGAN-GP, a more stable training process and better quality of generated samples can be achieved.

[0047] In a possible implementation, the pre-trained generative adversarial networks of each microgrid satisfy the following generation constraints: the transition probabilities of the uncertainty environmental quantities of each microgrid generated by the pre-trained generative adversarial networks of each microgrid respectively fall within the probability distribution fuzzy sets constructed with the empirical probability distribution of the uncertainty environmental quantity as the center and a preset Wasserstein distance as the radius for each microgrid.

[0048] Specifically, this generation constraint can limit the uncertainty of the environmental quantity within a specific range to achieve a balance between economy and robustness, and avoid the final optimal scheduling decision being too conservative.

[0049] Among them, the preset Wasserstein distance can be given according to historical experience. In practical applications, a specific value can be assumed for training first, and the uncertainty under this preset value can be evaluated through the training results and then adjusted.

[0050] In a possible implementation, during the iterative training step, after each training step is completed, the pre-trained generative adversarial networks of each microgrid are updated with the goal of minimizing the reward value function of the optimal scheduling reinforcement learning model of each microgrid.

[0051] Specifically, the training process of the pre-trained generative adversarial network is increased. The training objective of the pre-trained generative adversarial network is to find the worst prediction environment of the microgrid under certain constraints for the optimized scheduling reinforcement learning model to learn, so as to improve the robustness of the optimized scheduling reinforcement learning model as much as possible.

[0052] Therefore, in this embodiment, the pre-trained generative adversarial network of each microgrid is updated with the goal of minimizing the reward value function of the optimized scheduling reinforcement learning model of each microgrid. That is, the generative adversarial network can adopt a reward function setting and iteration direction that are completely opposite to those of the optimized scheduling reinforcement learning model.

[0053] In addition, combined with the settings of the previous embodiment, the pre-trained generative adversarial network of each updated microgrid also satisfies the corresponding generation constraints. Therefore, W represented by the Wasserstein distance, φ represented by the updated parameters of the pre-trained generative adversarial network, φ 0 represented by the initial parameters of the pre-trained generative adversarial network, the update of the pre-trained adversarial network satisfies:

[0054]

[0055] Among them, represents the expectation of, represents the probability distribution output by the updated pre-trained generative adversarial network, represents the probability distribution output by the initial pre-trained generative adversarial network, represents and the Wasserstein distance between, represents the upper limit of the Wasserstein distance.

[0056] In a possible implementation manner, the training of the optimized scheduling reinforcement learning model of each microgrid by using the coalition game according to the prediction environment of each microgrid includes:

[0057] The optimization scheduling reinforcement learning models of each microgrid interact with the prediction environment of each microgrid respectively to obtain the historical experience of each microgrid; based on the net power demand data in the historical experience of each microgrid, the alliance allocation benefits and unit electricity prices of each microgrid are obtained based on the market intermediate price mechanism and the Shapley value method; according to the alliance allocation benefits and unit electricity prices of each microgrid, the reward function value of the optimization scheduling reinforcement learning model of each microgrid is obtained; through the judgment network of the optimization scheduling reinforcement learning model of each microgrid, according to the reward function value of the optimization scheduling reinforcement learning model of each microgrid and the historical data stored in the experience replay pool in the historical experience, the state value function of the optimization scheduling reinforcement learning model of each microgrid is obtained, and the generalized advantage estimation is carried out according to the state value function of the optimization scheduling reinforcement learning model of each microgrid to obtain the advantage function of the optimization scheduling reinforcement learning model of each microgrid; according to the state value function of the optimization scheduling reinforcement learning model of each microgrid, the judgment network of the optimization scheduling reinforcement learning model of each microgrid is updated; according to the advantage function of the optimization scheduling reinforcement learning model of each microgrid, the policy network of the optimization scheduling reinforcement learning model of each microgrid is updated by using the constrained policy optimization method.

[0058] In this embodiment, the optimization scheduling reinforcement learning model for obtaining each microgrid can directly adopt an existing optimization scheduling reinforcement learning model or directly construct the optimization scheduling reinforcement learning model for each microgrid.

[0059] Specifically, in this embodiment, the following method is adopted to construct the optimization scheduling reinforcement learning model for each microgrid:

[0060] 1. Design of the action space. The action of the optimization scheduling reinforcement learning model is defined as the internal scheduling decision variable of the microgrid, including the active power output of the gas turbine, the reactive power output of the gas turbine, the active power output of the energy storage system, and the reactive power output of the energy storage system:

[0061]

[0062] Among them, is n the action of the optimization scheduling reinforcement learning model of the microgrid at t time, is n the active power output of the gas turbine of the microgrid at t time, is n the reactive power output of the gas turbine of the microgrid at t time, is n the active power output of the energy storage system of the microgrid at t time, is n the reactive power output of the energy storage system of the microgrid at tReactive power output of the energy storage system at a certain moment.

[0063] 2. Design of the state space. The optimal scheduling reinforcement learning model needs to issue scheduling instructions in advance. Therefore, the historical information that the optimal scheduling reinforcement learning model can perceive from the environment is used as the state quantity to avoid the result deviation caused by inaccurate source-load prediction data. The active power output of the gas turbine and the energy of the energy storage system at the previous moment, as well as the new energy output and load in the previous several moments, are used as the state , in this embodiment, 15 minutes is set as one moment, and the new energy output and load in the previous 4 moments are used:

[0064]

[0065] Among them, is n the state of the microgrid's optimal scheduling reinforcement learning model at t a certain moment, is n the active power output of the gas turbine of the microgrid at t -1 moment, is n the energy of the energy storage system of the microgrid at t -1 moment, is n the new energy output of the microgrid at t -4 moment, is n the new energy output of the microgrid at t -1 moment, is n the load of the microgrid at t -4 moment, is n the load of the microgrid at t -1 moment.

[0066] 3. Design of the reward function. The reward values of the optimal scheduling reinforcement learning models of each microgrid include: 1. Reward for the operation cost of the microgrid, which is oriented by the actual optimization goal to ensure the training trend of each agent. 2. Reward for the interest of the coalition contribution, which distributes interests according to the marginal contribution of participating in the coalition to ensure the stability of the microgrid group coalition.

[0067] First, consider the determination of the coalition distribution interest and the unit electricity price. Specifically, according to the net power demand data in the historical experience of each microgrid, the net load power t , net power generation and net power demand of the energy Internet at a certain moment are obtained through the following formula:

[0068] , ,

[0069] wherein, is n the net power demand of the microgrid at t moment, , is the microgrid alliance, is the set of microgrids with positive net power demand, is the set of microgrids with negative net power demand.

[0070] When , , ; when , , ; when , , ; wherein, is t the intermediate value of the purchase electricity price of the superior power grid unit at and t the selling electricity price of the superior power grid unit at moment; is the unit purchase electricity price between microgrids, is the unit selling electricity price between microgrids.

[0071] And when t at the moment, the ij branches connect the n microgrid sells electricity to the m microgrid:

[0072] ,

[0073] wherein, is n the tie-line power of the microgrid node i at t moment, is m the tie-line power of the microgrid node j at t moment, is n the unit selling electricity price of the microgrid at t moment, is m the unit purchase electricity price of the microgrid at t moment.

[0074] Taking the total cost reduction of collaborative optimization compared with individual optimization as the benefit function, the alliance allocation benefits of each microgrid are obtained through the following formula:

[0075]

[0076]

[0077]

[0078] Among them, represents t the benefit obtained by the microgrid by forming a microgrid alliance at time for the microgrid alliance of the net load power, for the microgrid alliance of the net power generation, is n the marginal contribution of the microgrid to joining the microgrid alliance at time N is the total number of microgrids, S for the microgrid alliance the number of microgrids, is n the benefit allocated by the microgrid in the alliance at t time is excluding n the microgrid, the remaining microgrid alliance.

[0079] Based on the above, t at time n the reward value of the optimal scheduling reinforcement learning model of the microgrid is:

[0080]

[0081] Among them, is t at time n the microgrid operation cost reward of the microgrid, is t at time n the alliance contribution benefit reward of the microgrid, is a preset proportionality coefficient.

[0082]

[0083]

[0084] Among them, is t at time n the power purchase cost of the microgrid, is t at time nThe power generation cost of the microgrid is t at time n the profit distribution of the microgrid alliance

[0085]

[0086]

[0087] wherein is t at time n the net power demand of the microgrid is t at time n the i active power output of the gas turbine of the microgrid and are preset coefficients is n the set of gas turbines of the microgrid, [·] + / − indicates taking positive / negative values

[0088] Therefore, the following reward function is set

[0089]

[0090] wherein is n the average value of all reward values generated by the microgrid under the control strategy π n , T is the moment of the optimal scheduling period is the discount coefficient of the preference degree of each microgrid for rewards at different times is n the t reward value of the microgrid at time

[0091] 4. Design of the evaluation network and the policy network. Specifically, for the evaluation network and the policy network, conventional reinforcement learning neural networks can be adopted currently

[0092] See Figure 2 , which shows a multi-agent cooperation framework. The optimal scheduling reinforcement learning models of each microgrid interact with the prediction environment of each microgrid respectively, and then a coordination layer is set to achieve the coordination between each microgrid, that is, based on the net power demand data in the historical experience of each microgrid, the profit distribution and unit electricity price of each microgrid alliance are obtained based on the market intermediate price mechanism and the Shapley value method, and sent to each microgrid

[0093] See Figure 3, which shows the training framework of the optimized scheduling reinforcement learning model. The transfer probability of the uncertainty environmental quantity of each microgrid is generated by the pre-trained generative adversarial network of each microgrid, and the predicted environment of each microgrid is obtained based on the transfer probability. Then, the optimized scheduling reinforcement learning model of each microgrid is respectively interacted with the predicted environment of each microgrid to obtain the historical experience of each microgrid and store it in the experience replay pool. Furthermore, the optimized scheduling reinforcement learning model and the generative adversarial network are updated according to the historical experience.

[0094] The constrained policy optimization method is a typical algorithm in safe reinforcement learning, which quickly solves the non-convex problem of safe reinforcement learning through convex approximation and convex optimization. In this embodiment, a value function for the constraint function is set to support the update of the policy network based on the constrained policy optimization method.

[0095] Specifically, in order to reflect the constraint effect of the system operation state on reinforcement learning, it is set that the constraint function is transformed from the constraints of each device in the microgrid and the power flow constraint. Therefore, the value function of the constraint function:

[0096]

[0097] Among them, is to n take the mean of all constraint return values generated by the microgrid under the control policy π n , is n the constraint function of the microgrid t at time

[0098]

[0099] Among them, is n the node set of the microgrid, is n the branch set of the microgrid.

[0100]

[0101]

[0102]

[0103]

[0104]

[0105] Among them, is t the degree of voltage constraint satisfaction of node i at time ist The current constraint satisfaction degree of the time branch ij , is t the tie-line constraint satisfaction degree at the time is t the gas turbine output and ramp constraint satisfaction degree of node i at the time is t the energy storage constraint satisfaction degree of node i at the time

[0106] is the lower voltage limit of node i , is i the voltage of node t at the time is the upper voltage limit of node i , is the current of branch ij at the time t , is the upper current limit of branch ij , is n the tie-line power of microgrid node i at the time t , is n the upper tie-line power limit of microgrid node i .

[0107] and are respectively n the upper and lower active power limits of the gas turbine of microgrid node i ; is n the upper reactive power limit of the energy storage system of microgrid node i ; is n the maximum ramp power of the gas turbine of microgrid node i ; is n the energy value of the energy storage system of microgrid node i at the time t , and are respectively n the upper and lower energy value limits of the energy storage system of microgrid node i ; is n the active power of the energy storage system of microgrid node i at the time t , is nMicrogrid Node i Upper limit of the discharge power of the energy storage system; is n Microgrid Node i Upper limit of the charging power of the energy storage system; is n Microgrid Node i Reactive power of the energy storage system at t moment; is n Microgrid Node i Upper limit of the reactive power of the energy storage system.

[0108] In a possible implementation manner, the simulation verification of the energy Internet optimal operation method of the present invention is carried out with a typical park of a certain integrated energy service company as an example. The typical park of this integrated energy service company includes three parks. Each of the three parks includes photovoltaic, combined heat and power units and energy storage, and can directly trade electric energy with the superior power grid. In order to more intuitively verify that the dispatching result meets the optimality law, the lowest unit power generation cost is set for Park 1, and the highest unit power generation cost is set for Park 3.

[0109] Assume that the power generation costs of photovoltaic and energy storage are ignored, and the unit power generation cost of Park 1 is lower than that of other parks. Therefore, when the load is greater than the photovoltaic power generation, the park group preferentially meets the load demand through photovoltaic and energy storage, then uses the gas turbine of Park 1 to make up for it, and finally purchases electricity from the superior power grid by using the gas turbine of Park 2 to make up the difference. On the contrary, if the load is less than the photovoltaic power generation, the photovoltaic is preferentially used to meet the demand and charge the energy storage. At this time, any excess electric energy will be sold to the superior power grid to generate profits, and the output of the gas turbine usually remains at the lowest level.

[0110] By analyzing the convergence curves of the overall reward values in three scenarios: alternately training the optimal dispatching reinforcement learning model and the generative adversarial network, only training the optimal dispatching reinforcement learning model, and only training the generative adversarial network. It can be found that the optimal dispatching reinforcement learning model promotes the increase of the reward value, while for the generative adversarial network, the most adverse dispatching environment is found by minimizing the reward value to improve the decision-making robustness. And, by analyzing the voltage amplitude curves of all nodes during the training process, it can be found that the voltage amplitudes of all nodes are maintained within the constraint range, effectively ensuring the safety of real-time decision-making.

[0111] To further illustrate that this method can balance economy and robustness, see Table 1 for a comparison of the algorithm performance under different degrees of prediction errors.

[0112] Table 1

[0113]

[0114] It can be seen that when the prediction error is greater than 6%, the economy of the optimal operation method of the energy Internet of the present invention is superior to the current deterministic mechanism optimization method. When the actual scheduling scenario has a low requirement for robustness, the optimal operation method of the energy Internet of the present invention can further reduce costs and be superior to the current deterministic mechanism optimization method under a smaller prediction error.

[0115] The following is an apparatus embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For details not disclosed in the apparatus embodiment, please refer to the method embodiment of the present invention.

[0116] See Figure 4 , in another embodiment of the present invention, an energy Internet optimal operation system is provided, which can be used to implement the above-mentioned energy Internet optimal operation method. Specifically, the energy Internet optimal operation system includes a data acquisition module and a decision-making module.

[0117] Among them, the data acquisition module is used to acquire the state data of each microgrid in the energy Internet; the decision-making module is used to obtain the optimal scheduling decision of each microgrid through the preset policy network model of each microgrid according to the state data of each microgrid. Among them, the preset policy network model of each microgrid is obtained through the following method: obtaining the optimal scheduling reinforcement learning model of each microgrid; training step: generating the transition probability of the uncertainty environment quantity of each microgrid through the pre-trained generative adversarial network of each microgrid, and obtaining the prediction environment of each microgrid based on the transition probability, and training the optimal scheduling reinforcement learning model of each microgrid by using the method of coalition game according to the prediction environment of each microgrid; iteratively training the steps until the optimal scheduling reinforcement learning model of each microgrid is trained, and obtaining the policy network in the optimal scheduling reinforcement learning model of each microgrid to obtain the preset policy network model of each microgrid.

[0118] All relevant contents of each step involved in the embodiment of the above-mentioned energy Internet optimal operation method can be cited in the function description of the function module corresponding to the energy Internet optimal operation system in the embodiment of the present invention, and will not be repeated here.

[0119] The division of modules in the embodiment of the present invention is illustrative, only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present invention, each function module can be integrated in one processor, or can exist separately physically, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software function modules.

[0120] In another embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of the energy Internet optimization operation method.

[0121] In another embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the energy Internet optimization operation method in the above embodiments.

[0122] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0123] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0124] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. An energy Internet optimal operation method, characterized in that, Including: Obtaining the status data of each microgrid in the energy Internet; According to the status data of each microgrid, respectively through the preset policy network model of each microgrid, obtaining the optimal scheduling decision of each microgrid; Among them, the preset policy network model of each microgrid is obtained through the following method: Obtaining the optimal scheduling reinforcement learning model of each microgrid; Training steps: Generating the transition probability of the uncertain environmental quantity of each microgrid through the pre-trained generative adversarial network of each microgrid, and obtaining the predicted environment of each microgrid based on the transition probability, and training the optimal scheduling reinforcement learning model of each microgrid by means of coalition game according to the predicted environment of each microgrid; Iteratively training the steps until the optimal scheduling reinforcement learning model of each microgrid is trained, and obtaining the policy network in the optimal scheduling reinforcement learning model of each microgrid, and obtaining the preset policy network model of each microgrid; During the iterative training step, after each training step is completed, respectively aiming at minimizing the reward value function of the optimal scheduling reinforcement learning model of each microgrid, updating the pre-trained generative adversarial network of each microgrid; It also includes: Taking the state of the optimal scheduling reinforcement learning model of each microgrid as the input and the transition probability of the uncertain environmental quantity of each microgrid as the output, constructing the initial generative adversarial network of each microgrid; Respectively training the initial generative adversarial network of each microgrid according to the historical data of each microgrid, and obtaining the pre-trained generative adversarial network of each microgrid; Among them, the generative adversarial network is used to predict the probability distribution of the uncertain environmental quantity at the next moment by fixing the state of the uncertain environmental quantity at the previous moment, and then realizing the state prediction of the uncertain environmental quantity at the next moment by sampling the probability distribution; The status data includes the active power output of the gas turbine and the energy of the energy storage system at the previous moment, as well as the new energy output and load at several previous moments; The uncertain environmental quantity includes new energy output and load; The optimal scheduling decision of each microgrid includes the active power output of the gas turbine of each microgrid, the reactive power output of the gas turbine, the active power output of the energy storage system, and the reactive power output of the energy storage system.

2. The energy Internet optimal operation method according to claim 1, wherein The generative adversarial network is constructed based on the Wasserstein generative adversarial network with gradient penalty.

3. The method for optimizing the operation of the energy Internet according to claim 1, wherein, The pre-trained generative adversarial network of each microgrid satisfies the following generation constraints: The transition probability of the uncertain environmental quantity of each microgrid generated by the pre-trained generative adversarial network of each microgrid respectively falls within the probability distribution fuzzy set constructed with the empirical probability distribution of the uncertain environmental quantity as the center and the preset Wasserstein distance as the radius of each microgrid.

4. The method for optimizing the operation of the energy Internet according to claim 1, wherein The training of the optimal scheduling reinforcement learning model of each microgrid by means of coalition game according to the predicted environment of each microgrid includes: The optimal scheduling reinforcement learning model of each microgrid respectively interacts with the predicted environment of each microgrid to obtain the historical experience of each microgrid; According to the net power demand data in the historical experience of each microgrid, obtaining the coalition allocation benefit and unit electricity price of each microgrid based on the market intermediate price mechanism and the Shapley value method; According to the coalition allocation benefit and unit electricity price of each microgrid, obtaining the reward function value of the optimal scheduling reinforcement learning model of each microgrid; The evaluation network of the reinforcement learning model for the optimal dispatch of each microgrid is strengthened. Based on the reward function value of the reinforcement learning model for the optimal dispatch of each microgrid and the historical data stored in the experience replay pool in the historical experience, the state value function of the reinforcement learning model for the optimal dispatch of each microgrid is obtained, and the generalized advantage estimation is performed according to the state value function of the reinforcement learning model for the optimal dispatch of each microgrid to obtain the advantage function of the reinforcement learning model for the optimal dispatch of each microgrid; According to the state value function of the reinforcement learning model for the optimal dispatch of each microgrid, the evaluation network of the reinforcement learning model for the optimal dispatch of each microgrid is updated; According to the advantage function of the reinforcement learning model for the optimal dispatch of each microgrid, the policy network of the reinforcement learning model for the optimal dispatch of each microgrid is updated by using the constrained policy optimization method.

5. An optimized operation system for an energy Internet, characterized in that, It includes: A data acquisition module for acquiring the state data of each microgrid in the energy Internet; A decision-making module for obtaining the optimal dispatch decision of each microgrid respectively through the preset policy network model of each microgrid according to the state data of each microgrid; Among them, the preset policy network model of each microgrid is obtained through the following method: Obtain the reinforcement learning model for the optimal dispatch of each microgrid; Training steps: Generate the transition probability of the uncertain environmental quantity of each microgrid through the pre-trained generative adversarial network of each microgrid, obtain the predicted environment of each microgrid based on the transition probability, and train the reinforcement learning model for the optimal dispatch of each microgrid by using the coalition game method according to the predicted environment of each microgrid; Iteratively train the steps until the reinforcement learning model for the optimal dispatch of each microgrid is trained, and obtain the policy network in the reinforcement learning model for the optimal dispatch of each microgrid to obtain the preset policy network model of each microgrid; During the iterative training step, after each training step is completed, the pre-trained generative adversarial network of each microgrid is updated with the goal of minimizing the reward value function of the reinforcement learning model for the optimal dispatch of each microgrid; It also includes: Construct an initial generative adversarial network for each microgrid with the state of the reinforcement learning model for the optimal dispatch of each microgrid as the input and the transition probability of the uncertain environmental quantity of each microgrid as the output; Train the initial generative adversarial network of each microgrid respectively according to the historical data of each microgrid to obtain the pre-trained generative adversarial network of each microgrid; Among them, the generative adversarial network is used to predict the probability distribution of the uncertain environmental quantity at the next moment by fixing the state of the uncertain environmental quantity at the previous moment, and then realize the state prediction of the uncertain environmental quantity at the next moment by sampling the probability distribution; The state data includes the active power output of the gas turbine and the energy of the energy storage system at the previous moment, and the new energy output and load at several previous moments; The uncertain environmental quantity includes new energy output and load; The optimal dispatch decision of each microgrid includes the active power output of the gas turbine of each microgrid, the reactive power output of the gas turbine, the active power output of the energy storage system, and the reactive power output of the energy storage system.

6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it realizes the steps of the energy Internet optimal operation method according to any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the energy Internet optimal operation method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Distributed energy optimization scheduling method and device based on mixed learning

    CN116451880A

  • Active power distribution network optimization scheduling method, system and device and storage medium

    CN118017492A

  • Micro-grid energy optimization scheduling method

    CN118174355A