Scheduling method and device of integrated energy system, computer equipment and medium
By constructing Markov decision-making process and using neural network environment model, combined with MBPO algorithm, the agent optimizes the scheduling of the comprehensive energy system, solving the scheduling problems in complex high-dimensional environments, improving the efficiency and accuracy of scheduling, and reducing costs.
Patent Information
- Application Number
- CN202510305100.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The prior art is difficult to adapt to the scheduling of integrated energy systems in complex high-dimensional environments, the calculation cost of heuristic algorithms is high and it is difficult to find the global optimal solution, and the mathematical optimization method is difficult to apply to the access of multiple loads and renewable energy.
By constructing Markov decision-making process, using neural networks to build an environmental model, combining the model-based strategy optimization algorithm MBPO, the agent interacts with the environmental model, and by maximizing the reward optimization strategy, the operation of each device in the comprehensive energy system is scheduled.
It improves the multi-energy scheduling capabilities in the integrated energy system, enhances the generalization capabilities of the agent, ensures the accuracy and reliability of scheduling, reduces the number of interactions between the agent and the real world, improves sampling efficiency, and reduces the operating costs and the frequency and cost of manual intervention.
Smart Images

Figure CN120124976A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of energy scheduling, and particularly relates to a scheduling method, device, computer device and medium for an integrated energy system. Background Art
[0002] Under the background of the close coupling between the power grid and the natural gas pipeline network, various types of energy are complexly intertwined. To solve the problem of multi-type energy scheduling, the existing technology has proposed an integrated energy system. The integrated energy system can achieve efficient energy utilization and sustainable development. By coordinating different types of energy to meet diversified load demands, it realizes the integrated operation of energy utilization, energy conversion and energy storage.
[0003] There have been many studies on the operation optimization of integrated energy systems in the existing methods, such as using particle swarm optimization algorithm and linear programming and other methods to reduce costs and carbon emissions. These methods are mainly divided into heuristic algorithms and mathematical optimization methods, but both of these methods have certain deficiencies in practical applications. The heuristic algorithm can solve non-linear, multi-modal and high-dimensional problems, but there are problems such as high computational cost, difficult convergence and inability to guarantee finding the global optimal solution, resulting in poor scheduling optimization effect for the integrated energy system. The mathematical optimization method relies on an accurate mathematical model, but with the access of renewable energy and various loads, accurate modeling becomes increasingly difficult and it is difficult to apply to the integrated energy system. Therefore, the above-mentioned existing technology is limited in optimizing the scheduling method of the integrated energy system, resulting in difficulty in adapting to the integrated energy system in a complex multi-energy high-dimensional environment. Summary of the Invention
[0004] In order to solve the problem that the existing technology is difficult to adapt to the scheduling of the integrated energy system in a complex high-dimensional environment, the present invention provides a scheduling method, device, computer device and medium for an integrated energy system.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] First, a scheduling method for an integrated energy system is provided, including:
[0007] Obtain the state information of the integrated energy system and the operation actions of the equipment operation;
[0008] Construct a Markov decision process according to the integrated energy system; the Markov decision process includes a state space, an action space and a reward function; wherein, the state space is a set of state information constructed according to the state information, the action space is a set of operation actions constructed according to the operation actions, the reward function generates a reward according to the corresponding state space and action space, and the reward is negatively correlated with the operation cost of the integrated energy system;
[0009] Construct an agent for the model-based policy optimization algorithm MBPO based on the Markov decision process, use a neural network to construct an environment model, interact the agent with the environment model, and optimize the agent's policy by maximizing the reward; wherein, the environment model is used to predict the future values of exogenous states and rewards, and the exogenous states include load demand, energy price, and wind speed;
[0010] According to the optimized policy, the agent schedules the operation of each device in the integrated energy system.
[0011] Optionally, the constructing an environment model using a neural network includes:
[0012] Construct an environment model using an ensemble of deep neural networks. Each deep neural network in the ensemble makes independent predictions, and the ensemble of deep neural networks consists of multiple deep neural networks DNNs.
[0013] Optionally, the interacting the agent with the environment model and optimizing the agent's policy of MBPO by maximizing the reward includes:
[0014] Based on the pre-acquired real experience samples, train the environment model through the interaction between the agent and the environment model;
[0015] Use the trained environment model to generate simulated experience samples, and combine with the real experience samples to optimize the agent's policy with the goal of maximizing the reward, obtaining the optimized agent's policy.
[0016] Optionally, the training the environment model through the interaction between the agent and the environment model includes:
[0017] Initialize the policy network Actor and Q-value network Critic of MBPO, as well as two experience replay buffers, which are used to store real experience samples and simulated experience samples respectively;
[0018] The interaction between the agent and the environment model includes: the agent inputs the current state of the integrated energy system into the policy network, based on the environment model to simulate the state change of the integrated energy system, determines the probability distribution of each continuous behavior of the agent, randomly samples an action from the probability distribution, and transfers the integrated energy system simulated by the environment model to the next state and receives a reward;
[0019] Loop through multiple time steps of the interaction process to generate multiple simulated experience samples composed of state, action, reward, and next state, and store the multiple simulated experience samples in the experience replay buffer;
[0020] Train the environment model by minimizing the difference between the next state in the real experience samples and the next state in the simulated experience samples.
[0021] Optionally, the difference between the simulated experience sample and the next state in the real experience sample is characterized by the mean squared error (MSE), and the DNN is trained with the goal of minimizing the difference.
[0022] Optionally, the policy optimization of the agent with the goal of maximizing the reward includes:
[0023] The policy optimization of the agent is achieved by minimizing the real-time operating cost of the integrated energy system in each time slot.
[0024] Secondly, a scheduling device for an integrated energy system is provided, including:
[0025] An acquisition module, configured to acquire the state information of the integrated energy system and the operating actions of the devices.
[0026] A construction module, configured to construct a Markov decision process according to the integrated energy system; the Markov decision process includes a state space, an action space, and a reward function; wherein, the state space is a set of state information constructed according to the state information, the action space is a set of operating actions constructed according to the operating actions, and the reward function generates a reward according to the corresponding state space and action space, and the reward is negatively correlated with the operating cost of the integrated energy system.
[0027] An optimization module, configured to construct an agent based on the model-based policy optimization algorithm (MBPO) according to the Markov decision process, construct an environment model using a neural network, interact the agent with the environment model, and optimize the policy of the agent by maximizing the reward; wherein, the environment model is used to predict the future values of exogenous states and rewards, and the exogenous states include load demand, energy price, and wind speed.
[0028] A scheduling module, configured to schedule the operation of each device in the integrated energy system through the agent according to the optimized policy.
[0029] In addition, a computer-readable storage medium is provided, and the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned scheduling method for an integrated energy system is implemented.
[0030] Finally, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned scheduling method for an integrated energy system is implemented.
[0031] The scheduling method for an integrated energy system provided by the present invention has the following beneficial effects:
[0032] By abstracting the scheduling problem of the integrated energy system into a Markov decision process, it provides a basis for the learning and decision-making of the agent, facilitating optimization using algorithms such as reinforcement learning. On this basis, an agent and an environment model based on a model-based policy optimization algorithm are constructed. The simulated experience quickly generated by the environment model and allowing the agent to explore possible states and actions more widely can adapt to the complex and changeable integrated energy system environment, improve the scheduling ability of multiple energies in the integrated energy system, enhance the generalization ability of the agent, ensure the accuracy and reliability of the agent's scheduling of the integrated energy system, reduce the number of interactions between the agent and the real world, improve the sampling efficiency, thus being more efficient than learning only from real experience, and at the same time can also reduce the operating cost of the integrated energy system and the frequency and cost of manual intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] To more clearly illustrate the embodiments of the present invention and their design solutions, the accompanying drawings required for this embodiment will be briefly introduced below. The drawings in the following description are only partial embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0034] Figure 1 It is a schematic flowchart of a scheduling method for an integrated energy system provided by the present invention according to an exemplary embodiment.
[0035] Figure 2 It is a schematic structural diagram of an integrated energy system provided by the present invention according to an exemplary embodiment.
[0036] Figure 3 It is a schematic diagram of the interaction between an agent and an environment provided by the present invention according to an exemplary embodiment.
[0037] Figure 4 It is a schematic diagram of parameter training of a scheduling algorithm provided by the present invention according to an exemplary embodiment.
[0038] Figure 5 It is a schematic flowchart of a training process provided by the present invention according to an exemplary embodiment.
[0039] Figure 6 It is a block diagram of a scheduling device for an integrated energy system provided by the present invention according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To enable those skilled in the art to better understand the technical solutions of the present invention and implement them, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0041] The present invention is an integrated energy system optimal scheduling method based on model reinforcement learning, and the method includes the following:
[0042] Step 1: Establish an integrated energy system model including electricity, heat, and natural gas, and obtain renewable energy device models, combined heat and power device models, gas turbine device models, electric boiler device models, heat boiler device models, and energy storage device models.
[0043] Step 2: Considering the models of each device in Step 1, construct an operating cost model, clarify system constraints, and construct an operating cost model according to the energy consumption and conversion indexes and parameters of each device.
[0044] Step 3: Based on the optimization objective of the cost in Step 2, map the energy optimal scheduling problem to a Markov decision process model, and design a reinforcement learning state space, action space, and reward function.
[0045] Step 4: Adopt the MBPO algorithm to solve the Markov decision model problem of continuous behavior. Taking the operating cost at each moment as the target, on the basis of expanding and optimizing the network parameter training samples, solve the optimal scheduling scheme. The optimization network includes an Actor network and a Critic network.
[0046] The following will, with reference to the accompanying drawings, detail the technical solutions provided by each embodiment of the present invention.
[0047] First, the present invention provides a scheduling method for an integrated energy system, specifically as Figure 1 shown, including the following steps:
[0048] S101. Obtain the state information of the integrated energy system and the operating actions of the devices.
[0049] The state information includes energy load, consumption, conversion, storage, and price information.
[0050] The structure of the integrated energy system IES in the present invention is as Figure 2As shown, considering the wide application of renewable energy such as wind and solar, introducing power generation equipment such as photovoltaic and wind turbines in the architecture can help the system reduce its dependence on fossil fuels. In addition, the IES architecture also includes gas turbines, combined heat and power units, gas boilers, electric boilers, electricity storage equipment, and heat storage equipment. These can flexibly adjust the output power, quickly respond to demand changes, make up for the gap caused by the fluctuations of renewable energy, and enhance the flexibility of the system. Electric boilers and gas boilers, as heat source equipment, can provide heat energy and work in coordination with the power system, choosing to provide different forms of consumed energy at different times to optimize the comprehensive utilization of energy. Combined heat and power simultaneously produces heat and electricity, improving the energy utilization efficiency. Energy storage equipment can store excess energy and release it during peak demand periods to balance load fluctuations and improve the stability and reliability of the system. In addition, energy can be directly purchased from the main power grid and natural gas suppliers, and the electricity and heat energy produced in the system are respectively delivered to users through the power grid and the district heating network.
[0051] For example, the integrated energy system may include photovoltaic (PV), wind turbine (WT), combined heat and power (CHP), gas boiler (GB), electric boiler (EB), gas turbine (GT), electricity storage (ES), heat storage (HS), electric load (EL), and heat load (HL). Electricity and natural gas can be directly purchased from the power grid and natural gas network and then delivered to users through the regional network.
[0052] Among them, the state information includes but is not limited to various energy demands, storage, and price information, such as electric load and heat load, etc. These loads reflect the energy demands of the system at different time periods; the energy storage information covers the storage status of electricity storage (ES) and heat storage (HS) equipment; the price information includes real-time electricity prices and natural gas prices, and these price information are crucial for optimizing the energy purchase strategy.
[0053] and respectively represent the electricity generated by PV and WT, which is calculated from the wind speed. CHP generates heat and electricity by consuming natural gas. GB and GT respectively output heat and electricity by burning natural gas. EB generates heat by consuming electrical energy through resistance and is the energy change of the storage device within the time period t. If it is positive, it means charging; otherwise, it means discharging.
[0054] S102. Construct a Markov decision process according to the integrated energy system.
[0055] The Markov decision process MDP includes a state space, a behavior space and a reward function; wherein the state space is a set of state information constructed according to the state information, the behavior space is a set of operating actions constructed according to the operating actions, and the reward function generates a reward according to the corresponding state space and behavior space, and the reward is negatively correlated with the operating cost of the integrated energy system.
[0056] Specifically, the state space is composed of state information, and the behavior space includes the relevant parameters of the operating actions of each device in the integrated energy system, such as the power generation of PV and WT, the natural gas consumption of CHP, the heat production and power output of GB and GT, the power consumption of EB, the charging and discharging / thermal power of ES and HS, etc.; the reward function is calculated according to the total cost of the integrated energy system. The total cost includes the cost of purchasing energy and the cost of equipment maintenance. The reward function is designed to be the negative value of the total cost, so that maximizing the reward is equivalent to minimizing the total cost.
[0057] For example, in MDP, the state space may include power load, heat load, electricity price, natural gas price, wind speed, PV power generation, ES power and HS heat, etc.; the behavior space may include PV power generation, CHP natural gas consumption, ES charging and discharging power, etc.; the reward function is designed as: -(energy purchase cost + equipment maintenance cost).
[0058] In one embodiment, the Markov decision process generally includes a tuple (S, A, P, r, γ), where S, A, and γ represent the state space, action space, and discount factor of future rewards, respectively. IES is the environment of the present invention, which provides the agent with its own state s at time slot t t ∈S. The agent is based on the state s t Returns the scheduled action a t ∈A. In addition, P(s t+1 |s t ,a t ) is the state transfer function, r t (s t ,a t ) is the reward function. The Markov decision process is as follows:
[0059] Definition 1 (State Space). At time t, the state space of the integrated energy system is defined as follows:
[0060]
[0061] in, and The demand for electricity and heat; and are the electricity price and natural gas price; v t is the wind speed, through which the corresponding amount of electricity produced can be calculated ; The power generation of photovoltaic; and is the energy storage state of the electricity storage device and the heat storage device, expressed as a percentage.
[0062] Definition 2 (behavior space). At time t, the behavior space of the integrated energy system is defined as follows:
[0063]
[0064] where The electrical output of the combined heat and power equipment; is the thermal output of the gas boiler; is the electrical output of the gas turbine; is the thermal output of the electric boiler; P t ES and P t HS respectively represent the charging and discharging energy magnitudes of the electricity storage device and the heat storage device.
[0065] Definition 3 (reward function). The optimization objective of the present invention is to minimize the energy purchase cost and equipment maintenance cost of the integrated energy system. Therefore, the reward function at time step t can be calculated as:
[0066]
[0067]
[0068] where represents the total cost of the integrated energy system; and respectively represent the cost of purchasing energy and the cost of equipment maintenance; represents the electricity purchase quantity from the main power grid; represents the natural gas quantity purchased from the natural gas official website, represents the maintenance price coefficient of equipment j, is the output of the equipment at time t, where j = {WT, PV, CHP, EB, GB, GT, ES, HS}.
[0069] Definition 4 (optimization objective). The optimization objective of the present invention is as follows:
[0070]
[0071] subject to
[0072]
[0073]
[0074]
[0075] Among them, the discount factor is represented by the symbol γ. Formulas (1.6), (1.7), and (1.8) are the power constraint, heat constraint, and natural gas constraint of the integrated energy system, aiming to ensure the safe and efficient operation of the integrated energy system while meeting the demand.
[0076] S103. Construct an agent of the model-based policy optimization algorithm MBPO according to the Markov decision process, construct an environment model using a neural network, interact the agent with the environment model, and optimize the policy of the agent by maximizing the reward.
[0077] Among them, the environment model is used to predict the future values of exogenous states and rewards. The exogenous states include uncertain variables such as load demand, energy price, and wind speed. The environment model is constructed through a neural network and can learn and simulate the dynamic behavior of the integrated energy system. The MBPO agent then uses these simulated experience samples for policy optimization, reducing the need for interaction with the real environment and improving the learning efficiency.
[0078] Specifically, based on the pre-acquired real experience samples, the environment model is trained through the interaction between the agent and the environment model; the trained environment model is used to generate simulated experience samples, combined with the real experience samples, and the policy of the agent is optimized with the goal of maximizing the reward to obtain the optimized agent policy.
[0079] In the present invention, MBPO uses a neural network to construct a model approximating the environment dynamics and uses the model to generate simulated experience samples to optimize the policy. Generally speaking, the state transition function of the Markov decision process is unknown. In the present invention, according to whether a neural network is needed for prediction, the states are divided into two types:
[0080] The endogenous state is the internal state of the IES, such as and which can be directly calculated in the environment based on the previous state and action without neural network prediction.
[0081] The exogenous states include random variables such as load demand, energy price, and wind speed, such as v t and Since these exogenous random variables have strong uncertainty, an environment model needs to be used for prediction to generate experience samples for policy optimization.
[0082] The present invention uses a model-based strategy to optimize the MBPO algorithm to solve the IES energy scheduling problem. The agent interacts with the environment to obtain simulated experience samples for training the environment model, such as Figure 3 as shown. The simulated experience samples generated by using the trained environment model can be used to more effectively explore states and actions, rather than relying solely on real experience samples for policy optimization.
[0083] Among them, the training of the environment model can be as follows: Initialize the policy network Actor and Q-value network Critic of MBPO, as well as two experience replay buffers for storing real experience samples and simulated experience samples respectively; Interact the agent with the environment model, including: The agent inputs the current state of the integrated energy system into the policy network, simulates the state change of the integrated energy system based on the environment model, determines the probability distribution of each continuous action of the agent, randomly samples actions from the probability distribution, and transfers the integrated energy system simulated by the environment model to the next state and receives a reward; Loop through multiple time steps of the interaction process to generate multiple simulated experience samples composed of states, actions, rewards, and the next state, and store the multiple simulated experience samples in the experience replay buffer; Train the environment model by minimizing the difference between the next state in the real experience samples and the next state in the simulated experience samples, or train the environment model by minimizing the difference between the rewards in the real experience samples and the rewards in the simulated experience samples.
[0084] To solve the problem of performance degradation caused by model bias, MBPO uses an ensemble of deep neural networks (with an ensemble size of N) to construct the environment model Each deep neural network in the ensemble makes independent predictions. The ensemble of deep neural networks consists of multiple deep neural networks DNNs. The ensemble model in the present invention consists of seven deep neural networks (DNNs), which increases the complexity and computational cost of the model, but it can effectively cope with the uncertainty brought by environmental noise, as well as the uncertainty brought by insufficient model parameters and data. The ensemble model can maintain a high prediction accuracy in the case of scarce data or drastic environmental changes, and reduce the negative impact on scheduling decisions by suppressing the accumulation of prediction biases.
[0085] As shown in the following formula, the ensemble model is trained by the mean squared error (MSE) loss function. The difference between the simulated experience samples and the next state in the real experience samples is characterized by the mean squared error MSE, and the DNNs are trained with the goal of minimizing this difference:
[0086]
[0087] Among them, ω is the neural network parameter of the DNNs; N is the sample size; is the predicted value of DNNs; e n is the true state; σ n is the predicted standard deviation.
[0088] The SAC (Soft Actor-Critic, a deep learning-based reinforcement learning algorithm) algorithm uses a mixture of simulated samples and real experience samples to optimize the policy and adds an entropy term to the reward to ensure a high level of exploration and robustness in the face of random factors. SAC consists of five networks: a policy network (Actor), two Q-value networks (current Critic), and two target Q-value networks (target Critic). The parameters of these networks are φ, θ m=1,2 and The Actor is updated by minimizing the soft Bellman residual:
[0089]
[0090] where, is the experience replay buffer, which is used to store experience samples, is the state s at time t t and action a t corresponding to the Q-value of the current Critic algorithm, is the state s at time t t and action a t corresponding to the Q-value of the target Critic algorithm. The parameters of the target Critic are soft-updated by mixing a small fraction of the current Critic network parameters into the old target Critic network parameters. This method makes the changes in the target Critic network smoother, thus helping to stabilize the training process. During the algorithm policy improvement process, the update of the agent network is achieved based on maximizing the expected value of the weighted Q-value and entropy, where π is the current policy:
[0091]
[0092] where, logπ φ (a t |s t ) is the entropy. The higher the entropy, the more the agent can find more effective strategies through continuous exploration; α is a hyperparameter used to control the weight of the entropy term. Using the minimum of the two Q-values can stabilize the training process and reduce the estimation bias.
[0093] The network parameters φ and θ are updated by the gradient descent method:
[0094]
[0095]
[0096] where λ π and λ Q are the update steps of the parameters φ and θ
[0097] To further improve the sample efficiency, the baseline algorithm SAC in the present invention adopts a prioritized experience replay (PER) mechanism. By assigning higher priorities to the experience samples that have a greater impact on policy improvement, usually determined by the temporal error, it ensures that the more influential samples are used more frequently, thereby accelerating the learning process.
[0098] In one embodiment, the scheduling algorithm of the present invention improves the sample efficiency and the policy learning rate through two types of experience samples, thereby minimizing the cost of realizing the device actions of the control system, as Figure 4 shown. (1) Real experience samples. Real experience refers to the experience of the current state, current action, reward, and next state obtained from the interaction with the integrated energy system, and these experiences are directly used for training the prediction model and subsequent policy optimization. (2) Simulated experience samples. Simulated experience is generated by the model and is used to learn the optimal policy. It should be noted that this scheme can reduce the number of real interactions and improve the sampling efficiency. The training process of MBPO in the entire workflow is as Figure 5 and the algorithm shown in Table 1 below.
[0099] Table 1 Algorithm for the training process of MBPO
[0100]
[0101]
[0102] At each time step t, the real experience samples stored in the environmental experience replay pool that interact with the environment are regularly used to train the environmental model After the training is completed, the model selects N historical states as the initial states for prediction. After determining the states and actions, each DNN predicts the Gaussian distribution of the next state and reward. Then, the agent generates the next action based on the next state and repeats the above process k times. Finally, the ensemble model generates N×k simulated experience samples and stores them in the model replay experience pool . The simulated samples reduce the need for real environment interactions in the policy learning process.
[0103] Updating the model usually involves procedures such as calculating complex gradients and refitting data. If the model is updated too frequently, the computational overhead may be very high, resulting in a slowdown in the training speed. To balance the model accuracy and learning efficiency, the model in the present invention is retrained every 250 steps (model update frequency) to ensure the stability of the learning process. Finally, the simulated experience samples from and the simulated experience samples from Combine with the real experience samples and optimize the strategy by combining with the SAC algorithm.
[0104] Specifically, first, to store real and simulated experiences, initialize two experience replay buffers. To implement the selection, evaluation, and optimization of actions, initialize the policy network Actor and the Q-value network Critic. (1) The policy network π φ , which is used to make decisions in the system; (2) The Q-value network Q θ , which is used to evaluate the decisions made by the policy network π φ in the current state.
[0105] Subsequently, the agent observes the current state s t of the integrated energy system and inputs it into the policy network π φ . This network outputs the probability distribution a t ~ π(·|s t ) of each continuous action. The agent randomly samples a t from the probability distribution, returns to the integrated energy system environment, and the environment transfers to the next state and receives a reward. Repeat the above interaction process for several time steps to generate a series of data quadruples (s t , a t , r t , s t+1 ) composed of state, action, reward, and next state. These data are then stored in the environmental experience replay buffer D env . D env is used to train the environmental prediction model and collect experiences for training the model in the initial stage (lines 2 - 6 of Algorithm 1).
[0106] Secondly, regularly use the experience samples in D env to train the environmental model (lines 7 - 9 of Algorithm 1). During the training process of the prediction model, the following defined loss function is used for optimization:
[0107]
[0108] where ω are the parameters of the DNN; N is the size of the samples; is the predicted average value of the DNN; e i are the real state and reward values; σ i is the predicted standard deviation.
[0109] Meanwhile, randomly select experience samples (s env , a t , r t , s t ) from D t+1 and take st As the initial state of the model's rolling prediction. Starting from s t on, let the environment integrate the model and perform k-step rolling prediction of the model by N deep neural networks in it. Store the predicted results in the model experience replay buffer (lines 10 - 14 of Algorithm 1).
[0110] Subsequently, mix the experience samples in and D env , and then use these samples to update the parameters φ of the policy network π φ and the parameters θ of the Q-value network Q θ in a gradient update manner (lines 15 - 18 of Algorithm 1), as defined below:
[0111]
[0112]
[0113] where λ π and λ Q are the update step sizes of the parameters φ and θ respectively. J π (φ t ) and J Q (θ t ) are the objective functions of the policy network and the Q-value network respectively, and their definitions are as follows:
[0114]
[0115]
[0116] The policy learning and optimization part of the present invention is implemented using the model-free reinforcement learning Soft Actor-Critic (SAC) algorithm. SAC is an algorithm based on entropy maximization. Its objective function adds an information entropy term αlogπ(a t |s t ) to the traditional cumulative reward, which consists of the adaptive entropy coefficient α and the logarithmic probability logπ(a t |s t |s t ) of the selected action a. This ensures that it can still maintain high exploration and robustness in the face of random factors (such as electricity price fluctuations or load changes). To improve the model's adaptive learning and generalization ability, SAC contains five deep neural networks: a policy network, two Q-value networks, and two target Q-value networks, as shown Figure 4 on the left side, and their parameters are φ, θ m=1,2 and
[0117] Equation (1.12) is the improvement process of the Actor network in the algorithm, that is, the Actor network update is realized by maximizing the expected value of the weighted Q value and entropy. In the equation, is the experience replay buffer, which stores a mixed sample of real experience and simulated experience (s t , a t , r t , s t+1 ). That is the Critic network, which will concatenate the state and action as the input, and output a Q value estimate after a series of linear transformations and activation functions in the neural network. The stability of training is improved by taking the minimum value operation of the two Q values.
[0118] Equation (1.13) is that the Critic network in the algorithm updates the parameters by minimizing the mean square error (MSE) loss function of the target Q value y t and the real Q value . The significance of the target Critic network lies in that the parameters of the Q value network are continuously updated. If the Q value output by the current Q value network is directly used to update itself, it will lead to instability in the training process. The calculation method of the target Q value is defined as follows:
[0119]
[0120] where r t is the current reward, a t+1 ~π(·|s t+1 ) is the probability distribution of selecting the action a t+1 according to the next state. By selecting the minimum value from the Q value estimates of the next state output by the two target Q value networks, the training stability is improved. And the influence of the policy entropy is comprehensively considered, that is, the logarithmic probability of the policy network π φ taking the action a t+1 .
[0121] The parameters of the target Critic network are updated as follows:
[0122]
[0123] where τ is a coefficient less than 1, which is used to control the update speed. The parameters of the target Critic network are softly updated by mixing a small part of the parameters of the current Critic network with the old target Critic network parameters. This method makes the update of the target Critic network more smooth, thus helping to stabilize the training process.
[0124] S104. According to the optimized policy, the operation of each device in the integrated energy system is scheduled through the agent.
[0125] Based on the optimized policy, the agent adjusts the operating parameters of each device in the integrated energy system in real time to meet the load demand while minimizing the total cost. These adjustments may include changing the power generation of PV and WT, adjusting the natural gas consumption of CHP, controlling the charging / discharging / thermal power of ES and HS, etc.
[0126] For example, under the optimized policy, the agent may decide to increase the electricity purchase volume when the electricity price is low, while reducing the power generation of PV and WT to save costs; when the heat load is high, it may increase the natural gas consumption of CHP to provide more heat.
[0127] By adopting the above method, abstracting the scheduling problem of the integrated energy system into a Markov decision process provides a basis for the learning and decision-making of the agent, facilitating optimization using algorithms such as reinforcement learning. Based on this, an agent and an environment model of a model-based policy optimization algorithm are constructed. Utilizing the simulated experience quickly generated by the environment model and allowing the agent to explore possible states and actions more extensively, it can adapt to the complex and changeable integrated energy system environment, improve the multi-energy scheduling ability in the integrated energy system, enhance the generalization ability of the agent, ensure the accuracy and reliability of the agent's scheduling of the integrated energy system, reduce the number of interactions between the agent and the real world, improve the sampling efficiency, thus being more efficient than only learning from real experience. At the same time, it can also reduce the operating cost of the integrated energy system and the frequency and cost of manual intervention.
[0128] Secondly, the present invention also provides a scheduling device for an integrated energy system, as Figure 6 shown, including:
[0129] An acquisition module 601 for acquiring the state information of the integrated energy system and the operating actions of the devices.
[0130] A construction module 602 for constructing a Markov decision process according to the integrated energy system; the Markov decision process includes a state space, an action space, and a reward function; wherein, the state space is a set of state information constructed according to the state information, the action space is a set of operating actions constructed according to the operating actions, and the reward function generates a reward according to the corresponding state space and action space, and the reward is negatively correlated with the operating cost of the integrated energy system.
[0131] An optimization module 603 for constructing an agent of a model-based policy optimization algorithm MBPO according to the Markov decision process, constructing an environment model using a neural network, interacting the agent with the environment model, and optimizing the policy of the agent by maximizing the reward; wherein, the environment model is used to predict the future values of exogenous states and rewards, and the exogenous states include load demand, energy price, and wind speed.
[0132] A scheduling module 604, configured to operate each device in the integrated energy system through an agent according to the optimized policy.
[0133] The scheduling process implemented by each module of the integrated energy system is as follows:
[0134] 1. Take the current environmental state of the integrated energy system as the environment for interacting with the reinforcement learning algorithm. The acquisition module 601 collects various state information from the integrated energy system, including energy load information (such as the load magnitudes of different energies at different times), consumption information (the consumption amounts of various energies), conversion information (such as the conversion situations of energy conversion devices like power-to-gas, combined heat and power, etc.), storage information (the energy storage levels of energy storage devices), and price information (the market prices of various energies).
[0135] 2. After integrating the key states, the construction module 602 constructs a Markov decision process model based on the obtained information of the integrated energy system. The state space contains the state information obtained from the acquisition module, comprehensively reflecting the current state of the integrated energy system. The action space incorporates the operation information of the devices, representing the operations that can be taken in this system. The reward function is calculated based on the total cost of the integrated energy system. Here, the total cost takes into account the cost of purchasing energy (such as the cost of purchasing electricity, natural gas, etc.) and the cost of equipment maintenance (the expenses required for maintaining various devices). The design of the reward function aims to guide the system to schedule in the direction of cost optimization. When the total cost decreases, a higher reward is given; otherwise, a lower reward is given, providing a quantitative evaluation criterion for subsequent optimization.
[0136] 3. Based on the constructed Markov decision process model, the optimization module 603 starts to construct an agent for the model-based policy optimization algorithm MBPO. A neural network is used to construct an environmental model that can predict the future values of exogenous states and rewards. The exogenous states include load demands (such as possible future changes in electrical load, thermal load, etc.), energy prices (future price fluctuations of different energies), and wind speeds (especially important for systems relying on wind energy). Using the empirical samples simulated by the environmental model, the policy of the MBPO agent is optimized. Through continuous simulation and optimization, the agent can learn which actions (device operation operations) to take in different states to obtain better rewards, and thus gradually adjust and optimize its policy to achieve better scheduling performance.
[0137] 4. After the optimization module 603 completes the policy optimization of the MBPO agent, the scheduling module 604 will use the optimized policy. The agent will schedule the operation of each device in the integrated energy system according to the current state of the integrated energy system and its optimized policy, and decide the output operation of the device to achieve the overall optimized operation of the system.
[0138] During the actual operation process, according to the feedback of the system and new state information, the above modules may continuously work in a loop to achieve dynamic and continuous optimized scheduling of the integrated energy system. The acquisition module will continuously update the state and operation information, the optimization module will continuously optimize the agent policy, and the scheduling module will adjust the device operation according to the new optimized policy, forming a closed-loop optimized scheduling process.
[0139] By adopting the above device, abstracting the scheduling problem of the integrated energy system into a Markov decision process provides a basis for the learning and decision-making of the agent, facilitates optimization using algorithms such as reinforcement learning, and on this basis constructs an agent and an environment model based on the model-based policy optimization algorithm. Utilizing the simulated experience quickly generated by the environment model and allowing the agent to explore possible states and actions more extensively, it can adapt to the complex and changeable integrated energy system environment, improve the multi-energy scheduling ability in the integrated energy system, enhance the generalization ability of the agent, ensure the accuracy and reliability of the agent's scheduling of the integrated energy system, reduce the number of interactions between the agent and the real world, improve the sampling efficiency, and thus be more efficient than only learning from real experience. At the same time, it can also reduce the operation cost of the integrated energy system and the frequency and cost of manual intervention.
[0140] The present invention also provides a computer-readable storage medium, which stores a computer program that can be used to execute the steps of the scheduling method of the integrated energy system provided above. Figure 1 of the integrated energy system provided above.
[0141] The present invention also provides a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the steps of the scheduling method of the integrated energy system provided above. Figure 1 of the integrated energy system provided above.
[0142] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0143] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0144] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0146] It should be noted that the above specific implementation method can enable those skilled in the art to understand the invention more comprehensively, but does not limit the invention in any way. Therefore, although the present invention has been described in detail in this specification, those skilled in the art should understand that the invention can still be modified or replaced by equivalents; and all technical solutions and improvements that do not deviate from the spirit and scope of the invention are included in the protection scope of the patent for the invention. Any figure mark in the claims should not be regarded as limiting the claims involved.
Claims
1. A scheduling method for an integrated energy system, characterized in that: include: Obtaining the status information of the integrated energy system and the operation actions of the equipment; constructing a Markov decision process based on the integrated energy system; The Markov decision process includes a state space, a behavior space and a reward function; wherein the state space is a set of state information constructed according to the state information, the behavior space is a set of operating actions constructed according to the operating actions, and the reward function generates a reward according to the corresponding state space and behavior space, and the reward is negatively correlated with the operating cost of the integrated energy system; An agent of the model-based policy optimization algorithm MBPO is constructed according to the Markov decision process, an environmental model is constructed using a neural network, the agent interacts with the environmental model, and the agent's strategy is optimized by maximizing rewards; wherein the environmental model is used to predict the future value and reward of exogenous states, including load demand, energy price, and wind speed; According to the optimized strategy, the operation of each device in the integrated energy system is scheduled by the intelligent agent.
2. A scheduling method for an integrated energy system according to claim 1, characterized in that: The use of a neural network to construct an environment model comprises: An environment model is constructed using a deep neural network set, each deep neural network in the set performs independent predictions, and the deep neural network set consists of multiple deep neural networks DNN.
3. A scheduling method for an integrated energy system according to claim 2, characterized in that: The agent strategy of interacting with the environment model and optimizing the MBPO agent strategy by maximizing the reward includes: Based on pre-acquired real experience samples, the environment model is trained through the interaction between the agent and the environment model; The trained environment model is used to generate simulated experience samples, which are combined with real experience samples to optimize the strategy of the agent with the goal of maximizing the reward, thereby obtaining an optimized agent strategy.
4. A method for dispatching a comprehensive energy system according to claim 3, characterized in that: The training of the environment model through the interaction between the agent and the environment model includes: Initialize MBPO's policy network Actor and Q-value network Critic, as well as two experience playback buffers, which are used to store real experience samples and simulated experience samples respectively; Interacting the agent with the environment model, including: the agent inputting the current state of the integrated energy system into the policy network, simulating the state change of the integrated energy system based on the environment model, determining the probability distribution of each continuous behavior of the agent, randomly sampling actions from the probability distribution, and transferring the integrated energy system simulated by the environment model to the next state and receiving a reward; Multiple time steps of the cyclic interaction process are generated, multiple simulated experience samples consisting of state, behavior, reward and next state are generated, and the multiple simulated experience samples are stored in the experience replay buffer; The environment model is trained by minimizing the difference between the next state in the real experience sample and the next state in the simulated experience sample.
5. A method for dispatching a comprehensive energy system according to claim 4, characterized in that: The mean square error (MSE) is used to characterize the difference between the next state in the simulated experience sample and the real experience sample, and the DNN is trained with the goal of minimizing the difference.
6. A scheduling method for a comprehensive energy system according to claim 3, characterized in that: The strategy optimization of the agent with the goal of maximizing the reward includes: The strategy optimization of the agent is achieved by minimizing the real-time operation cost of the comprehensive energy system in each time slot.
7. A dispatching device for an integrated energy system, characterized in that: include: An acquisition module is used to obtain the status information of the integrated energy system and the operation actions of the equipment; A building module for building a Markov decision process according to the integrated energy system; The Markov decision process includes a state space, a behavior space and a reward function; wherein the state space is a set of state information constructed according to the state information, the behavior space is a set of operating actions constructed according to the operating actions, and the reward function generates a reward according to the corresponding state space and behavior space, and the reward is negatively correlated with the operating cost of the integrated energy system; An optimization module is used to construct an agent of the model-based policy optimization algorithm MBPO according to the Markov decision process, use a neural network to build an environmental model, interact the agent with the environmental model, and optimize the agent's strategy by maximizing rewards; wherein the environmental model is used to predict the future value and reward of exogenous states, including load demand, energy price and wind speed; The scheduling module is used to schedule the operation of each device in the integrated energy system through the intelligent agent according to the optimized strategy.
8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
9. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Comprehensive energy system low-carbon economic dispatching strategy based on CT-TD3 algorithm
CN117787609A
Low-carbon economic dispatching method and device for integrated energy system, and storage medium
CN117910775A
Comprehensive energy system economic dispatching model method based on deep reinforcement learning
CN119273066A
Cited By
Heat supply system and power grid interaction method, device and equipment and storage medium
CN120355199A
A multi-energy microgrid cluster distributed energy management method based on RC-MAPPO
CN122617053A