Operation scheduling method of combined energy storage power system and related device

By adopting the operation scheduling generation model with a dual Q network collaborative architecture in the joint energy storage power system, the problems of low operation efficiency and slow response speed in the existing technology are solved, and more efficient operation scheduling and supply and demand balance are achieved.

CN120073806AActive Publication Date: 2025-05-30GUANGDONG POWER GRID CO LTD DONGGUAN POWER SUPPLY BUREAU +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510250102.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-30
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

When the prior art uses energy storage technology to optimize the operation and scheduling of the power system, there are problems such as low operating efficiency, slow response speed, and non-convex optimization operation, especially in the optimization of wind storage and photoelectric storage combined power systems.

Method used

The operation scheduling generation model adopts a dual Q network collaborative architecture, and realizes effective operation scheduling by obtaining the operating status of the joint energy storage power system, screening candidate scheduling actions, and using the coupling relationship between the online Q network and the target Q network to separate the scheduling actions and predict the value of the action, to achieve effective operation scheduling.

Benefits of technology

The operation efficiency and response speed of the combined energy storage power system are improved, non-convex optimization operation is avoided, and a more effective supply and demand balance is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120073806A_ABST
    Figure CN120073806A_ABST
Patent Text Reader

Abstract

The invention provides an operation scheduling method of a combined energy storage power system and a related device, and relates to the field of power system optimization control. The method comprises the following steps: acquiring an operation state of the combined energy storage power system; based on the operation state, candidate scheduling actions conforming to the operation state are selected from preset scheduling actions; inputting the operation state and the candidate scheduling action into a pre-trained operation scheduling generation model to obtain a predicted action value of executing the candidate scheduling action in the operation state, the operation scheduling generation model being an online Q network obtained based on double-Q network training; and in the selected candidate scheduling actions, scheduling operation of the combined energy storage power system according to the candidate scheduling action with the maximum predicted action value. According to the invention, the operation of the combined energy storage power system is effectively scheduled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of optimal control of power systems, and particularly to an operation scheduling method and related devices for a combined energy storage power system. Background Art

[0002] The field of optimal control of power systems is an important field involving the operation, control, and optimization of power systems. Its core goal is to achieve the efficient, stable, and reliable operation of power systems through scientific methods. And the balance between supply and demand is an important guarantee for improving the stability and reliability of power systems. Moreover, with the rapid development of renewable energy such as wind energy and solar energy, the power system needs to be continuously improved to better cope with the challenges of supply-demand balance. Among them, energy storage technology, as an effective means, stores excess electric energy and releases it when needed, thereby balancing the supply and demand of the power system.

[0003] Currently, to solve the problem of optimizing the operation scheduling of power systems using energy storage technology, most traditional mathematical methods such as operations research optimization are adopted. However, for the optimization of combined energy storage power systems such as wind-storage and photovoltaic-storage combined power systems, if traditional mathematical methods are used, there are problems such as low operation efficiency, slow response speed, and non-convex optimization operation.

[0004] Therefore, there is an urgent need to provide an optimization scheme for effectively operating and scheduling combined energy storage power systems. Summary of the Invention

[0005] The present application provides an operation scheduling method and related devices for a combined energy storage power system to effectively schedule the operation of the combined energy storage power system.

[0006] In a first aspect, the present application provides an operation scheduling method for a combined energy storage power system, including:

[0007] Obtain the operation state of the combined energy storage power system;

[0008] Based on the operation state, select candidate scheduling actions that conform to the operation state from preset scheduling actions;

[0009] Input the operation state and the candidate scheduling actions into a pre-trained operation scheduling generation model to obtain the predicted action value of executing the candidate scheduling actions in the operation state. The operation scheduling generation model is an online Q network trained based on a double Q network;

[0010] Among the selected multiple candidate scheduling actions, schedule the operation of the combined energy storage power system according to the candidate scheduling action with the maximum predicted action value.

[0011] In a possible implementation, the operating state and candidate scheduling actions are input into a pre-trained operating scheduling generation model to obtain the predicted action value of executing the candidate scheduling actions in the operating state, including:

[0012] The operating state is input into the front-end network of the operating scheduling generation model for feature extraction to obtain high-dimensional features of the motion state. Among them, the front-end network includes network units, and the network units are used to extract the high-dimensional features of the input operating state;

[0013] The high-dimensional features are input into the flattening layer of the operating scheduling generation model for feature-to-vector conversion to obtain a one-dimensional feature vector corresponding to the high-dimensional features;

[0014] The one-dimensional feature vector is input into the value function network of the operating scheduling generation model for calculating the state estimation value to obtain the state estimation value corresponding to the operating state;

[0015] The one-dimensional feature vector and the candidate scheduling actions are input into the concatenation layer of the operating scheduling generation model for feature concatenation to obtain a joint feature vector;

[0016] The joint feature vector is input into the advantage function network of the operating scheduling generation model for calculating the scheduling action advantage value to obtain the advantage value of executing the candidate scheduling actions in the operating state;

[0017] The state estimation value and the advantage value of the candidate scheduling actions are input into the output layer of the operating scheduling generation model for summation calculation to obtain the predicted action value corresponding to executing the candidate scheduling actions in the operating state.

[0018] In a possible implementation, the pre-trained operating scheduling generation model is obtained through the following method:

[0019] Obtain the historical operating information of the combined energy storage power system. The historical operating information includes the operating state of the combined energy storage power system at time t-1, the scheduling actions executed in the operating state at time t-1, the immediate reward obtained by executing the scheduling actions, and the operating state at time t;

[0020] According to the historical operating information, generate a sample data set containing Markov Decision Process (MDP) data units. The MDP data units include the operating state at time t-1, the scheduling actions executed in the operating state at time t-1, the immediate reward obtained by executing the scheduling actions corresponding to the operating state at time t-1, and the operating state at time t;

[0021] According to the sample data set, train the operating scheduling generation model to obtain the pre-trained operating scheduling generation model.

[0022] In a possible implementation manner, according to historical operation information, a sample data set including Markov decision process (MDP) data units is generated, including:

[0023] Construct an objective function aiming at the operation scheduling optimization of the combined energy storage power system and define a set of constraint conditions for the objective function. The set of constraint conditions includes the electric field output constraint condition and the energy storage power station output constraint condition;

[0024] According to the set of constraint conditions, screen out the target scheduling actions that meet the constraint conditions from the executed scheduling actions;

[0025] Based on the objective function, calculate the immediate reward of each target scheduling action, where the immediate reward is related to the optimization objective of the objective function;

[0026] Encode the operation state at time t - 1, the operation state at time t, each target scheduling action, and the immediate reward of each target scheduling action into MDP data units;

[0027] Aggregate multiple MDP data units to generate a sample data pool;

[0028] Extract a set number of MDP data units from the sample data pool to generate a sample data set.

[0029] In a possible implementation manner, extracting a set number of MDP data units from the sample data pool to generate a sample data set includes:

[0030] Randomly extract a set number of MDP data units from the sample data pool to generate a sample data set;

[0031] Or,

[0032] Determine the priority of each MDP data unit in the sample data pool;

[0033] According to the priority, extract a set number of MDP data units from the sample data pool to generate a sample data set.

[0034] In a possible implementation manner, the double Q - network includes an online Q - network and a target Q - network. The online Q - network and the target Q - network have the same network architecture. According to the sample data set, train the operation scheduling generation model to obtain a pre - trained operation scheduling generation model, including:

[0035] Input the operation state at time t - 1 into the online Q - network to determine the predicted action value corresponding to each target scheduling action;

[0036] Input the operation state at time t into the target Q - network, and according to the immediate reward, determine the output value of the target Q - network. The output value is the target action value;

[0037] Calculate the mean squared error loss between the predicted action value and the target action value, and update the parameters of the online Q-network through an optimization algorithm to minimize the mean squared error loss;

[0038] Based on a preset training end condition, obtain a pre-trained operation scheduling generation model according to the updated parameters of the online Q-network.

[0039] In a possible implementation manner, based on a preset training end condition, obtaining a pre-trained operation scheduling generation model according to the updated parameters of the online Q-network includes:

[0040] Determine whether the number of updates of the parameters of the online Q-network reaches a threshold number of times;

[0041] When the number of updates of the parameters of the online Q-network reaches the threshold number of times, synchronize the parameters of the online Q-network to the target Q-network based on an update strategy;

[0042] Determine whether the online Q-network is in a converged state;

[0043] When the online Q-network is in a converged state, use the current online Q-network as the pre-trained operation scheduling generation model;

[0044] When the online Q-network is not in a converged state, continue iterative training until the online Q-network is in a converged state.

[0045] In a possible implementation manner, when the number of updates of the parameters of the online Q-network reaches the threshold number of times, synchronizing the parameters of the online Q-network to the target Q-network based on an update strategy includes:

[0046] When the number of updates of the parameters of the online Q-network reaches the threshold number of times, copy the parameters of the online Q-network to the target Q-network;

[0047] Or,

[0048] When the number of updates of the parameters of the online Q-network reaches the threshold number of times, perform a weighted sum of the parameters of the online Q-network and the parameters of the current target Q-network according to a set weight coefficient to obtain the parameters of the target Q-network.

[0049] In a second aspect, the present application provides an operation scheduling device for a combined energy storage power system, including:

[0050] An acquisition module, configured to acquire the operation state of the combined energy storage power system;

[0051] A selection module, configured to select a candidate scheduling action that conforms to the operation state from preset scheduling actions based on the operation state;

[0052] A determination module, configured to input the operating state and candidate scheduling actions into a pre-trained operating scheduling generation model, and obtain a predicted action value for executing the candidate scheduling actions in the operating state. The operating scheduling generation model is an online Q network trained based on a double Q network;

[0053] A scheduling module, configured to schedule the operation of the combined energy storage power system according to the candidate scheduling action with the maximum predicted action value among the selected multiple candidate scheduling actions.

[0054] In a possible implementation manner, the determination module is specifically configured to:

[0055] Input the operating state into the front-end network of the operating scheduling generation model for feature extraction to obtain high-dimensional features of the motion state. The front-end network includes network units for extracting high-dimensional features of the input operating state;

[0056] Input the high-dimensional features into the flattening layer of the operating scheduling generation model for feature-to-vector conversion to obtain a one-dimensional feature vector corresponding to the high-dimensional features;

[0057] Input the one-dimensional feature vector into the value function network of the operating scheduling generation model for calculating the state estimation value to obtain the state estimation value corresponding to the operating state;

[0058] Input the one-dimensional feature vector and the candidate scheduling actions into the splicing layer of the operating scheduling generation model for feature splicing to obtain a combined feature vector;

[0059] Input the combined feature vector into the advantage function network of the operating scheduling generation model for calculating the scheduling action advantage value to obtain the advantage value of executing the candidate scheduling actions in the operating state;

[0060] Input the state estimation value and the advantage value of the candidate scheduling actions into the output layer of the operating scheduling generation model for summation calculation to obtain the predicted action value corresponding to executing the candidate scheduling actions in the operating state.

[0061] In a possible implementation manner, the pre-trained operating scheduling generation model is obtained through the following method:

[0062] Obtain the historical operation information of the combined energy storage power system, where the historical operation information includes the operating state of the combined energy storage power system at time t-1, the scheduling actions executed in the operating state at time t-1, the immediate reward obtained by executing the scheduling actions, and the operating state at time t;

[0063] Generate a sample data set containing Markov decision process (MDP) data units based on historical operation information. The MDP data units include the operation state at time t-1, the scheduling action executed under the operation state at time t-1, the immediate reward obtained by executing the scheduling action corresponding to the operation state at time t-1, and the operation state at time t.

[0064] Train an operation scheduling generation model based on the sample data set to obtain a pre-trained operation scheduling generation model.

[0065] In a possible implementation manner, the operation scheduling device of the combined energy storage power system further includes a processing module, and the processing module is specifically configured to:

[0066] Construct an objective function with the optimization of the operation scheduling of the combined energy storage power system as the goal and define a set of constraint conditions for the objective function. The set of constraint conditions includes an electric field output constraint condition and a storage power station output constraint condition;

[0067] According to the set of constraint conditions, screen out the target scheduling actions that meet the constraint conditions from the executed scheduling actions;

[0068] Based on the objective function, calculate the immediate reward of each target scheduling action, where the immediate reward is related to the optimization goal of the objective function;

[0069] Encode the operation state at time t-1, the operation state at time t, each target scheduling action, and the immediate reward of each target scheduling action into MDP data units;

[0070] Aggregate multiple MDP data units to generate a sample data pool;

[0071] Extract a set number of MDP data units from the sample data pool to generate a sample data set.

[0072] In a possible implementation manner, the processing module is further configured to:

[0073] Randomly extract a set number of MDP data units from the sample data pool to generate a sample data set;

[0074] Or,

[0075] Determine the priority of each MDP data unit in the sample data pool;

[0076] According to the priority, extract a set number of MDP data units from the sample data pool to generate a sample data set.

[0077] In a possible implementation manner, the double Q-network includes an online Q-network and a target Q-network, and the online Q-network and the target Q-network have the same network architecture. The determination module is specifically configured to:

[0078] Input the operating state at time t-1 into the online Q-network to determine the predicted action values corresponding to each target scheduling action;

[0079] Input the operating state at time t into the target Q-network, and determine the output value of the target Q-network according to the immediate reward. The output value is the target action value;

[0080] Calculate the mean square error loss between the predicted action value and the target action value, and update the parameters of the online Q-network through an optimization algorithm to minimize the mean square error loss;

[0081] Based on the preset training end condition, obtain the pre-trained operation scheduling generation model according to the updated parameters of the online Q-network.

[0082] In a possible implementation manner, the determination module is specifically configured to:

[0083] Judge whether the update times of the parameters of the online Q-network reach the times threshold;

[0084] When the update times of the parameters of the online Q-network reach the times threshold, synchronize the parameters of the online Q-network to the target Q-network based on the update strategy;

[0085] Judge whether the online Q-network is in a converged state;

[0086] When the online Q-network is in a converged state, use the current online Q-network as the pre-trained operation scheduling generation model;

[0087] When the online Q-network is not in a converged state, continue iterative training until the online Q-network is in a converged state.

[0088] In a possible implementation manner, the processing module is specifically configured to:

[0089] When the update times of the parameters of the online Q-network reach the times threshold, copy the parameters of the online Q-network to the target Q-network;

[0090] Or,

[0091] When the update times of the parameters of the online Q-network reach the times threshold, perform weighted summation on the parameters of the online Q-network and the parameters of the current target Q-network according to the set weight coefficient to obtain the parameters of the target Q-network.

[0092] In a third aspect, the present application provides an electronic device, including: a memory, a processor;

[0093] The memory stores computer execution instructions;

[0094] The processor executes the computer-executable instructions stored in the memory, such that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0095] In a fourth aspect, the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the above first aspect and / or various possible implementations of the first aspect when executed by a processor.

[0096] In a fifth aspect, the present application provides a computer program product including a computer program, which implements the above first aspect and / or various possible implementations of the first aspect when executed by a processor.

[0097] The operation scheduling method and related devices for a combined energy storage power system provided by the present application relate to the optimal control of a power system. The method includes: obtaining the operation state of the combined energy storage power system; based on the operation state, selecting candidate scheduling actions that conform to the operation state from preset scheduling actions; inputting the operation state and the candidate scheduling actions into a pre-trained operation scheduling generation model to obtain the predicted action value of executing the candidate scheduling actions in the operation state, where the operation scheduling generation model is an online Q-network trained based on a double Q-network; among the selected multiple candidate scheduling actions, scheduling the operation of the combined energy storage power system according to the candidate scheduling action with the maximum predicted action value. After obtaining the operation state of the combined energy storage power system, the present application filters out candidate scheduling actions that conform to the operation state based on the preset scheduling actions. This filtering process restricts the action exploration space and avoids interference caused by invalid scheduling actions to the estimation of the predicted action value. When evaluating the predicted action value of the candidate scheduling actions, by constructing a double Q-network collaborative architecture, the coupling relationship between the scheduling action selection and the predicted action value is separated. Among them, the online Q-network is used to filter out the candidate scheduling action with the maximum predicted action value that conforms to the operation state from the candidate scheduling actions, and another Q-network is used to perform target action value evaluation on the candidate actions based on the updated independent parameters. This division mechanism cuts off the error transmission path. Select the candidate scheduling action with the maximum predicted action value among multiple candidate scheduling actions, and according to the candidate scheduling action with the maximum predicted action value, effectively schedule the operation of the combined energy storage power system. Description of the Drawings

[0098] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0099] Figure 1 Flow diagram of the operation scheduling method for the combined energy storage power system provided by the present application Figure 1 ;

[0100] Figure 2 Flow schematic of the operation scheduling method for the wind-solar energy storage power system provided by the embodiment of the present application Figure 1 ;

[0101] Figure 3 Structural schematic diagram of the operation scheduling device for the combined energy storage power system provided by the present application;

[0102] Figure 4 Structural schematic diagram of the electronic device provided by an embodiment of the present application.

[0103] Through the above-mentioned drawings, the specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Specific embodiments

[0104] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of the devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0105] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0106] At present, to solve the problem of optimizing the operation and dispatching of power systems by means of energy storage, most traditional mathematical methods such as operational optimization are adopted. However, when using traditional mathematical methods to optimize wind-storage and photovoltaic-storage combined power stations in a market environment, there are problems such as low operation efficiency, slow response speed, and non-convex optimization operation. To address these challenges, in recent years, artificial intelligence algorithms have been used to optimize the operation of scenarios such as wind, photovoltaic, and storage combined power stations, achieving good results. The field of artificial intelligence applications has developed rapidly in recent years. The Deep Reinforcement Learning (DRL) algorithm is one of the algorithms that has been widely studied and applied. Using the DRL algorithm to optimize the operation and dispatching problem of wind, photovoltaic, and storage combined power stations also presents advantages such as high dispatching operation efficiency, fast response speed, and accurate output decisions, and has certain application prospects.

[0107] However, existing DRL algorithms such as Proximal Policy Optimization (PPO), Deep Deterministic Policy Gradient (DDPG), and Deep Q-Network (DQN) still have difficulties in dealing with high-dimensional state spaces and action spaces, which will limit the application of DRL algorithms in complex power systems.

[0108] To address the above problems, this application proposes an operation and dispatching method for a combined energy storage power system. By constructing a dual Q-network collaborative architecture, the coupling relationship between dispatching action selection and predicted action value is separated, and the error transmission path is cut off to effectively dispatch the operation of the combined energy storage power system.

[0109] The following will specifically describe the technical solutions of this application and how the technical solutions of this application solve the above technical problems through specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0110] Figure 1 Flow schematic of the operation and dispatching method for the combined energy storage power system provided by this application Figure 1 As Figure 1 shown, the method includes:

[0111] S101. Obtain the operation state of the combined energy storage power system.

[0112] Among them, the combined energy storage power system refers to a power energy storage system formed by combining multiple different types of energy storage technologies or devices, which work together to improve the stability, reliability and efficiency of the power system, and at the same time meet diverse power demands.

[0113] The operating state of the combined energy storage power system refers to a comprehensive description of the working conditions and performance of the combined energy storage power system at a specific moment. The operating state reflects the real-time working conditions of each energy storage device, power equipment and the overall combined energy storage power system. The operating state includes but is not limited to the ratio of the current remaining power to the total power of the energy storage device, the charge and discharge power, and the charge and discharge efficiency.

[0114] S102. Based on the operating state, select candidate scheduling actions that conform to the operating state from the preset scheduling actions.

[0115] In this step, it can be understood that the preset scheduling actions refer to the operations that the combined energy storage power system can perform. By analyzing the operating state, it is possible to determine which scheduling actions in the preset scheduling actions are feasible and effective, and select candidate scheduling actions that conform to the operating state from them. For example, assume that the state of charge of the energy storage device is low and the grid frequency is low. At this time, select the preset scheduling action of discharging as the candidate scheduling action.

[0116] S103. Input the operating state and the candidate scheduling actions into a pre-trained operating scheduling generation model to obtain the predicted action value of executing the candidate scheduling actions in the operating state. The operating scheduling generation model is an online Q-network trained based on the double Q-network.

[0117] After determining the candidate scheduling actions that conform to the operating state through S102, it is necessary to input the determined operating state and the candidate scheduling actions into a pre-trained operating scheduling generation model, so as to obtain the predicted action value of executing the candidate scheduling actions in the operating state.

[0118] Exemplarily, the operating state and candidate scheduling actions are input into a pre-trained operating scheduling generation model to obtain the predicted action value of executing the candidate scheduling actions in the operating state, including: inputting the operating state into the front-end network of the operating scheduling generation model for feature extraction to obtain the high-dimensional features of the motion state, where the front-end network includes network units for extracting the high-dimensional features of the input operating state; inputting the high-dimensional features into the flattening layer of the operating scheduling generation model for feature-to-vector conversion to obtain the one-dimensional feature vector corresponding to the high-dimensional features; inputting the one-dimensional feature vector into the value function network of the operating scheduling generation model for calculating the state estimation value to obtain the state estimation value corresponding to the operating state; inputting the one-dimensional feature vector and the candidate scheduling actions into the concatenation layer of the operating scheduling generation model for feature concatenation to obtain the joint feature vector; inputting the joint feature vector into the advantage function network of the operating scheduling generation model for calculating the scheduling action advantage value to obtain the advantage value of executing the candidate scheduling actions in the operating state; inputting the state estimation value and the advantage value of the candidate scheduling actions into the output layer of the operating scheduling generation model for summation calculation to obtain the predicted action value corresponding to executing the candidate scheduling actions in the operating state.

[0119] In this example, it can be understood that the pre-trained operating scheduling generation model is an online Q-network trained based on the double Q-network.

[0120] S104. Among the selected multiple candidate scheduling actions, schedule the operation of the combined energy storage power system according to the candidate scheduling action with the maximum predicted action value.

[0121] In S104, it can be understood that different candidate scheduling actions correspond to different predicted action values. The candidate scheduling action with the maximum predicted action value is selected from the predicted action values corresponding to multiple different candidate scheduling actions, and the candidate scheduling action with the maximum predicted action value is used as the operation scheduling scheme of the combined energy storage power system, so as to effectively schedule the operation of the combined energy storage power system.

[0122] After the operation state of the combined energy storage power system is obtained in the embodiment of the present application, candidate scheduling actions that conform to the operation state are screened based on preset scheduling actions. This screening process restricts the action exploration space to avoid interference of invalid scheduling actions on the prediction of action value estimation. When evaluating the predicted action value of the candidate scheduling actions, a dual Q-network collaborative architecture is constructed to separate the coupling relationship between the selection of scheduling actions and the prediction of action value. Among them, the online Q-network is used to screen out candidate scheduling actions with the maximum predicted action value that conform to the operation state from the candidate scheduling actions, and the other Q-network is used to evaluate the target action value of the candidate actions based on the updated independent parameters. This division of labor mechanism cuts off the error transmission path. The candidate scheduling action with the maximum predicted action value is selected from multiple candidate scheduling actions, and the operation of the combined energy storage power system is effectively scheduled according to the candidate scheduling action with the maximum predicted action value.

[0123] Based on the above embodiment, the pre-trained operation scheduling generation model is obtained in the following manner: Obtain the historical operation information of the combined energy storage power system, where the historical operation information includes the operation state of the combined energy storage power system at time t-1, the scheduling actions executed under the operation state at time t-1, the immediate reward obtained by executing the scheduling actions, and the operation state at time t; According to the historical operation information, generate a sample data set containing MDP data units, where the MDP data unit includes the operation state at time t-1, the scheduling actions executed under the operation state at time t-1, the immediate reward obtained by executing the scheduling actions corresponding to the operation state at time t-1, and the operation state at time t; According to the sample data set, train the operation scheduling generation model to obtain the pre-trained operation scheduling generation model.

[0124] In this embodiment, it can be understood that before training the operation scheduling generation model, it is necessary to determine the sample data set, and the sample data set is generated based on the historical operation information of the combined energy storage power system.

[0125] Specifically, generating a sample data set containing MDP data units according to the historical operation information includes: constructing an objective function with the optimization of the operation scheduling of the combined energy storage power system as the goal and defining a set of constraint conditions for the objective function, where the set of constraint conditions includes the electric field output constraint condition and the energy storage power station output constraint condition; According to the set of constraint conditions, screen out the target scheduling actions that conform to the constraint conditions from the executed scheduling actions; Based on the objective function, calculate the immediate reward of each target scheduling action, where the immediate reward is related to the optimization goal of the objective function; Encode the operation state at time t-1, the operation state at time t, each target scheduling action, and the immediate reward of each target scheduling action into MDP data units; Aggregate multiple MDP data units to generate a sample data pool; Extract a set number of MDP data units from the sample data pool to generate a sample data set.

[0126] Among them, the objective function constructed with the optimization of the operation and dispatch of the combined energy storage power system can be constructed according to actual needs.

[0127] In one example, taking the wind-solar combined energy storage power system as an example, an objective function with the minimization of the total operating cost of the wind-solar combined energy storage power system is constructed. Among them, the total cost includes the wind-solar tracking assessment cost, the curtailment cost of wind and light, and the minimum of the energy storage operation cost. The total cost calculation formula is as follows: ; where C wpb is the total combined cost of the wind-solar combined energy storage power system, C k is the wind-solar tracking assessment cost, C q is the curtailment cost of wind and light, C bt is the energy storage operation cost.

[0128] The set of constraint conditions includes the output constraint conditions of the wind farm, the output constraint conditions of the photovoltaic power plant, and the output constraint conditions of the energy storage power station. The output constraint conditions of the wind farm are:

[0129]

[0130] where, V wt (t) is the difference in the output power of the wind farm at time t and time t - 1; P wt (t) is the output power of the wind farm at time t - 1; V wtmax is the maximum value of the theoretical output power of the wind farm.

[0131] The output constraint conditions of the photovoltaic power plant are:

[0132]

[0133] where, V pv (t) is the difference in the output power of the photovoltaic power plant at time t and time t - 1; P pv (t) is the output power of the photovoltaic power plant at time t - 1; V pvmax is the maximum value of the theoretical output power of the photovoltaic power plant.

[0134] The output constraint conditions of the energy storage power station are:

[0135]

[0136] where, P btmax is the maximum charge-discharge power of the energy storage device in the energy storage power station; H socmin and H socmax are the upper and lower limits of the state of charge (SOC) of the energy storage; H soc (t) is the state of charge of the energy storage device in the energy storage power station at time t.

[0137] Furthermore, the operating state at time t-1, the operating state at time t, each target scheduling action and the immediate reward of each target scheduling action are encoded into an MDP data unit. The MDP data unit can be represented as a multi-tuple, namely S, A, , R, where S represents the Markov state space; A represents the scheduling action space; represents the discount factor, which is between 0 and 1 and determines the impact of future rewards; R represents the reward function.

[0138] Establish Markov state S: The combined power plant tracks the planned value S plan 、Charge and discharge power of energy storage S bt , S soc , Wind power prediction processing wt And the predicted output of photovoltaic pv As a state space with five input channels, it is expressed as: .

[0139] Establish dispatch action space A: Increase wind power output by A wt , Photovoltaic output increment A pv 、Energy storage output increment A bt As a scheduling action space, it is expressed as: .

[0140] Create a reward function: For the agent in a certain state Select Schedule Action Instant reward at the time. In the entire scheduling period T, the cumulative reward function is: The dispatch cycle refers to the time range from when the wind-solar-energy storage system starts from one state to when it enters a new state after several dispatches, such as one hour, one day, or a time range where the interval between continuous dispatch actions does not exceed the preset interval or other time periods with business significance.

[0141] In this embodiment, the first scheduling instruction execution state s 0 Start to end of scheduling operation state s n , forming MDP data units one by one according to the n-time scheduling order , , load into the sample data pool. In the entire scheduling period T, the scheduling period reward function of the sample is:

[0142] .

[0143] Among them, r + and r -They represent positive and negative reward values respectively, and tar is the scheduling target. The scheduling cycle reward can be used to determine the priority weights of samples.

[0144] In this embodiment, by generating a sample data set containing MDP data units according to historical operation information, the operation rules and dynamic characteristics of the actual combined energy storage power system can be reflected, making the trained operation scheduling generation model closer to the real scenario, which helps to improve the effective operation of the combined energy storage power system.

[0145] Furthermore, after aggregating multiple MDP data units to generate a sample data pool, it is necessary to extract the MDP data units of the set data from the sample data pool to generate a sample data set. Specifically, extracting a set number of MDP data units from the sample data pool to generate a sample data set includes: randomly extracting a set number of MDP data units from the sample data pool to generate a sample data set; or, determining the priority of each MDP data unit in the sample data pool; and according to the priority, extracting a set number of MDP data units from the sample data pool to generate a sample data set.

[0146] This means that there are different implementation methods for generating the sample data set. The first implementation method is to randomly extract a set number of MDP data units from the sample data pool, where the set number can be set according to actual needs. For example, the number can be set to 3.

[0147] The second implementation method first needs to determine the priority of each MDP data unit in the sample data pool, and then extract a set number of MDP data units from the sample data pool according to the priority. Among them, the calculation of the priority can be carried out through the weight calculation rule, and the weight calculation rule can be comprehensively evaluated based on dimensions such as the reduction range of the operating cost, the reduction degree of the curtailment rate of wind and light, and the improvement of the energy storage utilization rate. For example, for a sample in a scheduling cycle with a significant reduction in operating cost and a high energy storage utilization rate, its priority weight will be set relatively high so that it can be sampled more frequently in subsequent training.

[0148] By providing different ways to generate the sample data set in the embodiments of the present application, the flexibility of the operation scheduling method of the combined energy storage power system can be improved.

[0149] Further, the double Q-network described in S103 includes an online Q-network and a target Q-network. The online Q-network and the target Q-network have the same network architecture. According to the sample data set, the operation scheduling generation model is trained to obtain a pre-trained operation scheduling generation model, including: inputting the operation state at time t-1 into the online Q-network to determine the predicted action values corresponding to each target scheduling action; inputting the operation state at time t into the target Q-network, and determining the output value of the target Q-network according to the immediate reward, where the output value is the target action value; calculating the mean square error loss between the predicted action value and the target action value, and updating the parameters of the online Q-network through an optimization algorithm to minimize the mean square error loss; based on the preset training end condition, obtaining a pre-trained operation scheduling generation model according to the updated parameters of the online Q-network.

[0150] In this embodiment, it can be understood that both the online Q-network and the target Q-network are multi-input channels and double-output subnets. Among them, the number of input channels is the same as the number of channels of the previously established state space S. The double-output subnet includes a front-end network and an output subnet. The input of the network enters the output subnet through the front-end network. The output subnet includes a value function network Vn and an advantage function network An. The value function network is responsible for evaluating the value of the operation state s, and the advantage function network is responsible for evaluating the advantages and disadvantages of each scheduling action in the operation state s. The final network value is:

[0151]

[0152] Among them, represents the network parameters of the double-output subnet; a, respectively represent the parameters of the value function network Vn and the advantage function network An; A represents the set of all target scheduling actions; represents the target scheduling action of the next operation state, .

[0153] The front-end network is composed of multiple network units. Each layer of network units is composed of a convolutional layer, a normalization layer, and an activation function. The output subnet is built by a fully connected layer. The front-end network is connected to the output subnet through a flattening layer.

[0154] After determining the structure of the double Q-network, it is necessary to train the operation scheduling generation model according to the sample data set to obtain a pre-trained operation scheduling generation model. Specifically, it is necessary to calculate the predicted action value and the target action value for each training data in the sample data set, that is, input the operation state at time t-1 into the online Q-network to determine the predicted action value corresponding to each target scheduling action, that is, the Q value, input the operation state at time t into the target Q-network, and combine the immediate reward value r to obtain the target action value, that is, the Y value.

[0155] The loss function for the predicted action value and the target action value is expressed as follows:

[0156] .

[0157] in, .

[0158] j is the training step length, E[] 2 represents the mean square error calculation, Indicates the next running state s of the online Q network in the dual Q network i+1 The most valuable action is a m ; represents the discount factor; Represents the sample data in the sample dataset.

[0159] It should be noted that the online Q network and the target Q network need to be initialized before training the dual Q network.

[0160] After determining the average error loss between the predicted action value and the target action value, the optimization algorithm is used to update the parameters of the online Q network to minimize the mean square error loss. Then, based on the preset training end conditions, the pre-trained operation scheduling generation model is obtained according to the updated parameters of the online Q network.

[0161] The embodiment of the present application separates the coupling relationship between scheduling action selection and predicted action value by constructing a dual Q network collaborative architecture, wherein the online Q network is used to screen out candidate scheduling actions with the maximum predicted action value that meet the operating state from candidate scheduling actions, and the target Q network is used to evaluate the target action value of the candidate actions based on the updated independent parameters. This division of labor mechanism cuts off the error transmission path, thereby improving the accuracy of the operation scheduling generation model.

[0162] Furthermore, based on preset training end conditions, a pre-trained operation scheduling generation model is obtained according to the parameters of the updated online Q network, including: judging whether the number of updates of the parameters of the online Q network reaches a number threshold; when the number of updates of the parameters of the online Q network reaches the number threshold, based on the update strategy, synchronizing the parameters of the online Q network to the target Q network; judging whether the online Q network is in a convergence state; when the online Q network is in a convergence state, using the current online Q network as a pre-trained operation scheduling generation model; when the online Q network is not in a convergence state, continuing iterative training until the online Q network is in a convergence state.

[0163] In this embodiment, it can be understood that when determining the pre-trained operation scheduling generation model, it is necessary to judge whether the number of updates of the parameters of the online Q-network reaches the number threshold, or it can also be understood as judging whether the training frequency of the parameters of the online Q-network reaches the frequency threshold.

[0164] When the number of updates of the parameters of the online Q-network reaches the number threshold, based on the update strategy, synchronize the parameters of the online Q-network to the target Q-network. Exemplarily, when the number of updates of the parameters of the online Q-network reaches the number threshold, based on the update strategy, synchronize the parameters of the online Q-network to the target Q-network, including: when the number of updates of the parameters of the online Q-network reaches the number threshold, copy the parameters of the online Q-network to the target Q-network; or, when the number of updates of the parameters of the online Q-network reaches the number threshold, perform weighted summation on the parameters of the online Q-network and the parameters of the current target Q-network according to the set weight coefficient to obtain the parameters of the target Q-network.

[0165] In this example, it can be understood that the update strategy includes the following methods: directly copying the parameters of the online Q-network directly into the target Q-network, that is, hard update, or performing weighted summation on the parameters of the online Q-network and the parameters of the current target Q-network according to the set weight coefficient to obtain the parameters of the target Q-network, that is, soft update. This example provides different parameter update methods to meet different user needs, thereby enhancing the flexibility of the joint energy storage power system operation scheduling method.

[0166] Furthermore, after synchronizing the parameters of the online Q-network to the target Q-network, it is also necessary to determine whether the online Q-network is in a convergent state and perform corresponding measures according to whether the online Q-network is in a convergent state. When the online Q-network is in a convergent state, use the current online Q-network as the pre-trained operation scheduling generation model; when the online Q-network is not in a convergent state, continue iterative training until the online Q-network is in a convergent state.

[0167] Furthermore, judging whether the online Q-network is in a convergent state is also to judge whether the online Q-network has reached a stable state, that is, the parameters or output values of the online Q-network no longer change significantly and approach a stable value. The specific implementation method for judging whether the online Q-network is in a convergent state can be selected according to the actual scenario.

[0168] In one implementation, observe the change range of the predicted action value of the online Q-network, that is, the change range of the Q value. If the Q value fluctuates very little or tends to be stable within a period of time, it can be considered that the online Q-network converges.

[0169] In another implementation, check the loss function value of the online Q-network. If the loss function value of the online Q-network gradually decreases and tends to be stable, it can be explained that the online Q-network converges.

[0170] Next, taking the combined energy storage power system as the wind-solar energy storage power system and the dual Q-network as the dueling double deep reinforcement learning (D3QN) network as an example, this paper will explain how to use the operation scheduling method of the combined energy storage power system provided by the embodiments of this application. Figure 2 The flow chart of the operation scheduling method of the wind-solar energy storage power system provided by the embodiments of this application is shown in Figure 2 Figure , and the method includes the following steps: Figure 1 as Figure 2 shown, and the method includes the following steps:

[0171] 1. Collect operation information from the wind-solar combined energy storage system and record it as data;

[0172] 2. Establish an objective function, with the minimum of the wind-solar tracking assessment cost, the curtailment cost of wind and solar energy, and the energy storage operation cost as the objective function;

[0173] 3. Determine the constraint conditions, which include the output constraint of the wind farm, the output constraint of the photovoltaic power station, and the output constraint of the energy storage power station;

[0174] 4. Based on the objective function and the constraint conditions, process the operation information collected in step 1 into Markov decision process data units to form a sample data pool. The specific construction principle has been elaborated in detail in the previous embodiments, so it will not be repeated here;

[0175] 5. Construct an online Q-network and a target Q-network, where both Q-networks have a network architecture with multiple input channels and dual output subnets;

[0176] 6. Initialize the online Q-network and the target Q-network;

[0177] 7. Extract several Markov decision process data units from the sample data pool in step 4 to train the online Q-network, and copy the parameters of the online Q-network to the parameters of the target Q-network at a fixed frequency;

[0178] 8. Calculate the predicted action value and the target action value for the sample data in the sample data set, and update the parameters of the online Q-network through an optimization algorithm to minimize the mean square error loss between the predicted action value and the target action value;

[0179] 9. Repeat steps 7 and 8 until the online Q-network converges, and then end the model training;

[0180] 10. Call the trained online Q-network, use it as the knowledge network, determine the corresponding predicted action values for multiple candidate scheduling actions, determine the candidate scheduling action with the maximum predicted action value as the operation scheduling method for the wind-solar combined energy storage power system, and output it, so as to obtain the optimal strategy with the minimum wind-solar tracking assessment cost, wind and light abandonment cost, and energy storage operation cost, optimize the operation efficiency of the wind-solar combined energy storage system, and improve economic benefits.

[0181] In summary, the embodiment of the present application proposes a DRL algorithm improved on the basis of the DQN basic network structure, that is, the D3QN model, which can avoid the overestimation problem in the optimization process of traditional algorithms and DQN algorithms, obtain more accurate results, improve the operation efficiency at the same time, speed up the operation response speed, optimize the scheduling operation strategy of the wind-solar combined energy storage system, reduce the operation cost, and improve economic benefits.

[0182] Figure 3 It is a schematic structural diagram of the operation scheduling device for the combined energy storage power system provided by the present application, as Figure 3 shown, the operation scheduling device 300 for the combined energy storage power system provided in this embodiment includes:

[0183] An acquisition module 301, configured to acquire the operation state of the combined energy storage power system;

[0184] A selection module 302, configured to select candidate scheduling actions that conform to the operation state from preset scheduling actions based on the operation state;

[0185] A determination module 303, configured to input the operation state and the candidate scheduling actions into a pre-trained operation scheduling generation model to obtain the predicted action value of executing the candidate scheduling action in the operation state, and the operation scheduling generation model is an online Q-network trained based on the double Q-network;

[0186] A scheduling module 304, configured to schedule the operation of the combined energy storage power system according to the candidate scheduling action with the maximum predicted action value among the selected multiple candidate scheduling actions.

[0187] In a possible implementation manner, the determination module 303 is specifically configured to:

[0188] Input the operation state into the front-end network of the operation scheduling generation model for feature extraction to obtain the high-dimensional features of the motion state, where the front-end network includes network units for extracting the high-dimensional features of the input operation state;

[0189] Input the high-dimensional features into the flattening layer of the operation scheduling generation model for feature-to-vector conversion to obtain the one-dimensional feature vector corresponding to the high-dimensional features;

[0190] Input the one-dimensional feature vector into the value function network of the operation scheduling generation model to calculate the state estimation value, and obtain the state estimation value corresponding to the operation state;

[0191] Input the one-dimensional feature vector and the candidate scheduling action into the splicing layer of the operation scheduling generation model for feature splicing to obtain a joint feature vector;

[0192] Input the joint feature vector into the advantage function network of the operation scheduling generation model to calculate the scheduling action advantage value, and obtain the advantage value of executing the candidate scheduling action in the operation state;

[0193] Input the state estimation value and the advantage value of the candidate scheduling action into the output layer of the operation scheduling generation model for summation calculation to obtain the predicted action value corresponding to executing the candidate scheduling action in the operation state.

[0194] In a possible implementation manner, the pre-trained operation scheduling generation model is obtained through the following method:

[0195] Obtain the historical operation information of the combined energy storage power system, where the historical operation information includes the operation state of the combined energy storage power system at time t-1, the scheduling action executed in the operation state at time t-1, the immediate reward obtained by executing the scheduling action, and the operation state at time t;

[0196] According to the historical operation information, generate a sample data set containing Markov decision process (MDP) data units. The MDP data unit includes the operation state at time t-1, the scheduling action executed in the operation state at time t-1, the immediate reward obtained by executing the scheduling action corresponding to the operation state at time t-1, and the operation state at time t;

[0197] According to the sample data set, train the operation scheduling generation model to obtain the pre-trained operation scheduling generation model.

[0198] In a possible implementation manner, the operation scheduling device of the combined energy storage power system further includes a processing module (not shown), and the processing module is specifically used for:

[0199] Construct an objective function with the optimization of the operation scheduling of the combined energy storage power system as the goal and define a set of constraint conditions for the objective function. The set of constraint conditions includes the electric field output constraint condition and the energy storage power station output constraint condition;

[0200] According to the set of constraint conditions, screen out the target scheduling actions that meet the constraint conditions from the executed scheduling actions;

[0201] Based on the objective function, calculate the immediate reward of each target scheduling action, where the immediate reward is related to the optimization goal of the objective function;

[0202] Encode the operating state at time t-1, the operating state at time t, each target scheduling action, and the immediate reward of each target scheduling action into an MDP data unit;

[0203] Aggregate multiple MDP data units to generate a sample data pool;

[0204] Extract a set number of MDP data units from the sample data pool to generate a sample data set.

[0205] In a possible implementation manner, the processing module is further configured to:

[0206] Randomly extract a set number of MDP data units from the sample data pool to generate a sample data set;

[0207] Or,

[0208] Determine the priority of each MDP data unit in the sample data pool;

[0209] According to the priority, extract a set number of MDP data units from the sample data pool to generate a sample data set.

[0210] In a possible implementation manner, the double Q network includes an online Q network and a target Q network. The online Q network and the target Q network have the same network architecture. The determination module is specifically configured to:

[0211] Input the operating state at time t-1 into the online Q network to determine the predicted action value corresponding to each target scheduling action;

[0212] Input the operating state at time t into the target Q network, and determine the output value of the target Q network according to the immediate reward. The output value is the target action value;

[0213] Calculate the mean square error loss between the predicted action value and the target action value, and update the parameters of the online Q network through an optimization algorithm to minimize the mean square error loss;

[0214] Based on a preset training end condition, obtain a pre-trained operation scheduling generation model according to the updated parameters of the online Q network.

[0215] In a possible implementation manner, the determination module 303 is specifically configured to:

[0216] Judge whether the update times of the parameters of the online Q network reach the times threshold;

[0217] When the update times of the parameters of the online Q network reach the times threshold, synchronize the parameters of the online Q network to the target Q network based on the update strategy;

[0218] Judge whether the online Q network is in a converged state;

[0219] When the online Q-network is in a converged state, use the current online Q-network as the pre-trained operation scheduling generation model;

[0220] When the online Q-network is not in a converged state, continue iterative training until the online Q-network is in a converged state.

[0221] In a possible implementation manner, the processing module is specifically configured to:

[0222] When the number of updates of the parameters of the online Q-network reaches the number threshold, copy the parameters of the online Q-network to the target Q-network;

[0223] Or,

[0224] When the number of updates of the parameters of the online Q-network reaches the number threshold, perform weighted summation on the parameters of the online Q-network and the parameters of the current target Q-network according to the set weight coefficient to obtain the parameters of the target Q-network.

[0225] The operation scheduling device of the combined energy storage power system provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0226] It should be noted that it should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the processing module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above processing module. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or can be independently implemented. Here, the processing element can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the hardware of the processor element or the instruction in the form of software.

[0227] For example, the above-mentioned modules can be one or more integrated circuits configured to implement the above methods, such as: one or more Application Specific Integrated Circuits (ASICs), or one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs), etc. For another example, when a certain above-mentioned module is implemented in the form of a processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. For another example, these modules can be integrated together and implemented in the form of a System-On-a-Chip (SOC).

[0228] Figure 4 Schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 4 shown, the electronic device 400 provided by the embodiment of the present application may include: a processor 401, and a memory 402 communicatively connected to the processor, where:

[0229] The memory stores computer-executable instructions;

[0230] The processor executes the computer-executable instructions stored in the memory to implement the method described in the foregoing method embodiments.

[0231] It should be understood that the processor 401 can be a Central Processing Unit (CPU), and can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly embodied as being executed and completed by a hardware processor, or can be executed and completed by a combination of hardware and software modules in the processor. The memory 402 may include a high-speed random access memory (Random Access Memory, RAM), and may also include a non-volatile memory NVM (non-volatile memory), such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0232] Optionally, the electronic device 400 may further include a communication interface 403. In specific implementation, if the communication interface 403, the memory 402, and the processor 401 are implemented independently, the communication interface 403, the memory 402, and the processor 401 may be connected to each other through a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus may be divided into an address bus, a data bus, a control bus, etc., but it does not mean that there is only one bus or one type of bus.

[0233] Optionally, in specific implementation, if the communication interface 403, the memory 402, and the processor 401 are integrated on a single chip, the communication interface 403, the memory 402, and the processor 401 may communicate through an internal interface.

[0234] The embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed, they are used to implement the method described in any of the foregoing embodiments.

[0235] It can be understood that the computer-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a Static Random Access Memory (SRAM), an Electrically Erasable Programmable Read Only Memory (EEPROM), an Erasable Programmable Read Only Memory (EPROM), a Programmable Read Only Memory (PROM), a Read Only Memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disc. The readable storage medium may be any available medium accessible by a general-purpose or special-purpose computer.

[0236] An exemplary computer-readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the computer-readable storage medium. Of course, the computer-readable storage medium can also be a component of the processor. The processor and the computer-readable storage medium can be located in an ASIC. Of course, the processor and the computer-readable storage medium can also exist as discrete components in an electronic device.

[0237] The integrated modules implemented in the form of software functional modules as described above can be stored in a computer-readable storage medium. The software functional modules stored in a computer-readable storage medium include several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present application.

[0238] The embodiments of the present application also provide a computer program product, including a computer program, which when executed implements the method described in any of the foregoing embodiments.

[0239] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0240] Furthermore, it should be noted that although the steps in the flowchart are displayed in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0241] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should all be considered as within the scope described in this specification.

[0242] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only illustrative, and the true scope and spirit of the present application are pointed out by the following claims.

[0243] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for operating and dispatching a combined energy storage power system, characterized in that: include: Obtaining the operating status of the combined energy storage power system; Based on the running state, selecting a candidate scheduling action that matches the running state from preset scheduling actions; Inputting the running state and the candidate scheduling action into a pre-trained running scheduling generation model to obtain a predicted action value of executing the candidate scheduling action under the running state, wherein the running scheduling generation model is an online Q network obtained by dual Q network training; Among the selected multiple candidate scheduling actions, the operation of the combined energy storage power system is scheduled according to the candidate scheduling action with the maximum predicted action value.

2. The method according to claim 1, characterized in that The step of inputting the running state and the candidate scheduling action into a pre-trained running scheduling generation model to obtain a predicted action value of executing the candidate scheduling action under the running state includes: Inputting the running state into the front-end network of the running scheduling generation model for feature extraction to obtain high-dimensional features of the motion state, wherein the front-end network includes a network unit, and the network unit is used to extract the high-dimensional features of the input running state; Inputting the high-dimensional features into the flattening layer of the operation scheduling generation model to convert features into vectors, and obtaining a one-dimensional feature vector corresponding to the high-dimensional features; Inputting the one-dimensional feature vector into the value function network of the operation scheduling generation model to calculate the state estimation value, so as to obtain the state estimation value corresponding to the operation state; Inputting the one-dimensional feature vector and the candidate scheduling action into the splicing layer of the operation scheduling generation model for feature splicing to obtain a joint feature vector; Inputting the joint feature vector into the advantage function network of the operation scheduling generation model to calculate the advantage value of the scheduling action, and obtaining the advantage value of executing the candidate scheduling action in the operation state; The state estimation value and the advantage value of the candidate scheduling action are input into the output layer of the operation scheduling generation model for summation calculation to obtain the predicted action value corresponding to executing the candidate scheduling action under the operation state.

3. The method according to claim 1 or 2, characterized in that: The pre-trained operation scheduling generation model is obtained in the following way: Acquire historical operation information of the combined energy storage power system, the historical operation information including the operation status of the combined energy storage power system at time t-1, the dispatching action executed under the operation status at time t-1, the instant reward obtained by executing the dispatching action, and the operation status at time t; Generate a sample data set including a Markov decision process MDP data unit according to the historical operation information, wherein the MDP data unit includes the operation status at time t-1, the scheduling action executed under the operation status at time t-1, the instant reward obtained by executing the scheduling action corresponding to the operation status at time t-1, and the operation status at time t; The operation scheduling generation model is trained according to the sample data set to obtain a pre-trained operation scheduling generation model.

4. The method according to claim 3, characterized in that The step of generating a sample data set including a Markov decision process MDP data unit according to the historical operation information includes: Constructing an objective function with the combined energy storage power system operation dispatch optimization as the goal and defining a set of constraints for the objective function, wherein the set of constraints includes electric field output constraints and energy storage power station output constraints; According to the constraint condition set, a target scheduling action that meets the constraint condition is selected from the executed scheduling actions; Based on the objective function, calculating the immediate reward of each target scheduling action, wherein the immediate reward is related to the optimization target of the objective function; Encode the running state at time t-1, the running state at time t, each target scheduling action, and the instant reward of each target scheduling action into an MDP data unit; Aggregate multiple MDP data units to generate a sample data pool; A set number of MDP data units are extracted from the sample data pool to generate a sample data set.

5. The method according to claim 4, characterized in that The step of extracting a set number of MDP data units from the sample data pool to generate a sample data set includes: Randomly extracting a set number of MDP data units from the sample data pool to generate a sample data set; or, Determining the priority of each MDP data unit in the sample data pool; According to the priority, a set number of MDP data units are extracted from the sample data pool to generate a sample data set.

6. The method according to claim 4, characterized in that The dual Q network includes an online Q network and a target Q network, the online Q network and the target Q network have the same network architecture, and the operation scheduling generation model is trained according to the sample data set to obtain a pre-trained operation scheduling generation model, including: Inputting the operating state at time t-1 into the online Q network to determine the predicted action value corresponding to each target scheduling action; Input the running state at time t into the target Q network, and determine the output value of the target Q network according to the instant reward, wherein the output value is the target action value; Calculating the mean square error loss between the predicted action value and the target action value, and updating the parameters of the online Q network through an optimization algorithm to minimize the mean square error loss; Based on the preset training end conditions, the pre-trained operation scheduling generation model is obtained according to the updated parameters of the online Q network.

7. The method according to claim 6, characterized in that The method of obtaining a pre-trained operation scheduling generation model based on the preset training end condition and the updated parameters of the online Q network includes: Determine whether the number of updates of the parameters of the online Q network reaches a number threshold; When the number of updates of the parameters of the online Q network reaches a number threshold, synchronizing the parameters of the online Q network to the target Q network based on an update strategy; Determining whether the online Q network is in a convergence state; When the online Q network is in a converged state, the current online Q network is used as a pre-trained operation scheduling generation model; When the online Q network is not in a converged state, iterative training is continued until the online Q network is in a converged state.

8. The method according to claim 7, characterized in that When the number of updates of the parameters of the online Q network reaches a number threshold, synchronizing the parameters of the online Q network to the target Q network based on an update strategy includes: When the number of updates of the parameters of the online Q network reaches a number threshold, copying the parameters of the online Q network to the target Q network; or, When the number of updates of the parameters of the online Q network reaches a threshold number, the parameters of the online Q network and the parameters of the current target Q network are weighted and summed according to a set weight coefficient to obtain the parameters of the target Q network.

9. An operation and dispatching device for a combined energy storage power system, characterized in that: include: An acquisition module, used to acquire the operating status of the combined energy storage power system; A selection module, configured to select, based on the running state, a candidate scheduling action that matches the running state from preset scheduling actions; A determination module, configured to input the running state and the candidate scheduling action into a pre-trained running scheduling generation model to obtain a predicted action value of executing the candidate scheduling action under the running state, wherein the running scheduling generation model is an online Q network obtained by dual Q network training; The scheduling module is used to schedule the operation of the combined energy storage power system according to the candidate scheduling action with the maximum predicted action value among the selected multiple candidate scheduling actions.

10. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 8 when executed by a processor.

12. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 8 when being executed by a processor.

Citation Information

Patent Citations

  • Region coordinated control system and region coordinated control method based on active power distribution network

    CN103151796A

  • Micro-power-grid energy storage scheduling method and device based on deep Q-value network (DQN) reinforcement learning

    CN109347149A

  • Decision optimization method for energy storage in transaction market based on double-Q learning algorithm

    CN110598925A

  • Post-disaster power distribution network dynamic first-aid repair method and system

    CN113627733A

  • Emergency generator tripping decision-making method based on knowledge fusion and deep reinforcement learning

    CN115566665A