Energy storage system multi-market joint bidding optimization method, device, equipment and medium
By optimizing the multi-market bidding strategy of energy storage systems using a deep differentiable reinforcement learning framework and a doubly differentiable reward function, the problems of low sample efficiency and insufficient understanding of market physical rules in traditional methods are solved, enabling the efficient operation of energy storage systems in highly volatile market environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-03-20
AI Technical Summary
In existing technologies, energy storage systems face problems such as low sample efficiency and slow convergence speed when participating in the joint operation of the energy market and ancillary services market. Furthermore, traditional reinforcement learning methods lack an understanding of the physical rules of the market and cannot effectively embed complex constraints, resulting in poor performance in highly volatile market environments.
A deep differentiable reinforcement learning framework is adopted to transform the bidding curve into a continuous and differentiable form through a neural network supply function, establish a multi-market sequential clearing mechanism, construct a doubly differentiable reward function, and combine it with the physical constraint penalty term of the energy storage system to optimize the multi-market joint bidding strategy of the energy storage system.
It significantly enhances the economic value and flexible adjustment capabilities of energy storage systems in high-proportion renewable energy power systems, improves training efficiency, ensures strict compliance with market constraints, and reduces the optimization complexity caused by market coupling.
Smart Images

Figure CN120823023B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power market and artificial intelligence, and particularly relates to a multi-market joint bidding optimization method and device for an energy storage system, equipment and a medium. BACKGROUND
[0002] With the large-scale grid connection of renewable energy in power systems, the fluctuation and uncertainty of renewable energy output bring serious challenges to power system operation. As a key flexible regulation resource, the energy storage system faces multiple technical problems when participating in the joint operation of energy market and ancillary service market.
[0003] The current clearing mechanism of the power market generally adopts a step bidding curve. This discrete market rule is essentially in conflict with the gradient optimization-based reinforcement learning method, resulting in the problems of low sample efficiency and slow convergence speed of the traditional model-independent reinforcement learning algorithm in the training process. At the same time, the energy storage system must strictly follow the state of charge constraint and the charge and discharge power limit and other physical operation boundaries, and the existing methods are difficult to effectively embed these complex constraints into the agent learning process. More importantly, there is a strong coupling relationship between the energy market and the ancillary service market such as frequency modulation and standby, and the interaction between markets further increases the optimization complexity.
[0004] The existing research mainly has three limitations: the optimization method based on mathematical programming is severely dependent on the accuracy of price prediction and performs poorly in a high volatility market environment; the traditional reinforcement learning lacks understanding of market physical rules and has low exploration efficiency; and the simple physical information reinforcement learning method fails to fundamentally solve the non-differentiable characteristics of the market mechanism. SUMMARY
[0005] The present application provides a multi-market joint bidding optimization method and device for an energy storage system, equipment and a medium to solve the problems in the related art that the optimization method based on mathematical programming is severely dependent on the accuracy of price prediction and performs poorly in a high volatility market environment; the traditional reinforcement learning lacks understanding of market physical rules and has low exploration efficiency; and the simple physical information reinforcement learning method fails to fundamentally solve the non-differentiable characteristics of the market mechanism.
[0006] The first aspect of the present application provides a method for optimizing multi-market joint bidding of an energy storage system, comprising the following steps: obtaining at least one operating parameter of the energy storage system and a power market; constructing a deep differentiable reinforcement learning framework based on the at least one operating parameter; converting a bidding curve into a continuous differentiable form by a neural network supply function based on the deep differentiable reinforcement learning framework, establishing a multi-market sequential clearing mechanism based on market priority, and constructing a double differentiable reward function containing a physical constraint penalty term, so as to construct a deep differentiable reinforcement learning framework adapted to a power multi-market scenario according to the continuous differentiable form, the multi-market sequential clearing mechanism and the double differentiable reward function; constructing a target function whose sum of energy storage multi-market revenues satisfies a preset maximization condition based on the deep differentiable reinforcement learning framework adapted to the power multi-market scenario; training a preset model using the target function to generate an energy storage system multi-market joint bidding optimization model, so as to output a joint bidding optimization result of the energy storage system and the power market based on the energy storage system multi-market joint bidding optimization model.
[0007] Optionally, in an embodiment of the present application, the at least one operating parameter of the energy storage system and the power market comprises: obtaining energy storage physical parameters including initial state of charge, charge and discharge power limit, charge and discharge efficiency and state of charge safety margin; obtaining multi-market price data including energy market historical price sequence, frequency regulation market historical price sequence and standby market historical price sequence; obtaining reinforcement learning training parameters including total number of training steps, learning rate, discount factor and state of charge violation penalty coefficient.
[0008] Optionally, in an embodiment of the present application, the deep differentiable reinforcement learning framework comprises: obtaining an initial time system state of the energy storage system, and determining an initial bidding action according to a policy network; calculating a single-step revenue of the energy storage system according to a reward function, the initial time system state and the initial bidding action; generating a next time state according to the initial time system state and the initial bidding action based on the single-step revenue and a state transition function; traversing a time step based on the next time state, and accumulating a discount reward in the time step; calculating a gradient by back propagation to iteratively optimize a bidding strategy based on the discount reward until the accumulated reward converges, thereby constructing the deep differentiable reinforcement learning framework.
[0009] Optionally, in an embodiment of the present application, the deep reinforcement learning framework for the adaptive power multi-market scenario is constructed according to the continuously differentiable form, the multi-market sequential clearing mechanism and the doubly differentiable reward function, comprising: converting the discrete bidding curve of the power market into the continuously differentiable form by using the neural network supply function to generate a mapping from the market state vector to the multi-market bidding power, obtaining the differentiable bidding power; performing a sequential clearing operation on the differentiable bidding power according to the transaction rules of the power market and the energy storage physical constraints to dynamically update the available power boundary and calibrate the actual clearing power, generating the calibrated actual clearing power; based on the calibrated actual clearing power, establishing a doubly differentiable reward function containing the physical constraint penalty term, and converting the physical constraint into a regulation factor of the doubly differentiable reward function, to generate a collaborative optimization result of maximizing multi-market revenue and satisfying physical constraints; based on the collaborative optimization result, constructing a time sequence evolution model of the power market and the energy storage system, and calculating the state of charge and market state at the next time according to the time sequence evolution model, to construct the deep reinforcement learning framework for the adaptive power multi-market scenario according to the state of charge and the market state at the next time.
[0010] Optionally, in an embodiment of the present application, the mapping formula from the market state vector to the multi-market bidding power is:
[0011]
[0012] wherein, is the multi-market bidding power, is the parameter of the neural network supply function, is the neural network with as the parameter, is the market state;
[0013] The calculation formula of the calibrated actual clearing power is:
[0014]
[0015] wherein, is the calibrated actual clearing power, is the energy market bidding power, is the discharge direction auxiliary service market bidding power, is the charge direction auxiliary service market bidding power, is the direction reversal amount of the energy market bidding power, is the maximum discharge power, is the maximum charge power;
[0016] The calculation formula of the state of charge at the next time is:
[0017]
[0018] in, The state of charge at the next moment. For time period State of charge of energy storage For time step, For energy storage charging efficiency, For energy storage and discharge efficiency, The actual output power in the charging direction. This represents the actual clearing power in the discharge direction.
[0019] Optionally, in one embodiment of the present invention, the formula for constructing the objective function is:
[0020]
[0021] in, The parameters of the supply function for the neural network, For the time period ,market The clearing price For the time period ,market The energy storage bidding power, For the time period of Violation penalties This is the discount factor.
[0022] A second aspect of the present invention provides a multi-market joint bidding optimization device for an energy storage system, comprising: an acquisition module for acquiring at least one operating parameter of the energy storage system and the electricity market; a first framework construction module for constructing a deep differentiable reinforcement learning framework based on the at least one operating parameter; a second framework construction module for, based on the deep differentiable reinforcement learning framework, transforming the bidding curve into a continuously differentiable form through a neural network supply function, establishing a multi-market sequential clearing mechanism based on market priority, and constructing a double differentiable reward function containing a physical constraint penalty term, so as to construct a deep differentiable reinforcement learning framework adapted to the multi-market electricity scenario according to the continuously differentiable form, the multi-market sequential clearing mechanism, and the double differentiable reward function; a function construction module for constructing an objective function that satisfies a preset maximization condition for the sum of revenue from multiple energy storage markets based on the deep differentiable reinforcement learning framework adapted to the multi-market electricity scenario; and an optimization module for training a preset model using the objective function to generate a multi-market joint bidding optimization model for the energy storage system, so as to output the joint bidding optimization result obtained by the energy storage system and the electricity market based on the multi-market joint bidding optimization model for the energy storage system.
[0023] Optionally, in an embodiment of the present application, the obtaining module comprises: a first obtaining unit, configured to obtain energy storage physical parameters including an initial value of state of charge, a limit value of charging and discharging power, a charging and discharging efficiency, and a state of charge safety margin; a second obtaining unit, configured to obtain multi-market price data including a historical price sequence of an energy market, a historical price sequence of a frequency regulation market, and a historical price sequence of a reserve market; and a third obtaining unit, configured to obtain reinforcement learning training parameters including a total number of training steps, a learning rate, a discount factor, and a state of charge violation penalty coefficient.
[0024] Optionally, in an embodiment of the present application, the first framework construction module comprises: a determination unit, configured to obtain an initial time point system state of the energy storage system, and determine an initial bidding action according to a policy network; a calculation unit, configured to calculate a single-step return of the energy storage system according to a reward function, the initial time point system state, and the initial bidding action; a generation unit, configured to generate a next time point state according to the initial time point system state and the initial bidding action based on the single-step return and a state transition function; a traversal unit, configured to traverse a time step based on the next time point state, and accumulate a discounted reward in the time step; and a first construction unit, configured to construct the deep-differentiable reinforcement learning framework by iteratively optimizing a bidding policy through back propagation to calculate a gradient based on the discounted reward until a cumulative reward converges.
[0025] Optionally, in an embodiment of the present application, the second framework construction module comprises: a transformation unit, configured to transform a discrete bidding curve of the electricity market into the continuous differentiable form by using the neural network supply function to generate a mapping from a market state vector to multi-market bidding power, to obtain a differentiable bidding power; an update unit, configured to perform a sequential clearing operation on the differentiable bidding power according to a transaction rule of the electricity market and an energy storage physical constraint, to dynamically update an available power boundary and calibrate an actual clearing power, to generate a calibrated actual clearing power; an establishment unit, configured to establish a double differentiable reward function including the physical constraint penalty term based on the calibrated actual clearing power, and transform the physical constraint into a regulation factor of the double differentiable reward function, to generate a collaborative optimization result of maximizing multi-market returns and satisfying the physical constraint; and a second construction unit, configured to construct a time sequence evolution model of the electricity market and the energy storage system based on the collaborative optimization result, and calculate a next time point state of charge and market state according to the time sequence evolution model, to construct the deep-differentiable reinforcement learning framework for the adaptive electricity multi-market scenario according to the next time point state of charge and the market state.
[0026] Optionally, in an embodiment of the present application, the mapping formula from the market state vector to the multi-market bidding power is:
[0027]
[0028] wherein, is the multi-market bidding power, is a parameter of a neural network supply function, is a neural network with parameters, is the market state;
[0029] The calculation formula of the calibrated actual clearing power is:
[0030]
[0031] wherein, is the calibrated actual clearing power, is the energy market bidding power, is the discharging direction ancillary service market bidding power, is the charging direction ancillary service market bidding power, is the direction reversal amount of the energy market bidding power, is the maximum discharging power, is the maximum charging power;
[0032] The calculation formula of the next time state of charge is:
[0033]
[0034] wherein, is the next time state of charge, is the time period state of charge of the energy storage, is the time step, is the charging efficiency of the energy storage, is the discharging efficiency of the energy storage, is the charging direction actual clearing power, is the discharging direction actual clearing power.
[0035] Optionally, in an embodiment of the present application, the construction formula of the objective function is:
[0036]
[0037] wherein, is a parameter of the neural network supply function, is the clearing price of the market in the time period, is the energy storage bidding power of the market in the time period, is the energy storage bidding power of the market in the time period, is the time period of a violation penalty term, is a discount factor.
[0038] The third aspect of the present application provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the energy storage system multi-market joint bidding optimization method as described in the above embodiments.
[0039] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the energy storage system multi-market joint bidding optimization method as described above.
[0040] The fifth aspect of the present application provides a computer program product, which stores a computer program, and the program is executed by a processor to implement the energy storage system multi-market joint bidding optimization method as described above.
[0041] The energy storage system multi-market joint bidding optimization method based on the derivable reinforcement learning of the embodiments of the present application effectively solves the three technical problems faced by the traditional energy storage bidding method through the innovative design of the derivable framework. First, the neural network supply function is used to realize the continuous derivable parameterization of the discrete bidding curve, which breaks through the bottleneck of the gradient optimization failure of the traditional reinforcement learning under the step market rule; second, the designed double penalty term reward function ingeniously converts the physical constraints of energy storage into a derivable optimization target, which guarantees the training efficiency while strictly meeting the operation safety requirements; finally, the multi-market sequential clearing mechanism significantly reduces the optimization complexity caused by market coupling through the priority dynamic adjustment strategy. Thus, the problems in the related art are solved, such as the optimization method based on mathematical programming which seriously depends on the accuracy of price prediction and performs poorly in a high volatility market environment; the traditional reinforcement learning lacks understanding of the market physical rules and has low exploration efficiency; and the simple physical information reinforcement learning method fails to fundamentally solve the non-derivable characteristics of the market mechanism.
[0042] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS
[0043] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, taken in conjunction with the accompanying drawings, in which:
[0044] Figure 1 is a flowchart of an energy storage system multi-market joint bidding optimization method according to an embodiment of the present application;
[0045] Figure 2 A schematic diagram of core logic and operation mechanism according to an embodiment of the present application;
[0046] Figure 3 A schematic diagram of four functional modules according to an embodiment of the present application;
[0047] Figure 4 A structural schematic diagram of a storage system multi-market joint bidding optimization device according to an embodiment of the present application;
[0048] Figure 5 A structural schematic diagram of an electronic device according to an embodiment of the present application.
[0049] Among them, 10 is a storage system multi-market joint bidding optimization device; 100 is an acquisition module, 200 is a first framework construction module, 300 is a second framework construction module, 400 is a function construction module, 500 is an optimization module; 501 is a memory, 502 is a processor, and 503 is a communication interface. DETAILED DESCRIPTION
[0050] The embodiments of the present application will be described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application.
[0051] The storage system multi-market joint bidding optimization method, device, equipment and medium of the embodiments of the present application are described below with reference to the accompanying drawings. In the related technologies mentioned in the above background art, the optimization method based on mathematical programming is seriously dependent on the accuracy of price prediction and performs poorly in a high volatility market environment; traditional reinforcement learning lacks understanding of market physical rules and has low exploration efficiency; the simple physical information reinforcement learning method fails to fundamentally solve the problem of non-differentiable characteristics of market mechanisms, and the present application provides a storage system multi-market joint bidding optimization method, in which an innovative solution based on differentiable reinforcement learning is proposed, a continuous differentiable parameterization of the bidding curve is realized through a neural network supply function, a reward function with boundary perception ability is designed, and a clearing mechanism considering market priority order is established, which significantly improves the revenue of the storage system in the electricity market and greatly improves the training efficiency, thereby providing reliable technical support for the participation of the storage in multi-market operation under the new power system. Thus, the problems in the related technologies, such as the optimization method based on mathematical programming being seriously dependent on the accuracy of price prediction and performing poorly in a high volatility market environment, traditional reinforcement learning lacking understanding of market physical rules and having low exploration efficiency, and the simple physical information reinforcement learning method failing to fundamentally solve the problem of non-differentiable characteristics of market mechanisms, are solved.
[0052] Specifically, Figure 1 A flowchart of a multi-market joint bidding optimization method for an energy storage system provided by an embodiment of the present application is shown.
[0053] As Figure 1 shown, the multi-market joint bidding optimization method for the energy storage system includes the following steps:
[0054] In step S101, at least one operating parameter of the energy storage system and the power market is obtained.
[0055] In actual execution, by obtaining the operating parameters of the energy storage system and the power market, the present application takes the maximization of multi-market revenue of the energy storage system as the optimization target from the perspective of the fusion of the power market mechanism and artificial intelligence, fully utilizes the gradient optimization characteristics of the guided reinforcement learning, and combines the dynamic response capability of the energy storage system to realize the collaborative optimization of the energy storage resource in the energy market, the frequency regulation market and the standby market, and significantly improves the economic value and flexible adjustment capability of the energy storage in the high-proportion renewable energy power system.
[0056] Optionally, in an embodiment of the present application, obtaining at least one operating parameter of the energy storage system and the power market includes: obtaining energy storage physical parameters including an initial value of state of charge, charge and discharge power limits, charge and discharge efficiency and state of charge safety margin; obtaining multi-market price data including an energy market historical price sequence, a frequency regulation market historical price sequence and a standby market historical price sequence; obtaining reinforcement learning training parameters including a total number of training steps, a learning rate, a discount factor and a state of charge violation penalty coefficient.
[0057] In actual execution, the operating parameters of the energy storage system and the power market in the embodiment of the present application specifically include: energy storage physical parameters: the initial value of state of charge is set to ; the charge and discharge power limits are determined to be , , the charge and discharge efficiency , ; the state of charge safety margin is set. Multi-market price data: the energy market historical price sequence , the frequency regulation market historical price sequence , the standby market historical price sequence are collected, and the time window is set to steps. Reinforcement learning training parameters: the total number of training steps is determined to be steps, the learning rate , the discount factor , and the state of charge violation penalty coefficient .
[0058] In step S102, a deep differentiable reinforcement learning framework is constructed based on at least one operating parameter.
[0059] Specifically, embodiments of the present invention can construct a deep differentiable reinforcement learning framework based on at least one operating parameter to achieve multi-market joint bidding optimization for energy storage systems. Its core logic and operating mechanism are as follows: Figure 2 As shown, the framework comprises four core components that work together to complete the closed loop of "state awareness - policy decision-making - reward feedback - state transition".
[0060] The blue box represents the policy network ( The energy storage bidding strategy is represented by a neural network, with parameters denoted as follows: The input is the current state. The output is the bidding action. Furthermore, the bidding action affects network parameters. Differentiable, supporting gradient optimization; the green diamond represents the reward function ( ), calculate single-step profit It integrates the benefits and penalties of bidding and clearing across multiple markets. The input is the current state. Bidding actions The output is a single-step reward signal, guiding the optimization direction of the policy network; the red diamond represents the state transition ( According to the current state Bidding actions Calculate the system state at the next moment. The framework rewards cumulative discounts. , The discount factor balances the weights of current and future returns. As the optimization objective, the objective function is calculated through backpropagation with respect to the policy network parameters. The gradient is then updated iteratively through gradient ascent. This achieves global optimization by maximizing multi-market returns and satisfying constraints.
[0061] Optionally, in one embodiment of the present invention, constructing a deep differentiable reinforcement learning framework includes: obtaining the initial system state of the energy storage system and determining the initial bidding action according to the policy network; calculating the single-step revenue of the energy storage system based on the reward function, the initial system state, and the initial bidding action; generating the next time-step state based on the single-step revenue and the state transition function, according to the initial system state and the initial bidding action; traversing time steps based on the next time-step state and accumulating the discounted reward in each time step; calculating the gradient through backpropagation based on the discounted reward to iteratively optimize the bidding strategy until the accumulated reward converges, thereby constructing the deep differentiable reinforcement learning framework.
[0062] In actual implementation, the operational flow in this embodiment of the invention includes:
[0063] (1) Initial state input: Obtain the system state at the initial moment. Input policy network ;
[0064] (2) Bidding action generation: strategy network Output initial bidding action ;
[0065] (3) Single-step reward calculation: reward function Combination and Calculate the single-step profit ;
[0066] (4) State transition update: State transition function in accordance with and Generate the state for the next time step. ;
[0067] (5) Timing Iteration Optimization: Repeat steps (2)-(4) above, traversing the time steps. Cumulative discount rewards ;
[0068] (6) Policy network update: The gradient is calculated through backpropagation, and the policy network parameters are updated along the gradient ascent direction. The bidding strategy is iteratively optimized until the cumulative reward converges.
[0069] In step S103, based on the deep differentiable reinforcement learning framework, the bidding curve is transformed into a continuously differentiable form through the neural network supply function, a multi-market sequential clearing mechanism based on market priority is established, and a dual differentiable reward function containing physical constraint penalty terms is constructed. In order to build a deep differentiable reinforcement learning framework adapted to the multi-market scenario of electricity based on the continuously differentiable form, the multi-market sequential clearing mechanism and the dual differentiable reward function.
[0070] In practical implementation, this invention innovatively constructs a differentiable multi-market operating environment, resolving the contradiction between the non-differentiability of market mechanisms and the requirements of reinforcement learning training in traditional energy storage bidding. Specifically, it employs a neural network supply function to provide a differentiable parameterized expression for the stepped bidding curve, designs a doubly differentiable reward function including Sigmoid reward decay and a quadratic penalty term to handle energy storage physical constraints, and establishes a multi-market sequential clearing mechanism based on market priority to replace the traditional joint clearing algorithm. Technically, the electricity market clearing rules and the dynamic characteristics of the energy storage system are first embedded into a differentiable computational graph. Then, the bidding strategy neural network is trained through end-to-end gradient optimization. Finally, the trained model is deployed to a real electricity market environment.
[0071] The present application breaks through the technical bottlenecks of traditional optimization methods relying on price prediction and conventional reinforcement learning with low exploration efficiency, starting from the deep integration of power market mechanism and artificial intelligence. According to the actual measurement data of the power market, the present application can increase the income of energy-frequency-reserve multi-market joint bidding of energy storage by more than 14%, improve the training efficiency by 4 times compared with the traditional reinforcement learning method, and ensure the strict satisfaction of energy storage operation constraints. The technology is particularly suitable for maximizing the economic value of energy storage systems in high-proportion renewable energy power systems, and provides an innovative solution for the commercial operation of energy storage in new power systems.
[0072] Optionally, in an embodiment of the present application, a deep differentiable reinforcement learning framework suitable for the power multi-market scenario is constructed according to the continuous differentiable form, multi-market sequential clearing mechanism and double differentiable reward function, comprising: converting the discrete bidding curve of the power market into a continuous differentiable form by using a neural network supply function to generate a mapping from the market state vector to the multi-market bidding power, obtaining the differentiable bidding power; according to the transaction rules of the power market and the physical constraints of the energy storage, performing a sequential clearing operation on the differentiable bidding power to dynamically update the available power boundary and calibrate the actual clearing power, generating the calibrated actual clearing power; based on the calibrated actual clearing power, a double differentiable reward function containing a physical constraint penalty term is established, and the physical constraint is converted into a regulation factor of the double differentiable reward function, to generate a collaborative optimization result of maximizing multi-market revenue and satisfying physical constraints; based on the collaborative optimization result, a time evolution model of the power market and the energy storage system is constructed, and the state of charge and the market state at the next time are calculated according to the time evolution model, to construct a deep differentiable reinforcement learning framework suitable for the power multi-market scenario according to the state of charge and the market state at the next time.
[0073] Specifically, the embodiment of the present application can construct a "state perception-strategy decision-reward feedback-state transition" closed-loop core framework based on the deep differentiable reinforcement learning framework, and realize the multi-market bidding optimization of the energy storage system through the cooperation of the four functional modules, as shown in the following figure. Figure 3
[0074] (1) Differentiable multi-market bidding curve generation module (DMO-B)
[0075] Break through the limitations of traditional discrete bidding mode in the power market, construct a continuous differentiable bidding strategy generation mechanism suitable for deep differentiable reinforcement learning, realize the gradient traceability of bidding power to strategy network parameters, and provide a differentiable strategy space for multi-market joint bidding optimization. Based on the neural network supply function, the discrete bidding curve is converted into a continuous differentiable form, realizing the mapping from the market state vector to the differentiable bidding power.
[0076] Implementation process:
[0077] Input layer
[0078] Market state vector: comprehensive state of electricity multi-market , including energy market historical price sequence , frequency regulation market historical price sequence , reserve market historical price sequence , and physical parameters such as energy storage state of charge, charge and discharge power limit Network initialization parameters: initial weights of policy network , using random initialization or pre-training method.
[0079] Operation layer
[0080] Using neural network supply function, a differentiable mapping from market state to multi-market bidding power is constructed, and the mathematical relationship is:
[0081]
[0082] where, is the parameter of the neural network supply function, is a neural network with as a parameter, and satisfies the differentiable condition: is the partial derivative of with respect to , which exists and is continuous, .
[0083] Output layer
[0084] Output continuous and differentiable multi-market bidding power , as the input of DMO-C module, which provides an initial strategy for "gradient optimization" for joint bidding of electricity multi-market, realizing the paradigm shift from "discrete bidding experience decision" to "continuous and differentiable intelligent optimization".
[0085] (2) Multi-market sequential clearing module (DMO-C)
[0086] Linking "differentiable theoretical bidding" and "actual clearing of electricity market", according to the trading rules of "energy priority, auxiliary coordination" of electricity market, the differentiable bidding power output by DMO-B is calibrated for actual clearing, ensuring the operability and compliance of bidding strategy, and providing real operation data for reward calculation and state transition. According to the priority order of "energy market → auxiliary service market", the bidding power is cleared, and the available power boundary is dynamically updated to calibrate the actual clearing power.
[0087] Implementation process:
[0088] Input layer
[0089] Theoretical bidding power: multi-market steerable bidding power output by the DMO-B module , covering bidding power for energy market, frequency regulation market, reserve market, etc. Market and device constraints: energy market clearing priority rules, auxiliary service market coordination rules, and storage charging and discharging power limits , state of charge safety margin .
[0090] Operation layer
[0091] According to the power market clearing priority and device constraints, the market bidding power is sequentially settled and calibrated, and the mathematical relationship is:
[0092]
[0093] Among them, is the energy market bidding power, is the discharging direction auxiliary service market bidding power, is the charging direction auxiliary service market bidding power, is the calibrated actual clearing power, is the direction reversal amount of the energy market bidding power, is the maximum discharging power, is the maximum charging power.
[0094] Output layer
[0095] Output the actual clearing power that meets the market rules and device constraints , which is the basis for DMO-R revenue calculation and DMO-T state evolution input, providing real and effective market feedback data for subsequent modules.
[0096] (3) Punishment and reward function module (DMO-R)
[0097] Establish a "revenue optimization + constraint satisfaction" dual-oriented reward mechanism to solve the "physical constraints and revenue target split" problem in traditional reinforcement learning. Convert physical constraints such as storage state of charge overrun into regulatory factors of the reward function to guide the policy network to actively avoid risks in multi-market bidding decisions, ensuring safe operation of the storage system while maximizing revenue. Construct a reward function containing a physical constraint penalty term to achieve the coordinated optimization of multi-market revenue maximization and physical constraint satisfaction.
[0098] Implementation process:
[0099] Input layer
[0100] Market side data: multi-market clearing price , actual clearing power calibrated by DMO-C Device-side data: real-time state of charge, safety margin of energy storage , penalty coefficient , maximum discharge power , time step .
[0101] Operation layer
[0102] The penalty term is calculated based on the SOC state segmentation, realizing the dynamic regulation and control of "no penalty in the safe interval, gradient penalty in the out-of-limit interval":
[0103]
[0104] wherein, : penalty term , encouraging the strategy network to maintain the state; , Triggering the quadratic function penalty term, the deeper the out-of-limit degree, the greater the penalty strength, forcing the strategy network to learn "constraint avoidance behavior", wherein, represents the actual state of charge of the battery at time t, represents the minimum state of charge, represents the maximum state of charge, represents the state of charge safety margin.
[0105] Output layer
[0106] Embedding the penalty term into the single-step revenue function, constructing an integrated feedback signal of "multi-market revenue-constraint penalty":
[0107]
[0108] wherein, is the multi-market clearing price, is the calibrated actual clearing power, is the time period of the violation penalty term, is the time period, is the market.
[0109] Through this feedback signal, the strategy network needs to consider both "maximizing multi-market revenue" and "minimizing constraint penalty" during the optimization process, realizing the coordinated optimization of power market revenue target and safe operation of energy storage devices, and providing a guiding reward signal for the strategy iteration of the framework.
[0110] (4) State transition function module (DMO-T)
[0111] The time sequence evolution model of the power multi-market and the energy storage system is constructed, and the state of charge and the market state at the next moment are calculated according to the current clearing power and the state of charge, so as to provide the state information required for time sequence decision-making for the DDRL framework.
[0112] Implementation process:
[0113] Input layer
[0114] Current system state: comprehensive state of power multi-market ; Actual clearing power: multi-market actual clearing power output by the DMO-C module ; Evolution model parameters: energy storage charging and discharging efficiency , Power market price evolution coefficient.
[0115] Operation layer
[0116] Based on the time sequence evolution of the power market price and the energy storage state of charge evolution model, the system state at the next moment is calculated , and the mathematical relationship is:
[0117] The calculation formula of the state of charge at the next moment is:
[0118]
[0119] Among them, is the state of charge at the next moment, is the time period of the state of charge of the energy storage, is the time step, is the charging efficiency of the energy storage, is the discharging efficiency of the energy storage.
[0120] The evolution formula of the market price at the next moment is:
[0121]
[0122] Among them, is the power clearing price of market M at the future t+1 moment, is the clearing price of market , is the price evolution coefficient of market M, is the discount factor, is a random disturbance term. Output layer
[0123] Output the system state at the next moment,
[0124] including the updated state of charge of the energy storage and the predicted multi-market price.
[0125] The state is taken as an input of the next iteration cycle under the deep derivable reinforcement learning framework, drives the strategy network to continuously optimize, and realizes the closed-loop iteration of ''historical data-current decision-future state''.
[0126] The deep derivable reinforcement learning closed loop is constructed through the four modules of DMO-B, DMO-C, DMO-R and DMO-T. Firstly, the neural network supply function is used to convert the multi-dimensional state of the electricity market into continuous derivable bidding power, breaking through the optimization limitation of discrete bidding. Secondly, the bidding power is sequentially cleared according to the rules of ''energy priority, multi-market coordination'' and the physical constraints of energy storage, realizing the transformation from theory to operation. Thirdly, the segmented penalty function is used to embed the overcharge constraint into the reward calculation, driving the strategy to balance the income and safety. Then, the market price time sequence rule and the energy storage physical evolution model are fused to deduce the system state, providing dynamic input for strategy iteration. On this basis, the whole process iterative training mechanism is used to drive the dynamic state time sequence data, and the strategy network parameters are optimized by back propagation. The strategy learns the correlation of multi-time step decision and the adaptability of market constraints, and gradually converges to the optimal bidding strategy of ''income-safety''.
[0127] The embodiment of the present application effectively solves the three technical problems faced by the traditional energy storage bidding method through the innovative derivable framework design. Firstly, the method uses a neural network supply function to realize the continuous derivable parameterization of the discrete bidding curve, breaking through the bottleneck of gradient optimization failure of traditional reinforcement learning under the step market rule. Secondly, the designed double penalty item reward function skillfully converts the physical constraints of energy storage into a derivable optimization target, ensuring training efficiency while strictly meeting the operation safety requirements. Finally, the multi-market sequential clearing mechanism significantly reduces the optimization complexity caused by market coupling through dynamic adjustment of the priority strategy.
[0128] In an embodiment of the present application, the mapping formula from the market state vector to the multi-market bidding power is:
[0129]
[0130] Among them, is the multi-market bidding power, is the neural network with as the parameter, is the market state;
[0131] The calculation formula of the calibrated actual clearing power is:
[0132]
[0133] Among them, is the calibrated actual clearing power, is the energy market bidding power, is the discharge direction auxiliary service market bidding power, a direction of charging auxiliary service market bidding power, a direction of energy market bidding power, a maximum discharging power, a maximum charging power;
[0134] The calculation formula of the state of charge at the next moment is:
[0135]
[0136] wherein, is the state of charge at the next moment, is a time period a state of charge of the energy storage, is a time step, is a charging efficiency of the energy storage, is a discharging efficiency of the energy storage, is an actual clearing power in the charging direction, is an actual clearing power in the discharging direction.
[0137] a clearing price, is a price evolution coefficient of the market M, is a discount factor, is a random disturbance term.
[0138] The embodiment of the present application can improve the accuracy of calculation according to the formula, and further provides reliable technical support for the energy storage participating in multi-market operation under the new power system.
[0139] In step S104, based on the deep derivable reinforcement learning framework adapted to the power multi-market scene, a target function is constructed, wherein a sum of multi-market benefits of the energy storage satisfies a preset maximum condition.
[0140] In the actual execution process, the embodiment of the present application can construct a target function for maximizing the sum of multi-market benefits of the energy storage based on the deep derivable reinforcement learning framework adapted to the power multi-market scene, thereby providing support for subsequent multi-market joint bidding optimization of the energy storage system.
[0141] wherein, in an embodiment of the present application, the construction formula of the target function is:
[0142]
[0143] wherein, is a parameter of a neural network supply function, used for constructing a fitting bidding strategy; is a multi-market set, is an energy market, is a discharging direction auxiliary service market, is a charging direction auxiliary service market; for the time period , market clearing price; for the time period , market storage bidding power, energy market net output , auxiliary market bidding power , ; for the time period penalty term for violation of rules; is a discount factor, reflecting the global optimization of the "current income + future income".
[0144] In step S105, the preset model is trained using the objective function to generate a storage system multi-market joint bidding optimization model, so as to output a joint bidding optimization result of the storage system and the power market based on the storage system multi-market joint bidding optimization model.
[0145] It can be understood that the storage system multi-market joint bidding optimization model in the embodiment of the application can be a trained model.
[0146] Specifically, the embodiment of the application can train the preset model using the objective function to generate a storage system multi-market joint bidding optimization model, so as to output a joint bidding optimization result of the storage system and the power market based on the storage system multi-market joint bidding optimization model, that is, to apply the trained model data to the real power market: first, the market price sequence and the storage state of charge collected in real time are forward propagated through the neural network supply function to generate a continuous supply curve covering the entire price interval, and the power-price mapping relationship is constructed by dense sampling; then the supply curve is subjected to monotonicity processing, and is discretized based on the iterative greedy algorithm - taking the square deviation integral as the error function, the price-power anchor points are updated through iteration to generate N groups of price-power pairs that meet the market rules, and a standardized high-dimensional bidding file is formed.
[0147] The embodiment of the application solves the joint bidding optimization problem of the storage system in the energy market, the frequency regulation market and the standby market by fusing the physical rules of the power market and the guided reinforcement learning technology, and is particularly suitable for flexible scheduling and economic value maximization of storage resources in a high-proportion renewable energy power system.
[0148] The multi-market joint bidding optimization method for energy storage systems proposed in this invention, based on differentiable reinforcement learning, effectively solves three major technical challenges faced by traditional energy storage bidding methods through an innovative differentiable framework design. First, this method uses a neural network supply function to achieve continuous differentiable parameterization of the discrete bidding curve, overcoming the bottleneck of gradient optimization failure in traditional reinforcement learning under tiered market rules. Second, the designed double-penalty reward function cleverly transforms the physical constraints of energy storage into a differentiable optimization objective, ensuring training efficiency while strictly meeting operational safety requirements. Finally, the established multi-market sequential clearing mechanism significantly reduces the optimization complexity caused by market coupling through a priority dynamic adjustment strategy. Thus, it addresses the problems in related technologies: optimization methods based on mathematical programming heavily rely on price prediction accuracy and perform poorly in highly volatile market environments; traditional reinforcement learning lacks understanding of market physical rules and has low exploration efficiency; and simple physical information reinforcement learning methods fail to fundamentally solve the problem of the non-differentiable nature of market mechanisms.
[0149] Next, referring to the accompanying drawings, we describe the multi-market joint bidding optimization device for energy storage systems proposed according to an embodiment of the present invention.
[0150] Figure 4 This is a schematic diagram of the structure of the multi-market joint bidding optimization device for the energy storage system according to an embodiment of the present invention.
[0151] like Figure 4 As shown, the multi-market joint bidding optimization device 10 for the energy storage system includes: an acquisition module 100, a first framework construction module 200, a second framework construction module 300, a function construction module 400, and an optimization module 500.
[0152] Specifically, the acquisition module 100 is used to acquire at least one operating parameter of the energy storage system and the electricity market.
[0153] The first framework building module 200 is used to build a deep differentiable reinforcement learning framework that includes state awareness, policy network, reward function and state transition function based on at least one running parameter.
[0154] The second framework construction module 300 is used to transform the bidding curve into a continuously differentiable form through a neural network supply function based on a deep differentiable reinforcement learning framework, establish a multi-market sequential clearing mechanism based on market priority, and construct a double differentiable reward function containing physical constraint penalty terms. This allows for the construction of a deep differentiable reinforcement learning framework adapted to the multi-market scenario of electricity based on the continuously differentiable form, the multi-market sequential clearing mechanism, and the double differentiable reward function.
[0155] The function construction module 400 is configured to construct a target function of which a total sum of energy storage multi-market benefits meets a preset maximization condition based on the deep differentiable reinforcement learning framework suitable for the power multi-market scenario.
[0156] The optimization module 500 is configured to train a preset model by using the target function to generate an energy storage system multi-market joint bidding optimization model, so as to output a joint bidding optimization result of the energy storage system and the power market based on the energy storage system multi-market joint bidding optimization model.
[0157] Optionally, in an embodiment of the present application, the acquisition module 100 comprises a first acquisition unit, a second acquisition unit and a third acquisition unit.
[0158] The first acquisition unit is configured to acquire energy storage physical parameters including an initial value of a state of charge, a limit value of charging and discharging power, a charging and discharging efficiency and a state of charge safety margin.
[0159] The second acquisition unit is configured to acquire multi-market price data including a historical price sequence of an energy market, a historical price sequence of a frequency modulation market and a historical price sequence of a standby market.
[0160] The third acquisition unit is configured to acquire reinforcement learning training parameters including a total number of training steps, a learning rate, a discount factor and a state of charge violation penalty coefficient.
[0161] Optionally, in an embodiment of the present application, the first framework construction module 200 comprises a determination unit, a calculation unit, a generation unit, a traversal unit and a first construction unit.
[0162] The determination unit is configured to acquire an initial time system state of the energy storage system, and determine an initial bidding action according to a policy network.
[0163] The calculation unit is configured to calculate a single-step benefit of the energy storage system according to a reward function, the initial time system state and the initial bidding action.
[0164] The generation unit is configured to generate a next time state according to the initial time system state and the initial bidding action based on the single-step benefit and a state transition function.
[0165] The traversal unit is configured to traverse a time step based on the next time state, and accumulate a discounted reward in the time step.
[0166] The first construction unit is configured to construct the deep differentiable reinforcement learning framework by iteratively optimizing the bidding policy through back propagation to calculate a gradient based on the discounted reward until the accumulated reward converges.
[0167] Optionally, in an embodiment of the present application, the second framework construction module 300 comprises a conversion unit, an update unit, a building unit and a second construction unit.
[0168] The transformation unit is configured to transform discrete bidding curves of the electricity market into a continuously derivable form by using a neural network supply function to generate a mapping from a market state vector to multi-market bidding power, and obtain derivable bidding power.
[0169] The updating unit is configured to perform a sequential clearing operation on the derivable bidding power according to a transaction rule of the electricity market and an energy storage physical constraint, to dynamically update an available power boundary and calibrate an actual clearing power, and generate a calibrated actual clearing power.
[0170] The establishing unit is configured to establish a double derivable reward function containing a physical constraint penalty term based on the calibrated actual clearing power, and convert the physical constraint into a regulation factor of the double derivable reward function, to generate a collaborative optimization result of maximizing multi-market revenue and satisfying the physical constraint.
[0171] The second constructing unit is configured to construct a time sequence evolution model of the electricity market and the energy storage system based on the collaborative optimization result, and calculate a next time state of charge and market state according to the time sequence evolution model, to construct a deep derivable reinforcement learning framework suitable for the power multi-market scenario according to the next time state of charge and market state.
[0172] Optionally, in an embodiment of the present application, the mapping formula from the market state vector to the multi-market bidding power is:
[0173]
[0174] wherein, is the multi-market bidding power, is a neural network with as a parameter, is the market state;
[0175] The calculation formula of the calibrated actual clearing power is:
[0176]
[0177] wherein, is the calibrated actual clearing power, is the energy market bidding power, is the discharge direction auxiliary service market bidding power, is the charge direction auxiliary service market bidding power, is the direction reversal amount of the energy market bidding power, is the maximum discharge power, is the maximum charge power;
[0178] The calculation formula of the next time state of charge is:
[0179]
[0180] in, The state of charge at the next moment. For time period State of charge of energy storage For time step, For energy storage charging efficiency, For energy storage and discharge efficiency, The actual output power in the charging direction. This represents the actual clearing power in the discharge direction.
[0181] Optionally, in one embodiment of the present invention, the formula for constructing the objective function is:
[0182]
[0183] in, To provide parameters for the supply function of the neural network, For time period ,market The clearing price For time period ,market The energy storage bidding power, For time period of Violation penalties This is the discount factor.
[0184] It should be noted that the foregoing explanation of the multi-market joint bidding optimization method for energy storage systems also applies to the multi-market joint bidding optimization device for energy storage systems in this embodiment, and will not be repeated here.
[0185] The multi-market joint bidding optimization device for energy storage systems proposed in this invention innovatively proposes a solution based on differentiable reinforcement learning. It achieves continuous differentiable parameterization of the bidding curve through a neural network supply function, designs a reward function with boundary awareness, and establishes a clearing mechanism that considers market priority. Practical application shows that this method significantly improves the profitability of energy storage systems in the Australian electricity market, while also greatly improving training efficiency, providing reliable technical support for energy storage participation in multi-market operations under new power systems. This addresses the problems in related technologies, such as the heavy reliance on price prediction accuracy in mathematical programming optimization methods, which perform poorly in highly volatile market environments; the lack of understanding of market physical rules in traditional reinforcement learning, resulting in low exploration efficiency; and the failure of simple physical information reinforcement learning methods to fundamentally solve the problem of the non-differentiable nature of market mechanisms.
[0186] Figure 5A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. The electronic device can include:
[0187] The memory 501, the processor 502, and a computer program stored in the memory 501 and executable on the processor 502.
[0188] The processor 502 implements the energy storage system multi-market joint bidding optimization method provided by the above embodiment when executing the program.
[0189] Further, the electronic device further includes:
[0190] The communication interface 503 is used for communication between the memory 501 and the processor 502.
[0191] The memory 501 is used for storing a computer program executable on the processor 502.
[0192] The memory 501 can include a high-speed RAM memory, and can also include a non-volatile memory, for example, at least one disk memory.
[0193] If the memory 501, the processor 502, and the communication interface 503 are independently implemented, the communication interface 503, the memory 501, and the processor 502 can be connected to each other through a bus and complete communication between each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0194] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a chip, the memory 501, the processor 502, and the communication interface 503 can complete communication between each other through an internal interface.
[0195] The processor 502 can be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0196] The embodiment also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the energy storage system multi-market joint bidding optimization method.
[0197] The embodiment also provides a computer program product, which stores a computer program, and the computer program is executed by a processor to implement the energy storage system multi-market joint bidding optimization method.
[0198] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or N embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples without contradiction.
[0199] In addition, the terms "first", "second" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, for example, two, three, etc., unless otherwise specifically limited.
[0200] Any process or method descriptions in flow charts or otherwise described herein can be understood as representing code modules, segments, or portions of code that include one or more executable instructions for performing a step, a function of the custom logic or process, and that the scope of the preferred embodiments of the present application encompasses additional implementation in which the functions are performed in a different order, in substantially simultaneous fashion, or as part of a concurrent process, as will be understood by those skilled in the art.
[0201] The logic and / or steps represented in flow diagrams or otherwise described herein, for example, can be considered as a sequence of executable instructions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be a product of the manufacturing and / or processing. The computer-readable medium can include, but is not limited to, the following: an electronic connection (an electronic device with one or N wires), a portable computer diskette (a magnetic device), a RAM (random access memory), a ROM (read-only memory), an EPROM (erasable programmable ROM) or a Flash memory, an optical fiber, and a portable CD ROM. In addition, the computer-readable medium can even be paper or another suitable medium upon which the program is printed, since the program can be electronically captured, via the optically scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and stored in the computer memory.
[0202] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. As such, if implemented in hardware and in another embodiment, any of the following technologies, known in the art, or their combinations can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.
[0203] Those of ordinary skill in the art can understand that all or part of the steps carried out by the above-mentioned embodiments can be completed by programs instructing relevant hardware, and the programs can be stored in a computer-readable storage medium. When the programs are executed, one or a combination of the steps of the method embodiments is included.
[0204] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0205] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A multi-market joint optimization method for energy storage systems, characterized in that, Includes the following steps: Acquire at least one operating parameter of the energy storage system and the electricity market, wherein acquiring at least one operating parameter of the energy storage system and the electricity market includes: acquiring energy storage physical parameters including initial state of charge (SOC), charging and discharging power limits, charging and discharging efficiency, and SOC safety margin; acquiring multi-market price data including historical price series of the energy market, the frequency regulation market, and the reserve market; and acquiring reinforcement learning training parameters including total training steps, learning rate, discount factor, and SOC violation penalty coefficient. Based on the at least one operating parameter, a deep differentiable reinforcement learning framework is constructed. Based on the aforementioned deep differentiable reinforcement learning framework, discrete curves are transformed into continuously differentiable forms through a neural network supply function. A multi-market sequential clearing mechanism based on market priority is established, and a doubly differentiable reward function containing physical constraint penalties is constructed. This allows for the construction of a deep differentiable reinforcement learning framework adapted to the multi-market electricity scenario based on the continuously differentiable form, the multi-market sequential clearing mechanism, and the doubly differentiable reward function. Specifically, the construction of the deep differentiable reinforcement learning framework adapted to the multi-market electricity scenario based on the continuously differentiable form, the multi-market sequential clearing mechanism, and the doubly differentiable reward function includes: using the neural network supply function to transform the discrete curves of the electricity market into the continuously differentiable form to generate a mapping from the market state vector to multi-market power, thereby obtaining differentiable power. Based on the trading rules of the electricity market and the physical constraints of energy storage, a sequential liquidation operation is performed on the differentiable power to dynamically update the available power boundary and calibrate the actual cleared power, generating the calibrated actual cleared power. Based on the calibrated actual cleared power, a doubly differentiable reward function containing the physical constraint penalty term is established, and the physical constraint is transformed into a control factor of the doubly differentiable reward function to generate a collaborative optimization result that maximizes multi-market revenue and satisfies physical constraints. Based on the collaborative optimization result, a time-series evolution model of the electricity market and the energy storage system is constructed, and the state of charge and market state at the next time step are calculated according to the time-series evolution model, so as to construct the deep differentiable reinforcement learning framework adapted to the multi-market electricity scenario based on the state of charge and market state at the next time step. Based on the deep differentiable reinforcement learning framework adapted to the multi-market electricity scenario, an objective function is constructed to satisfy the preset maximization condition of the total revenue of energy storage in multiple markets. The objective function is used to train a preset model to generate a multi-market joint optimization model for the energy storage system. Based on the multi-market joint optimization model, the joint optimization results of the energy storage system and the electricity market are output. Based on the joint optimization results, the energy storage resources applicable to high-proportion renewable energy power systems are flexibly dispatched. The method of outputting the joint optimization results obtained from the energy storage system and the electricity market also includes: The physical constraint penalty term is calculated based on the SOC state segmentation, enabling dynamic adjustment of no penalty in the safe interval and gradient penalty in the out-of-limit interval: in, Penalty items , , Triggering a quadratic function penalty term increases the severity of the penalty as the constraint is exceeded, forcing the policy network to learn constraint-avoidance behavior. This represents the actual state of charge of the battery at time t. Indicates the lowest state of charge. Indicates the maximum state of charge. This indicates the safety margin of the stated state of charge; The formula for calculating the state of charge at the next moment is: in, The state of charge at the next moment. The state of charge of the energy stored in time period T. For time step, For energy storage charging efficiency, For energy storage and discharge efficiency, The actual output power in the charging direction. This represents the actual clearing power in the discharge direction.
2. The multi-market joint optimization method for energy storage systems according to claim 1, characterized in that, The construction of the deep differentiable reinforcement learning framework includes: Obtain the initial system state of the energy storage system and determine the initial action based on the policy network; The single-step revenue of the energy storage system is calculated based on the reward function, the initial system state, and the initial action. Based on the single-step reward and state transition function, the state at the next moment is generated according to the system state at the initial moment and the initial action; Based on the state at the next moment, traverse the time steps and accumulate the discount reward in the time steps; Based on the aforementioned discount reward, the gradient is calculated through backpropagation to iteratively optimize the strategy until the cumulative reward converges, thus constructing the deep differentiable reinforcement learning framework.
3. The multi-market joint optimization method for energy storage systems according to claim 1, characterized in that, The mapping formula from the market state vector to multi-market power is as follows: in, For the aforementioned multi-market power, To provide parameters for the supply function of the neural network, For Neural networks with parameters The market state is as described; The formula for calculating the actual clearing power after calibration is as follows: in, The actual purging power after calibration. For energy market power, For the power of the discharge direction auxiliary service market, Power for the charging side traction service market, The direction reversal amount of the energy market power. For maximum discharge power, This is the maximum charging power; The formula for calculating the state of charge at the next moment is: in, The state of charge at the next moment. The state of charge of the energy stored in time period T. For time step, For energy storage charging efficiency, For energy storage and discharge efficiency, The actual output power in the charging direction. This represents the actual clearing power in the discharge direction.
4. The multi-market joint optimization method for energy storage systems according to claim 3, characterized in that, The formula for constructing the objective function is as follows: in, The parameters of the supply function for the neural network, For the time period ,market The clearing price For the time period ,market energy storage capacity, For the time period of Violation penalties This is the discount factor.
5. A multi-market joint optimization device for an energy storage system, characterized in that, include: An acquisition module is used to acquire at least one operating parameter of the energy storage system and the electricity market. The acquisition module includes: a first acquisition unit for acquiring energy storage physical parameters including initial state of charge (SOC), charging / discharging power limits, charging / discharging efficiency, and SOC safety margin; a second acquisition unit for acquiring multi-market price data including historical price series from the energy market, frequency regulation market, and reserve market; and a third acquisition unit for acquiring reinforcement learning training parameters including total training steps, learning rate, discount factor, and SOC violation penalty coefficient. The first framework construction module is used to construct a deep differentiable reinforcement learning framework based on the at least one running parameter. The second framework construction module is used to, based on the deep differentiable reinforcement learning framework, transform discrete curves into continuously differentiable forms through a neural network supply function, establish a multi-market sequential clearing mechanism based on market priority, and construct a doubly differentiable reward function containing physical constraint penalties. This allows for the construction of a deep differentiable reinforcement learning framework adapted to multi-market electricity scenarios based on the continuously differentiable form, the multi-market sequential clearing mechanism, and the doubly differentiable reward function. Specifically, constructing the deep differentiable reinforcement learning framework adapted to multi-market electricity scenarios based on the continuously differentiable form, the multi-market sequential clearing mechanism, and the doubly differentiable reward function includes: using the neural network supply function to transform the discrete curves of the electricity market into the continuously differentiable form to generate a mapping from market state vectors to multi-market power. The process involves: 1) Reaching the differentiable power; 2) Performing a sequential clearing operation on the differentiable power according to the trading rules of the electricity market and the physical constraints of energy storage, dynamically updating the available power boundary and calibrating the actual cleared power to generate the calibrated actual cleared power; 3) Establishing a doubly differentiable reward function containing the physical constraint penalty term based on the calibrated actual cleared power, and transforming the physical constraints into a control factor of the doubly differentiable reward function to generate a collaborative optimization result that maximizes multi-market returns and satisfies physical constraints; 4) Constructing a time-series evolution model of the electricity market and the energy storage system based on the time-series evolution model, and calculating the state of charge and market state at the next time step, to construct a deep differentiable reinforcement learning framework adapted to the multi-market electricity scenario based on the state of charge and market state at the next time step; The function construction module is used to construct an objective function that satisfies the preset maximization condition for the sum of energy storage multi-market revenue based on the deep differentiable reinforcement learning framework adapted to the multi-market electricity scenario. The optimization module is used to train a preset model using the objective function, generate a multi-market joint optimization model for the energy storage system, output the joint optimization results of the energy storage system and the electricity market based on the multi-market joint optimization model for the energy storage system, and flexibly schedule energy storage resources applicable to high-proportion renewable energy power systems based on the joint optimization results. The method of outputting the joint optimization results obtained from the energy storage system and the electricity market also includes: The physical constraint penalty term is calculated based on the SOC state segmentation, enabling dynamic adjustment of no penalty in the safe interval and gradient penalty in the out-of-limit interval: in, Penalty items , , Triggering a quadratic function penalty term increases the severity of the penalty as the constraint is exceeded, forcing the policy network to learn constraint-avoidance behavior. This represents the actual state of charge of the battery at time t. Indicates the lowest state of charge. Indicates the maximum state of charge. This indicates the safety margin of the stated state of charge; The formula for calculating the state of charge at the next moment is: in, The state of charge at the next moment. The state of charge of the energy stored in time period T. For time step, For energy storage charging efficiency, For energy storage and discharge efficiency, The actual output power in the charging direction. This represents the actual clearing power in the discharge direction.
6. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the multi-market joint optimization method for energy storage systems as described in any one of claims 1-4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the multi-market joint optimization method for energy storage systems as described in any one of claims 1-4.
8. A computer program product, comprising a computer program, characterized in that, The computer program is executed to implement the multi-market joint optimization method for energy storage systems as described in any one of claims 1-4.
Citation Information
Patent Citations
Clearing method and system for electric power peak shaving market
CN112580850A
Multi-microgrid electric energy transaction pricing strategy and system based on reinforcement and imitation learning
CN113706197A