Energy storage operation optimization method and system based on hierarchical multi-agent reinforcement learning

By constructing a hierarchical multi-agent reinforcement learning decision-making architecture, an upper-level planning agent and a lower-level execution agent are built, which solves the problem of balancing short-term benefits and long-term lifespan loss in energy storage systems, and realizes efficient energy storage system optimization and lifespan management.

CN120952277BActive Publication Date: 2026-02-27BEIJING DINGCHENG HONGAN TECH DEV CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511483406.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-02-27
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing optimization methods for energy storage systems fail to effectively balance short-term economic benefits with long-term lifespan losses, leading to premature equipment degradation, reduced economic efficiency throughout the entire life cycle, and a single decision-making level that results in high computational complexity and difficulty in ensuring real-time performance.

Method used

A decision-making architecture based on hierarchical multi-agent reinforcement learning is adopted to construct an upper-level planning agent and a lower-level execution agent. A monetization cost quantification model and reward function for battery health status are designed to achieve hierarchical decision optimization of the energy storage system.

Benefits of technology

By coordinating optimization across multiple time scales through a hierarchical decision-making architecture, computational complexity is reduced, solution efficiency and real-time performance are improved, a balance is achieved between short-term gains and long-term lifespan losses, a unified economic evaluation standard is provided, and the scalability and modularity of the system are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952277B_ABST
    Figure CN120952277B_ABST
Patent Text Reader

Abstract

Energy storage operation optimization method and system based on hierarchical multi-agent reinforcement learning. First, initialize the hierarchical multi-agent reinforcement learning decision framework, and define the interlayer communication synchronization mechanism. Then, design an upper-layer reward function balancing electricity bill income and battery life loss cost to optimize the upper-layer planning strategy; design a lower-layer reward function integrating current smoothness and budget tracking accuracy to optimize the lower-layer execution strategy; finally, deploy the optimized upper-layer planning strategy and lower-layer execution strategy to the energy storage scheduling task, and the upper-layer planning agent generates charging and discharging power budget in the day-ahead scheduling stage, and the lower-layer execution agent accepts the budget instruction and adjusts the power. The application realizes the coordinated optimization of short-term economic benefits and long-term life protection of energy storage equipment by constructing a hierarchical decision architecture.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of power grid energy storage operation optimization, and particularly relates to an energy storage optimization method and system based on hierarchical multi-agent reinforcement learning. BACKGROUND

[0002] With the increasing penetration of renewable energy in power systems, energy storage systems, as key technical equipment for balancing power supply and demand and improving power system stability, their operation optimization problems are increasingly prominent. Traditional energy storage optimization methods mainly focus on maximizing short-term economic benefits, achieving peak-valley electricity price arbitrage through frequent charging and discharging operations, but ignore the long-term life loss of energy storage devices. This short-sighted optimization strategy leads to premature degradation of energy storage devices and significantly reduces the overall life cycle economy.

[0003] Patent application with publication number CN120165417A mainly coordinates hybrid energy storage systems through model optimization and frequency control, focusing on system investment and operation cost and reliability improvement. Patent application with publication number CN119864839A focuses on the configuration and scheduling of microgrid energy storage to improve new energy utilization, reduce costs and ensure system stability. Patent application with publication number CN117955133A optimizes energy storage site selection, capacity and daily operation plan through improved particle swarm algorithm to improve voltage stability, reduce cost and network loss.

[0004] The existing technology realizes the operation optimization of energy storage systems from multiple angles, but still has some shortcomings. First, the overall life cycle economy of energy storage devices is ignored, and the life loss model in the existing technology cannot accurately reflect the complex degradation mechanism of the battery under actual working conditions, resulting in a large deviation between the optimization results and the actual operation effect.

[0005] Secondly, the existing technology generally has the problem of single decision level, using a flat decision structure that cannot effectively handle the coupling relationship between different time scale decision variables. This mixing of long-term planning and short-term execution decisions leads to a sharp increase in computational complexity, making real-time performance difficult to guarantee and limiting the targeted application of different optimization algorithms, which fails to fully exploit the advantages of various algorithms in specific problems. SUMMARY

[0006] To solve the problems in the prior art, the present application provides an energy storage operation optimization method and system based on hierarchical multi-agent reinforcement learning, which realizes the coordinated optimization of short-term economic benefits and long-term life protection of energy storage devices by constructing a hierarchical decision architecture.

[0007] The first aspect of the present application provides an energy storage operation optimization method based on hierarchical multi-agent reinforcement learning, which adopts the following technical solution:

[0008] An initialization hierarchical multi-agent reinforcement learning decision framework is provided, an upper planning agent is used to include an upper planning agent and a lower execution agent, and an inter-layer communication mechanism is defined;

[0009] A battery health state monetization cost quantization model is constructed, an upper reward function balancing electricity bill income and battery life loss cost is designed, and an upper planning strategy is optimized;

[0010] A state space, an action space and a reward function of a lower execution agent are designed, a lower reward function integrating current smoothing and budget tracking accuracy is constructed, and a lower execution strategy is optimized;

[0011] Based on the optimized upper planning strategy and lower execution strategy, the energy storage scheduling task is executed, the upper planning agent generates a charge-discharge power budget in the day-ahead scheduling stage, and the lower execution agent accepts the budget instruction and adjusts the power.

[0012] Further, the upper planning agent generates a power budget allocation vector for the lower execution agent; the lower execution agent executes real-time power adjustment and feeds back the execution result to the upper planning agent;

[0013] The inter-layer message format between the upper planning agent and the lower execution agent is defined as follows:

[0014] ;

[0015] Wherein, is an inter-layer message, is a time stamp of the energy storage system, represents an agent identifier, represents an energy storage control message type, is a control load of the energy storage system.

[0016] Further, the upper reward function is represented by the difference between an electricity bill income incentive term and a battery loss penalty term;

[0017] The calculation method of the electricity bill income incentive term is: according to the electricity price and the charge-discharge power budget at each time, the electricity bill income of a single time period is calculated; the incomes of all 24 time periods are accumulated, and then multiplied by the income weight to obtain the total electricity bill income;

[0018] The calculation method of the battery loss penalty term is: combining the charge-discharge power budget, the state of charge and the state of health of the battery at each time, the loss cost of a single time period is calculated through a battery life loss cost function; the incomes of all 24 time periods are accumulated, and then multiplied by the life loss weight to obtain the cumulative loss cost;

[0019] The state of health SOH of the battery is the degree of battery degradation.

[0020] Further, for the balance relationship of the revenue weight and the life consumption weight , a dynamic weight self-adaptive adjustment mechanism is adopted to realize:

[0021] ;

[0022] ;

[0023] wherein, and are weight adjustment parameters, is the battery health state at the moment .

[0024] Further, the battery life consumption cost function is the product of a life consumption base coefficient, a power influence factor term, an SOC deviation penalty term, and a health state decay term;

[0025] The life consumption base coefficient is the base aging rate of the battery under the battery temperature and the historical cumulative cycle number;

[0026] The power influence factor term takes the absolute value of the charge-discharge power budget as the base number and the health state related power influence coefficient as the exponent;

[0027] The SOC deviation penalty term takes the natural constant e as the base number and the product of the temperature sensitive coefficient and the absolute deviation of the SOC as the exponent; the absolute deviation of the SOC is the absolute value of the difference between the current power and the optimal power point;

[0028] The health state decay term takes the remaining health loss of the battery as the base number and the health state penalty coefficient related to the charge-discharge cycle number as the exponent; the remaining health loss of the battery is represented as .

[0029] Further, the lower layer reward function is the sum of a current smoothing penalty term, a budget tracking penalty term, and a constraint violation penalty term;

[0030] The current smoothing penalty term is the product of the square of the current change value at adjacent moments and the current smoothing weight coefficient, taking a negative value;

[0031] The budget tracking penalty term is the product of the square of the remaining deviation at the current moment and the budget tracking weight coefficient, taking a negative value;

[0032] The constraint violation penalty term is the product of the result of the multi-constraint comprehensive judgment function and the constraint violation penalty coefficient, taking a negative value.

[0033] Further, the current smoothing weight coefficient is adaptively adjusted based on the battery state of health and the battery temperature; the adjusted current smoothing weight coefficient is the product of the basic current smoothing weight, the state of health adjustment factor and the temperature adjustment factor;

[0034] The state of health adjustment factor is expressed as , represents the time; the temperature adjustment factor takes the natural constant e as the base number and the ratio of the absolute value of the temperature difference and the temperature sensitive constant as the exponent; the temperature difference is the difference between the current temperature of the battery and the reference temperature.

[0035] Further, the budget tracking weight coefficient is dynamically adjusted based on the electricity price and the remaining budget deviation; the adjusted budget tracking weight coefficient is the product of the basic budget tracking weight, the electricity price sensitive factor and the deviation amplification factor;

[0036] The electricity price sensitive factor is the ratio of the real-time electricity price and the historical average electricity price, summed with 1; the deviation amplification factor is the ratio of the absolute value of the current time remaining deviation and the rated power of the energy storage system, summed with 1.

[0037] Further, the multi-constraint comprehensive judgment function includes four sub-constraints, including an SOC constraint indication function, a current constraint indication function, a voltage constraint indication function and a temperature constraint indication function, all of which are of Boolean type;

[0038] In the sliding time window, the ratio of the number of violations of each sub-constraint to the total number of samples is calculated, and the output of the sub-constraint is converted into a constraint violation severity index;

[0039] The multi-constraint comprehensive judgment function is expressed as:

[0040] ;

[0041] Among them, 、 、 and are the constraint violation severity indexes of the SOC constraint indication function, the current constraint indication function, the voltage constraint indication function and the temperature constraint indication function, respectively;

[0042] By hierarchical punishment of the violation severity, a constraint violation punishment coefficient is obtained .

[0043] The second aspect of the present application provides an energy storage operation optimization system based on hierarchical multi-agent reinforcement learning, which adopts the energy storage operation optimization method provided by the first aspect of the present application, and the system comprises:

[0044] A decision framework initialization module is configured to initialize a hierarchical multi-agent reinforcement learning architecture, an upper layer planning strategy is configured to include an upper layer planning agent and a lower layer execution agent, and an inter-layer communication synchronization mechanism is defined;

[0045] An upper layer strategy optimization module is configured to construct a monetization cost quantification model of battery health state, design an upper layer reward function balancing electricity bill income and battery life loss cost, and optimize the upper layer planning strategy;

[0046] A lower layer strategy optimization module is configured to design a state space, an action space and a reward function of the lower layer execution agent, construct a lower layer reward function integrating current smoothing and budget tracking accuracy, and optimize the lower layer execution strategy;

[0047] A hierarchical collaborative decision module is configured to deploy the optimized upper layer planning strategy and lower layer execution strategy to the energy storage scheduling task, the upper layer planning agent generates a charge-discharge power budget in a day-ahead scheduling stage, and the lower layer execution agent accepts the budget instruction and adjusts the power.

[0048] The application has the advantages that,

[0049] 1. The hierarchical decision architecture of the application realizes multi-time scale coordination of energy storage system operation optimization, effectively solving the technical problem that short-term income and long-term life loss cannot be balanced in traditional methods.

[0050] 2. In terms of battery life protection, the application innovatively introduces a monetization cost quantification mechanism of battery health state, which converts the abstract life loss concept into specific economic cost indicators, enabling the optimization algorithm to balance short-term electricity bill savings and long-term equipment depreciation under a unified cost framework.

[0051] 3. The multi-agent reinforcement learning framework of the application has excellent environmental adaptability and self-learning ability, and can continuously optimize decision strategies in a complex and variable power grid operation environment.

[0052] 4. In terms of computational efficiency, the hierarchical architecture design of the present application enables the upper layer planning agent to make decisions on a longer time scale, while the lower layer execution agent focuses on fine control on a short time scale, avoiding the dimension disaster problem caused by mixing optimization of different time scale decision variables in traditional methods. This design not only improves the solving efficiency, but also enhances the scalability and modularization of the system. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 An implementation architecture diagram of the energy storage life optimization system. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. The embodiments described in the present application are only a part of the embodiments of the present application, but not all the embodiments. Based on the spirit of the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.

[0055] As an embodiment of the present application, a specific implementation of a method and system for energy storage operation optimization based on hierarchical multi-agent reinforcement learning is described, referring to Figure 1 , Figure 1 An implementation architecture diagram of the energy storage operation optimization system.

[0056] Step 1: System architecture initialization and agent design;

[0057] 1.1: Initialize the hierarchical multi-agent reinforcement learning system architecture

[0058] Construct the hierarchical multi-agent reinforcement learning architecture of the energy storage life optimization system, including the network structure of the upper layer planning agent and the lower layer execution agent . Specifically:

[0059] 1.1.1: The upper layer planning agent adopts the ActorCritic architecture of the multi-agent proximal policy optimization algorithm (MAPPO), which is responsible for the coordinated decision of the day-ahead power budget of the energy storage system.

[0060] In the upper layer planning agent, is the policy network of the upper layer agent, the input is the upper layer state , and the output is the power budget allocation strategy; is the value network of the upper layer agent, which evaluates the expected return value under a given state;

[0061] 1.1.2: Lower-layer executing agent, adopting the Twin Delayed Deep Deterministic Policy Gradient (TD3) architecture, responsible for real-time power regulation control of the energy storage system.

[0062] In the lower-layer executing agent, For the Actor network of the lower-layer agent, the input is the lower-layer state , and the output is the deterministic power regulation action;

[0063] For the first Critic network of the lower-layer agent, it evaluates the Q value of the state-action pair, For the second Critic network of the lower-layer agent, it constitutes a double Q learning mechanism with the first Critic network;

[0064] For the parameter vector of the neural network, including the weight matrix and the bias vector, For the state vector at time , containing the current operating state information of the energy storage system; For the action vector at time , representing the control action selected by the agent in the current state.

[0065] 1.2: Establish inter-layer communication protocol and data exchange interface;

[0066] Design an inter-layer communication architecture for the energy storage system agent based on asynchronous message passing, and define a special message format for the energy storage life optimization. The inter-layer message Format definition:

[0067] ;

[0068] The parameter interpretation of the inter-layer message format is given in Table 1 below.

[0069] Table 1 Parameter interpretation of the energy storage system message format

[0070]

[0071] Inter-layer communication mechanism: The upper-layer agent generates a power budget allocation vector once every 24 hours (day-ahead) and sends it to the lower-layer agent through the message queue. The lower-layer agent executes real-time power regulation every 1 minute and feeds back the execution results to the upper-layer agent through the buffer mechanism, realizing information transmission for the hierarchical control of the energy storage system.

[0072] 1.3: Design the energy storage system agent network parameter initialization strategy;

[0073] The embodiment is based on the initialization of the energy storage control neural network weight by using the Xavier initialization method, and a special initialization strategy is designed for the energy storage system. The specific improvement is:

[0074] For the output layer of the energy storage system strategy network, a smaller initialization variance is used to avoid excessive aggressiveness of the initial power control strategy and prevent unnecessary damage to the battery life.

[0075] wherein the soft update rate is in the range of 0.001-0.01, and the value is designed for the long-time running characteristics of the energy storage system to ensure the stability of the target network parameter update and avoid the impact of the drastic fluctuations of the energy storage control strategy on the battery life.

[0076] The application constructs a hierarchical multi-agent reinforcement learning decision-making framework, which divides the energy storage optimization problem into two levels of upper planning layer and lower execution layer. The upper planning agent is responsible for long-term energy storage charging and discharging budget allocation, adopts a multi-agent proximal policy optimization algorithm, takes the battery health state monetization cost as the core objective function, and formulates a day-ahead charging and discharging plan. The lower execution agent is responsible for real-time charging and discharging power control, adopts a double-delay deep deterministic policy gradient algorithm, and realizes smooth control of charging and discharging current under the premise of meeting the upper budget constraint.

[0077] Step 2: Design and implementation of the upper planning agent;

[0078] 2.1: Define the state space, action space and reward function of the upper planning agent;

[0079] 2.1.1: State space of the upper planning agent The state vector at time is represented as:

[0080]

[0081] wherein , , are the battery health state (with a value range of [0, 1], 1 representing a brand-new battery and 0 representing complete failure), the state of charge of the battery (with a value range of [0, 1], representing the ratio of the current storage capacity to the rated capacity), and the ambient temperature (in Celsius); , and are the 24-hour hourly load prediction value (in kW), the photovoltaic output prediction value (in kW) and the electricity price information (in yuan / kWh) from time to , respectively.

[0082] As an optional approach in this embodiment, the battery state of health (SOH) can be determined by combining an online estimation method based on capacity decay and a frequency domain analysis method based on electrochemical impedance spectroscopy (EIS).

[0083] As an optional approach in this embodiment, the State of Charge (SOC) is obtained by: calculating the real-time SOC using the ampere-hour integration method, and then dynamically correcting the cumulative error of the ampere-hour integration using the extended Kalman filter method.

[0084] As an optional method in this embodiment, load forecast value The load forecasting model is obtained using a Long Short-Term Memory (LSTM) network combined with meteorological data; the load forecasting model is expressed as:

[0085] ;

[0086] in, For the first Load forecast at time of day Represents the LSTM model. for Historical load data for the 24 hours prior to the current time. and The first Forecast temperature and humidity values ​​at any given time; This is a weekday type.

[0087] As an optional method in this embodiment, photovoltaic power output prediction The method of obtaining satellite cloud imagery combined with meteorological data correction is employed.

[0088] ;

[0089] in, For the first Predicted photovoltaic output at any given time For photovoltaic installed capacity, For the first Satellite cloud image data at any given time. and They are respectively Predicted temperature and wind speed values ​​at any given time; For convolutional neural networks used to extract cloud map features, This is a correction factor for the attenuation of temperature and wind speed in numerical weather forecasts.

[0090] As an optional method in this embodiment, electricity price information A forecasting model based on real-time electricity prices and day-ahead electricity prices is used:

[0091] ;

[0092] wherein, is the predicted electricity price at time t, is the base electricity price, is the market supply-demand function, is the policy adjustment factor; and and denote the electricity demand and supply at time t.

[0093] 2.1.2: Upper agent action space The feasible region is defined as the 24-hour charge-discharge power budget allocation vector:

[0094]

[0095] wherein, is the charge-discharge power budget at time t, positive value represents discharging, negative value represents charging, unit is kW; and are the maximum allowed discharging power and the maximum allowed charging power of the energy storage system, respectively; is the action of the upper planning agent, in this embodiment, represents the charge-discharge power budget allocated to the energy storage system. 2.2: Establish a monetized cost quantification model of battery state of health, and design a reward function that comprehensively considers electricity bill income and battery life loss cost;

[0096]

[0097] ; wherein,

[0098] represents the reward value obtained by the upper agent at time t; is the income weight at time t; is the life loss weight at time t; is the time step, usually 1 hour; is the battery life loss cost function; 2.2.1: For the income weight and the life loss weight , a dynamic weight self-adaptive adjustment mechanism is adopted to achieve effective balance between short-term income and long-term life loss:

[0099]

[0100] ; ​​​​​

[0101] ;

[0102] wherein, , are weight adjustment parameters, which are set to 2 and 3 respectively in this embodiment. When the state of health is high (SOH close to 1), is increased, and more attention is paid to short-term benefits; when the state of health is low (SOH close to 0), is increased, and more attention is paid to life protection.

[0103] 2.2.2: Battery life loss cost function is defined as:

[0104] ;

[0105] wherein, is a life loss base coefficient, which is related to battery temperature and the number of charge and discharge cycles , wherein, is a base loss coefficient, which is obtained through accelerated aging experiments; is an activation energy, which is obtained by fitting aging data at different temperatures through the Arrhenius equation; is a cycle aging index, which is obtained by regression analysis of long-term cycle experiment data; is an ideal gas constant, is a rated cycle life.

[0106] is a power influence factor term, is a state of health related power influence coefficient, the lower the SOH, the greater the impact of charge and discharge power on life; in this embodiment, , which is obtained based on an online identification method of battery internal resistance change;

[0107] is an SOC deviation penalty, is a temperature related SOC deviation sensitive coefficient; in this embodiment, , which is obtained through thermodynamic modeling and experimental data fusion;

[0108] represents the optimal working state of charge point of the battery, based on the electrochemical characteristics of lithium ion batteries, the internal stress of the battery is minimum and the life loss is lowest when the SOC is near 50%;

[0109] is a state of health decay term, is a cycle number related state of health penalty coefficient; in this embodiment, ​Time-varying parameter estimation based on battery capacity degradation history data.

[0110] 2.3: Implement multi-agent proximal policy optimization training algorithm for upper-layer planning agent. Adopt distributed training architecture, each agent maintains independent policy network and value network. Policy network update adopts clipping objective function.

[0111] The application establishes a battery health state dynamic tracking and monetization cost quantification model. By monitoring key parameters such as battery capacity degradation and internal resistance growth in real time, a multi-dimensional evaluation index system reflecting the true health state of the battery is constructed. The battery life loss is converted into quantifiable economic cost, and the accurate mapping relationship between cycle number, discharge depth, charge-discharge rate and life loss is established, providing accurate cost reference for upper-layer planning decision.

[0112] Step 3: Design and implement lower-layer executive agent; design state space, action space and reward function of lower-layer executive agent;

[0113] 3.1: Design state space of lower-layer executive agent , then The state vector at time t is :

[0114] ;

[0115] Wherein, SOC is the state of charge of the battery, with a value range of [0, 1]; I is the battery current, positive value indicating discharge current and negative value indicating charging current; V is the battery voltage; T is the battery temperature; P is the upper-layer power budget, positive value indicating planned discharge and negative value indicating planned charge;

[0116] is the remaining budget deviation, indicating the cumulative budget execution deviation from the start of the current hour to the end of the hour, calculated as:

[0117] ;

[0118] The meaning of this formula is that from the start time of the current hour to the current time , the cumulative deviation of power budget and actual execution power at time t, is weighted and summed by minute step .

[0119] Positive value indicates actual execution is lower than budget, power needs to be increased; negative value indicates actual execution is higher than budget, power needs to be reduced.

[0120] 3.2: Design the action space of the lower-level execution agent The lower-level execution agent selects the action vector according to the state vector at each moment as the real-time charging and discharging power adjustment amount ;

[0121] ;

[0122] wherein, is the actual output power, the value range of , is the power adjustment upper limit; for the real-time adjustment amount , a comprehensive adjustment based on the current deviation and the predicted deviation is adopted for control, which is expressed as:

[0123] ;

[0124] wherein, is the feedback control gain, used to correct the existing deviation, which is set to 0.6 in this embodiment; is the feedforward control gain, used to prevent future deviation, which is set to 0.4 in this embodiment; is the natural power output based on the prediction of the current battery state.

[0125] As an optional way of this embodiment, the power adjustment upper limit needs to meet the joint restriction of multiple constraints, including the maximum charging and discharging power constraint of the battery, the SOC safety boundary constraint and the temperature protection constraint, and the minimum value of the three is taken as the global upper limit.

[0126] 3.3: Construct the reward function and current smoothing target of the lower-level execution agent

[0127] The lower-level reward function mainly focuses on current smoothing and budget tracking accuracy. The current smoothing is expressed as the square of the difference of the battery current at adjacent time points. The larger the square of the difference, the worse the current smoothing, and vice versa. The budget tracking accuracy is expressed as the square of the remaining budget deviation. The larger the square of the deviation, the worse the budget tracking accuracy, and vice versa. The lower-level reward function is expressed as:

[0128] ;

[0129] In the lower-level reward function:

[0130] (1) The current smoothing weighting coefficient has the dimension of A. -2 Adaptive adjustment based on battery health status and temperature:

[0131] ;

[0132] In the formula, The base current smoothing weight; As a health condition regulating factor, more attention is paid to current smoothing during battery aging; This is a temperature adjustment factor that increases smoothing requirements when the temperature deviates from the optimal value.

[0133] (2) The weighting factor for budget tracking is in kW. -2 Dynamic adjustments based on electricity price and budget deviations:

[0134]

[0135] In the formula, Track weights for the basic budget; Electricity price is a sensitive factor, and budget execution is given more attention when electricity prices are high; This is the deviation amplification factor; the larger the deviation, the heavier the penalty. For real-time electricity prices, The historical average electricity price This refers to the rated power of the energy storage system.

[0136] (3) To constrain the penalty coefficient for violations, To constrain the severity of violations; The multi-constraint comprehensive judgment function includes four sub-constraints, and is expressed as follows:

[0137] .

[0138] Each sub-constraint is a Boolean type, taking the value 0 or 1. Specifically:

[0139] SOC constraint indicator function is ;

[0140] The current constraint indication function is ;

[0141] The voltage constraint indication function is ;

[0142] Temperature constraint indication function is ;

[0143] In the expressions of the four sub-constraints, This is a SOC (State of Charge) limit violation flag. and respectively maximum and minimum allowed SOC; is a current out-of-limit flag, is a maximum allowed current; is a voltage out-of-limit flag, and respectively maximum and minimum allowed voltage; is a temperature out-of-limit flag, and respectively maximum and minimum working temperature.

[0144] The sub-constraint indicator function is of Boolean type, and the embodiment converts the discrete judgment into a constraint violation severity indicator through time dimension expansion . Specifically: in the sliding time window (such as 60 seconds, sampling once per second), the violation time proportion of each sub-constraint is calculated, that is, the ratio of the number of violations of each sub-constraint to the total number of samples.

[0145] The ratio is converted into a specific value of the constraint violation penalty coefficient through a hierarchical penalty based on the violation degree:

[0146]

[0147] As an optional embodiment of the present application, the battery state monitoring and safety guarantee mechanism provides real-time and accurate battery state information for the upper layer planning agent, and provides safety constraint guarantee for the lower layer execution agent. Specifically: a composite model based on cycle counting and calendar aging is used to update the state of health SOH of the battery in real time, and the update result is provided as input to the battery life loss cost function in step 2.2.

[0148] A safety verification module for the execution agent is constructed to judge in real time whether the control instruction violates the operation constraints such as voltage, power, SOC, etc., and to provide real-time feedback for the constraint violation indicator function in step 3.2.

[0149] An extended Kalman filter is introduced to dynamically estimate the battery SOC, ensuring that the reward function SOC-related term of the upper layer planning agent has actual reference value.

[0150] A battery temperature management subsystem is deployed to collect temperature information in real time and supplement the thermal state variables of the lower layer state space.

[0151] The embodiment deeply integrates the existing battery management technology with the reinforcement learning framework, ensuring that the agent decision is based on accurate physical state, and avoiding control failure caused by state estimation error.

[0152] As an optional embodiment of the present application, a cascade utilization battery coordination control mechanism is provided for the scenario of a plurality of battery modules with inconsistent health states in an energy storage system, to coordinate the power and safety control of each module, prolong the overall system life, and improve the system energy utilization rate and regulation stability. Specifically: the state space and control parameters constructed in step 1 for a single agent are extended to a multi-module architecture, so that each battery module independently maintains its parameter vector, achieving personalized modeling of state description and control strategy.

[0153] Based on the upper layer state space defined in step 2, the SOC, SOH and other indicators of multiple battery modules are introduced to form a joint state representation, enhancing the adaptability of the strategy network to system heterogeneity.

[0154] The lower layer execution agent in step 3 outputs the total system power. In this embodiment, a power decomposition mechanism is introduced to distribute control instructions according to module health and state, ensuring that control instructions are landed on a module level in the physical layer.

[0155] Based on the constraint mechanism in step 3, the SOC balancing constraint and power distribution balancing requirement between modules are added.

[0156] This embodiment can solve the problem of large differences between cascade battery modules and difficult control in the prior art. By introducing a multi-module difference perception and collaborative regulation mechanism under the reinforcement learning architecture, the system-level coupling of battery state estimation, power control, constraint execution and health management is achieved, effectively improving the compatibility and scheduling stability of the energy storage system for cascade batteries.

[0157] In the application of cascade utilization batteries, the application has developed a differentiated parameter adaptive mechanism to effectively address the technical challenge of inconsistent performance parameters of cascade utilization batteries. By establishing a personalized model parameter library and online learning mechanism, the system can automatically identify and adapt to the characteristic differences of different battery modules, achieving coordinated optimization control of the cascade battery pack.

[0158] As an optional embodiment of the present application, a multi-source information fusion and prediction mechanism is used to improve the environmental perception ability of the energy storage system agent, ensure the accuracy of the predicted variables in the state space, and provide high-confidence decision-making basis for the upper and lower layer agents. Specifically: the upper layer state space constructed in step 2.1 is provided with key prediction variables, including photovoltaic output, load change trend, electricity price fluctuation, etc., to ensure the accuracy and reliability of the predicted components. The prediction accuracy will directly affect the calculation accuracy of the upper layer reward function in step 2.2, thereby affecting the training stability and running effect of the reinforcement learning strategy.

[0159] The prediction result can provide a basis for making a day-ahead charging and discharging power budget in step 3.1, so that the system operation is forward-looking. Prediction errors (such as photovoltaic and load prediction deviations) provide a correction reference for real-time power regulation in step 3.2, enhancing the robustness of control.

[0160] The application integrates the multi-source information fusion mechanism of power grid load prediction and energy storage state estimation, combines historical operation data, weather information, load characteristics and other multi-dimensional information, and constructs a high-precision load prediction model. At the same time, a multi-state online estimation algorithm of energy storage equipment is established to update the key parameters such as battery state of charge, health state and temperature state in real time, providing accurate state information input for intelligent agent decision-making.

[0161] As an optional embodiment of the application, a training algorithm is designed for the intelligent agent to improve the learning efficiency and control performance of the distributed intelligent agent in energy storage scheduling, and to ensure that the intelligent agent can stably converge in a complex environment and obtain a robust optimal strategy. Specifically: for the double-layer intelligent agent architecture proposed in step 1, a cooperative training process of the upper-layer multi-agent proximal policy optimization (MAPPO) and the lower-layer deterministic policy gradient (TD3) algorithm is designed to ensure that the upper and lower layer strategies are improved simultaneously at their respective scales.

[0162] The asynchronous message passing mechanism in step 1.2 is improved to improve the training data synchronization efficiency between distributed intelligent agents and alleviate the impact of communication bottlenecks on the training process.

[0163] Combining the long-term planning characteristics of the upper-layer decision-making intelligent agent in step 2 and the fast response characteristics of the lower-layer control intelligent agent in step 3, global and local experience pools are designed respectively to realize a hierarchical and differentiated sample replay strategy and improve sample utilization efficiency.

[0164] For the system safety constraints involved in step 3.2, a penalty mechanism is introduced during the training process to improve the agent's perception ability of the constraint boundary, ensuring the deployability and safety of the training results.

[0165] Through the design of a provably safe constraint verification module, it is ensured that the energy storage system can strictly comply with the safety boundary under any operating condition. Unlike the soft constraint or penalty function method used in existing technologies, the safety verification mechanism of the application can mathematically prove the certainty of constraint satisfaction, eliminating the risk of constraint violation that may occur in traditional methods.

[0166] As an optional embodiment of the application, a system integration and performance evaluation mechanism is provided for integrating all the modules designed in the above embodiments of the application to build a complete energy storage life optimization system, and comprehensively evaluating the system performance under a unified framework. Specifically:

[0167] The integrated double-layer intelligent agent architecture, state space modeling and layered decision mechanism, state monitoring method, coordination control strategy, prediction mechanism and training algorithm are combined to build an end-to-end intelligent energy storage system.

[0168] A performance evaluation system covering indexes such as economy, reliability, response speed, life influence, etc. is designed to quantitatively evaluate the operation effect of the proposed control strategy and provide a decision basis for actual deployment.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them, although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific embodiments of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered within the protection scope of the claims of the present application.

Claims

1. A method for optimizing energy storage operation based on hierarchical multi-agent reinforcement learning, characterized in that, A hierarchical multi-agent reinforcement learning architecture for an energy storage lifetime optimization system is constructed, comprising an upper-layer planning agent and a lower-layer execution agent. The upper-layer planning agent adopts the ActorCritic architecture of the multi-agent near-end policy optimization algorithm and is responsible for coordinating and deciding the day-ahead power budget of the energy storage system. The lower-layer execution agent adopts a dual-delay deep deterministic policy gradient algorithm architecture and is responsible for the real-time power regulation and control of the energy storage system. The hierarchical multi-agent reinforcement learning architecture, including the upper-layer planning agent and the lower-layer execution agent, is initialized, and the inter-layer communication mechanism is defined. Construct a quantitative model of the monetization cost of battery health status, design an upper-level reward function that balances electricity revenue and battery life depreciation cost, and optimize the upper-level planning strategy. The upper-level reward function is represented by the difference between the electricity revenue incentive term and the battery loss penalty term; The calculation method for the electricity revenue incentive is as follows: calculate the electricity revenue for a single time period based on the electricity price and charging / discharging power budget at each moment; sum up the revenue of all 24 time periods, and then multiply by the revenue weight to obtain the total electricity revenue. The battery loss penalty is calculated as follows: combining the charging and discharging power budget, battery state of charge, and battery health status at each moment, the loss cost for a single time period is calculated through the battery life loss cost function; the gains of all 24 time periods are summed up and multiplied by the life loss weight to obtain the cumulative loss cost. The State of Health (SOH) of the battery refers to the degree of battery degradation. Design the state space, action space, and reward function of the lower-level executive agent, construct a lower-level reward function that integrates current smoothness and budget tracking accuracy, and optimize the lower-level execution strategy; Based on the optimized upper-layer planning strategy and lower-layer execution strategy, the energy storage scheduling task is executed. The upper-layer planning agent generates the charging and discharging power budget during the day-ahead scheduling phase, and the lower-layer execution agent receives the budget instructions and performs power adjustment.

2. The energy storage operation optimization method based on hierarchical multi-agent reinforcement learning according to claim 1, characterized in that, The upper-layer planning agent generates a power budget allocation vector and sends it to the lower-layer execution agent; the lower-layer execution agent performs real-time power adjustment and feeds back the execution result to the upper-layer planning agent. The inter-layer message format between the upper-layer planning agent and the lower-layer execution agent is defined as follows: ; in, For inter-layer messages, For the energy storage system's operating timestamp, Represents the agent identifier. Indicates the type of energy storage control message. For controlling the load of the energy storage system.

3. The energy storage operation optimization method based on hierarchical multi-agent reinforcement learning according to claim 1, characterized in that, For the yield weight and lifetime loss weight The balance relationship is achieved using a dynamic weight adaptive adjustment mechanism: ; ; In the formula, and For weight adjustment parameters, For a moment Battery health status.

4. The energy storage operation optimization method based on hierarchical multi-agent reinforcement learning according to claim 1, characterized in that, The battery life loss cost function is the product of the life loss base coefficient, the power influence factor, the SOC deviation penalty, and the health state decay. The life loss baseline coefficient is the basic aging rate of the battery under battery temperature and historical cumulative cycle number. The power impact factor item has the absolute value of the charging and discharging power budget as the base and the index as the power impact coefficient related to the health status. The SOC deviation penalty term has the natural constant e as the base and the exponent is the product of the temperature sensitivity coefficient and the absolute SOC deviation; the absolute SOC deviation is the absolute value of the difference between the current power level and the optimal power level. The health status degradation term uses the remaining health loss of the battery as the base and the exponent as a health status penalty coefficient related to the number of charge-discharge cycles; the remaining health loss of the battery is expressed as... .

5. The energy storage operation optimization method based on hierarchical multi-agent reinforcement learning according to claim 1, characterized in that, The lower-level reward function is the sum of the current smoothing penalty, the budget tracking penalty, and the constraint violation penalty. The current smoothing penalty term is the product of the square of the current change value at adjacent time points and the current smoothing weight coefficient, and is negative. The budget tracking penalty term is the product of the square of the remaining deviation at the current time and the budget tracking weight coefficient, and is negative. The constraint violation penalty term is the product of the result of the multi-constraint comprehensive judgment function and the constraint violation penalty coefficient, and is negative.

6. The energy storage operation optimization method based on hierarchical multi-agent reinforcement learning according to claim 5, characterized in that, The current smoothing weight coefficient is adaptively adjusted based on the battery health status and battery temperature; the adjusted current smoothing weight coefficient is the product of the base current smoothing weight, the health status adjustment factor, and the temperature adjustment factor. The health status regulation factor is represented as , The time is indicated; the temperature adjustment factor is based on the natural constant e, and the exponent is the ratio of the absolute value of the temperature difference to the temperature sensitivity constant; the temperature difference is the difference between the current battery temperature and the reference temperature.

7. The energy storage operation optimization method based on hierarchical multi-agent reinforcement learning according to claim 5, characterized in that, The budget tracking weight coefficient is dynamically adjusted based on the deviation between electricity price and remaining budget; the adjusted budget tracking weight coefficient is the product of the basic budget tracking weight, the electricity price sensitivity factor, and the deviation amplification factor. The electricity price sensitivity factor is the ratio of the real-time electricity price to the historical average electricity price, summed with 1; the deviation amplification factor is the ratio of the absolute value of the remaining deviation at the current moment to the rated power of the energy storage system, summed with 1.

8. The energy storage operation optimization method based on hierarchical multi-agent reinforcement learning according to claim 5, characterized in that, The multi-constraint comprehensive judgment function includes four sub-constraints, including the SOC constraint indication function, the current constraint indication function, the voltage constraint indication function, and the temperature constraint indication function, all of which are Boolean type; Within the sliding time window, the ratio of the number of violations of each sub-constraint to the total number of samples is calculated, and the output of the sub-constraint is converted into a constraint violation severity index. The multi-constraint comprehensive judgment function is expressed as: ; in, , , and These are the severity indicators of constraint violations for the SOC constraint indicator function, current constraint indicator function, voltage constraint indicator function, and temperature constraint indicator function, respectively. By classifying the penalties according to the severity of the violation, the constraint violation penalty coefficient is obtained. .

9. An energy storage operation optimization system based on hierarchical multi-agent reinforcement learning, running the energy storage operation optimization method as described in any one of claims 1-8, characterized in that, The system includes: The decision framework initialization module is used to initialize the hierarchical multi-agent reinforcement learning architecture, including the upper-layer planning agent and the lower-layer execution agent, and to define the inter-layer communication and synchronization mechanism. The upper-level strategy optimization module is used to build a quantitative model of the monetization cost of battery health status, design an upper-level reward function that balances electricity revenue and battery life loss cost, and optimize the upper-level planning strategy. The lower-level strategy optimization module designs the state space, action space, and reward function of the lower-level execution agent, constructs a lower-level reward function that integrates current smoothness and budget tracking accuracy, and optimizes the lower-level execution strategy. The hierarchical collaborative decision-making module is used to deploy the optimized upper-level planning strategy and lower-level execution strategy to the energy storage scheduling task. The upper-level planning agent generates the charging and discharging power budget during the day-ahead scheduling phase, and the lower-level execution agent receives the budget instructions and performs power adjustment.

Citation Information

Patent Citations

  • Power distribution network energy storage optimization configuration method and system

    CN117955133A

  • Micro-grid energy storage optimal configuration method and system

    CN119864839A

  • Optimization method for hybrid energy storage system of power distribution network

    CN120165417A

  • Dynamic charging and discharging strategy collaborative optimization method for prolonging service life of energy storage system

    CN120601494A

  • Industrial and commercial energy storage power station peak-valley arbitrage reinforcement learning strategy considering participation in frequency modulation

    CN120638383A