Robust optimization configuration method for shared energy storage based on genetic algorithm and reinforcement learning

By constructing a unified equivalent model and a two-layer collaborative optimization framework for shared energy storage systems, and combining genetic algorithms and reinforcement learning, the stability and economic issues of shared energy storage systems under load fluctuations and uncertainties are solved, achieving efficient collaborative optimization and robust allocation of energy storage resources.

CN121663580BActive Publication Date: 2026-05-19SICHUAN SIJI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN SIJI TECHNOLOGY CO LTD
Filing Date
2026-02-09
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In existing technologies, shared energy storage systems suffer from insufficient stability and economic impact under load fluctuations and uncertainties. Traditional methods are insufficient to achieve coordinated and optimized allocation of centralized and distributed energy storage resources.

Method used

A unified equivalent model for centralized electrochemical energy storage and distributed generalized energy storage is constructed using a genetic algorithm and reinforcement learning approach. A two-layer collaborative optimization framework is designed, which combines a multi-agent proximal policy optimization algorithm and a time-aware adaptive affine decision rule to achieve robust optimization configuration of the energy storage system.

Benefits of technology

It improves the overall economic efficiency and operational efficiency of shared energy storage systems, enhances the system's adaptability and stability in complex and uncertain environments, and improves resource utilization efficiency and engineering practical value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121663580B_ABST
    Figure CN121663580B_ABST
Patent Text Reader

Abstract

The application discloses a shared energy storage robust optimization configuration method based on a genetic algorithm and reinforcement learning, relates to the technical field of intelligent scheduling, and realizes unified description and collaborative configuration of multiple types of energy storage resources and user-side adjustable resources by constructing a centralized electrochemical energy storage and distributed generalized energy storage unified equivalent model, thereby avoiding the low resource utilization efficiency problem caused by traditional dispersed modeling; by designing a double-layer collaborative optimization framework of a shared energy storage operator and a user group, a genetic algorithm is combined with a multi-agent proximal policy optimization algorithm, coordinated optimization of investment configuration decisions and operation scheduling strategies is realized, and the overall economy and operation efficiency of the shared energy storage system are improved; and an adaptive affine decision rule with time sequence perception is introduced, the conservativeness of a traditional robust optimization method is effectively reduced while considering load uncertainty, and the stability and robustness of system operation benefits are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent scheduling technology, specifically to a robust optimization configuration method for shared energy storage based on genetic algorithms and reinforcement learning. Background Technology

[0002] With the increasing proportion of renewable energy integration and the growing diversification of energy consumption structures, load volatility and uncertainty in power system operation have significantly increased. Energy storage systems, as an important means to improve system flexibility and operational safety, are often configured independently and operated in a decentralized manner in practical applications, which can easily lead to problems such as redundant investment, low utilization rates, and insufficient coordinated dispatch capabilities. Shared energy storage, through centralized construction, unified operation, and service provision to multiple users, helps improve the overall utilization efficiency of energy storage resources. Meanwhile, on the user side, there are also broad energy storage resources such as distributed energy storage and adjustable loads, whose operating characteristics and constraints vary considerably. Therefore, it is urgent to establish a unified modeling and configuration method to achieve coordinated optimization of centralized and distributed energy storage resources.

[0003] In terms of optimization and scheduling methods, shared energy storage systems typically involve multi-level decision-making relationships between operators and user groups. Traditional deterministic programming or single-level optimization models are insufficient to effectively characterize this type of structural feature. While multi-agent reinforcement learning methods possess distributed decision-making and policy adaptation capabilities, they are prone to insufficient stability of scheduling strategies under conditions of load and operational uncertainty. Traditional robust optimization methods, although capable of ensuring feasibility in extreme scenarios, often employ conservative decision rules, impacting system economics. Therefore, there is an urgent need for a technical solution capable of unified modeling of energy storage resources, two-level collaborative decision-making, and robust optimization configuration under uncertain environments, to balance the system's economics, stability, and adaptability. Summary of the Invention

[0004] The purpose of this application is to provide a robust optimization configuration method for shared energy storage based on genetic algorithms and reinforcement learning, which solves the technical problems of insufficient stability and impact on system economy in the prior art.

[0005] This application is achieved through the following technical solution:

[0006] A robust optimization configuration method for shared energy storage based on genetic algorithms and reinforcement learning includes:

[0007] A unified equivalent model is constructed for centralized electrochemical energy storage and distributed generalized energy storage. The centralized electrochemical energy storage refers to the centralized electrochemical energy storage deployed by the shared energy storage operator, and the distributed generalized energy storage refers to the energy storage architecture composed of multiple generalized energy storage units provided by the user group. The generalized energy storage unit refers to the user-side adjustable resource coupled with the distributed generalized energy storage.

[0008] Based on the unified equivalent model, a two-layer collaborative optimization framework for shared energy storage operators and user groups is constructed. The upper layer of the two-layer collaborative optimization framework uses a genetic algorithm to determine the capacity and power configuration scheme of centralized electrochemical energy storage and the leasing ratio of distributed generalized energy storage. The lower layer models the user group as a multi-agent system and uses a multi-agent near-end strategy optimization algorithm to adaptively schedule distributed generalized energy storage and its corresponding user-side adjustable resources.

[0009] By introducing time-aware adaptive affine decision rules into a two-layer optimization framework, robust optimization configuration of shared energy storage systems can be achieved.

[0010] In one possible implementation, a unified equivalent model for centralized electrochemical energy storage and distributed generalized energy storage is constructed, including:

[0011] This paper constructs system power balance constraints, state-of-charge constraints for centralized electrochemical energy storage, and capacity-power matching constraints. These constraints, along with the capacity-power matching constraints, limit the operational characteristics of the shared energy storage system, thereby constructing a unified equivalent model for centralized electrochemical energy storage and distributed generalized energy storage. The shared energy storage system refers to an energy storage system composed of centralized electrochemical energy storage and distributed generalized energy storage. The unified equivalent model describes the operational characteristics of the shared energy storage system.

[0012] In one possible implementation, the system power balance constraint is:

[0013] ;

[0014] In the formula, P buy (t) represents the power P that the shared energy storage operator purchases from the grid at time t. sell (t) represents the power P sold to the grid by the shared energy storage operator at time t. ES (t) represents the equivalent charge / discharge power of centralized electrochemical energy storage at time t, and P ES (t)>0 indicates that the centralized electrochemical energy storage is in a discharge state; P ES (t)<0 indicates that the centralized electrochemical energy storage is in a charging state; P load (t) represents the equivalent aggregated load power formed by distributed generalized energy storage and its corresponding user-side adjustable resources at time t.

[0015] In one possible implementation, the state-of-charge constraint of the centralized electrochemical energy storage is:

[0016] ;

[0017] ;

[0018] In the formula, SOC(t) represents the state of charge of the stored energy at time t. min State of Charge (SOC) represents the minimum permissible state of charge threshold for centralized electrochemical energy storage. max P represents the maximum permissible state of charge threshold for centralized electrochemical energy storage. ch (t) represents the charging power, P dis (t) represents the discharge power, η c Indicates charging efficiency, η d Δt represents the discharge efficiency, and Δt represents the scheduling cycle duration.

[0019] In one possible implementation, the capacity and power matching constraint is:

[0020] ;

[0021] In the formula, E s P represents the rated capacity of the centralized electrochemical energy storage, φ represents the energy storage rate factor, and P represents the rated capacity of the centralized electrochemical energy storage. s This indicates the rated power of centralized electrochemical energy storage.

[0022] In one possible implementation, the upper layer of the two-layer collaborative optimization framework is configured as follows:

[0023] With the goal of maximizing the annual net revenue of shared energy storage operators, the revenue objective function is constructed as follows:

[0024] ;

[0025] In the formula, This represents the annual net revenue of shared energy storage operators. This refers to the benefits derived from the right to use distributed generalized energy storage. This represents the arbitrage profit obtained from the price difference between purchasing and selling electricity with the power grid; This represents the annualized investment cost of centralized electrochemical energy storage. This represents the operation and maintenance cost of centralized electrochemical energy storage. This represents the cost of distributed generalized energy storage leasing;

[0026] The decision variables are coded as follows:

[0027] ;

[0028] In the formula, x represents the code of the decision variable. This indicates the rated capacity of centralized electrochemical energy storage. This indicates the rated power of centralized electrochemical energy storage. This represents the leasing ratio of the i-th generalized energy storage unit, where i = 1, 2, ..., N, and N represents the number of generalized energy storage units.

[0029] Using the unified equivalent model as a constraint and the objective function as the optimization direction, a genetic algorithm is used to optimize the encoding of decision variables and determine the optimal investment configuration scheme for shared energy storage operators.

[0030] In one possible implementation, the lower layer of the two-layer collaborative optimization framework is configured as follows:

[0031] The user group in the shared energy storage system is modeled as a multi-agent system consisting of N agents; each agent in the agent system corresponds to a generalized energy storage unit.

[0032] The state vector of the agent at time t is defined as:

[0033] ;

[0034] In the formula, Let m represent the state vector of the m-th agent at time t, where m = 1, 2, ..., N. This represents the state of charge of the generalized energy storage unit corresponding to the m-th intelligent agent. This represents the current power demand of the generalized energy storage unit corresponding to the m-th intelligent agent. Indicates renewable energy output. Indicates the unit price of electricity. Indicates the unit price of electricity;

[0035] The action vector of the agent at time t is defined as:

[0036] ;

[0037] In the formula, This represents the action vector of the m-th agent at time t. This indicates the charging power selected by the agent at time t. This indicates the charging power selected by the agent at time t;

[0038] The reward function obtained by the agent executing the action vector under the state vector is defined as:

[0039] ;

[0040] In the formula, This represents the immediate reward function corresponding to the m-th agent. This refers to the economic benefits resulting from charging and discharging activities. Indicates carbon emission penalties, Constraints to ensure fairness in energy allocation among users; α represents the weighting coefficient of the carbon constraint, and β represents the weighting coefficient of the fairness constraint;

[0041] Based on the reward function, the cumulative discount reward maximizing over the entire scheduling period T is obtained as follows:

[0042] ;

[0043] In the formula, This indicates the cumulative discount return. This represents the mathematical expectation operator, and max represents finding the maximum value. This represents the discount factor between (0,1);

[0044] Based on the cumulative discount reward, a multi-agent proximal policy optimization algorithm is used to iterate the agents, and in the next selection of action vectors, the iterated agents are used for selection.

[0045] In one possible implementation, based on the cumulative discount reward, the agents are iteratively evaluated using a multi-agent proximal policy optimization algorithm, including:

[0046] Each agent is treated as an Actor network, and the same evaluation network is set for all agents to obtain the Critic network corresponding to all Actor networks, forming an optimized structure of centralized Critic and distributed Actor.

[0047] The joint state of all agents is processed using a Critic network to obtain an advantage function, and the objective function for policy optimization is obtained based on the advantage function:

[0048] ;

[0049] In the formula, This represents the policy optimization objective function. This indicates finding the minimum value. Represents the dominance function. For shearing function, The shearing threshold, The strategy ratio, and ; This indicates the agent's current policy. This represents the agent's strategy in the previous iteration. This represents the action vector of the agent at time t. This represents the local observation of the agent at time t. The local observation information includes the state of charge of the distributed generalized energy storage unit corresponding to the agent, renewable energy output information, and current electricity purchase price and electricity sales price information.

[0050] Based on the aforementioned strategy, the objective function is optimized, and the Actor network is updated using a gradient descent strategy to achieve iteration of the agent.

[0051] Based on the cumulative discount return, the loss function is obtained as follows:

[0052] ;

[0053] In the formula, Represents the loss function. The parameters representing the Critic network, This represents the fitted system state-value function. This represents the joint global state of all agents at time t. This indicates a cumulative discount report;

[0054] The Critic network is updated by minimizing the value loss function, and the updated Critic network is used to guide the update of the Actor network in the next iteration.

[0055] In one possible implementation, the time-aware adaptive affine decision rule introduced into the two-layer optimization framework is as follows:

[0056] ;

[0057] In the formula, This represents an adaptive affine decision rule. Represents the decision-making benchmark coefficient. Represents the linear response coefficient. This represents the second-order correction coefficient. The uncertain disturbance variable at time t can be used to characterize time-series uncertainties such as renewable energy output deviation, load forecasting error, or electricity price fluctuation. This represents the uncertain disturbance variable at time t-1.

[0058] In one possible implementation, it also includes:

[0059] Within the robust optimization configuration framework, taking the optimal operational benefit of the two-layer collaborative optimization model under uncertain scenarios as the evaluation object, the outer robust optimization objective function is constructed as follows:

[0060] ;

[0061] In the formula, The summation objective function value represents the outer robust optimization, and δ represents the uncertainty control parameter in the robust optimization. This represents the optimal operating benefit of the two-layer collaborative model. This indicates the operating scenario of a shared energy storage system. Indicates expected return. This indicates the calculation of variance. The carbon emission cost is represented by α, the economic weight, β, and γ. The uncertain scenario refers to a typical operating condition consisting of uncertain factors existing in the operation of the shared energy storage system. These uncertain factors include at least fluctuations in renewable energy output and changes in load demand.

[0062] The outer robust optimization objective function guides the operation and control of the shared energy storage system under uncertain environments.

[0063] This application provides a robust optimization configuration method for shared energy storage based on genetic algorithms and reinforcement learning. By constructing a unified equivalent model for centralized electrochemical energy storage and distributed generalized energy storage, it achieves unified description and collaborative configuration of multiple types of energy storage resources and user-side adjustable resources, avoiding the low resource utilization efficiency problem caused by traditional decentralized modeling. By designing a two-layer collaborative optimization framework for shared energy storage operators and user groups, and adopting a combination of genetic algorithms and multi-agent proximal strategy optimization algorithms, it achieves coordinated optimization of investment configuration decisions and operation scheduling strategies, improving the overall economy and operating efficiency of the shared energy storage system. By introducing time-aware adaptive affine decision rules, it effectively reduces the conservatism of traditional robust optimization methods while considering load uncertainty, improving the stability and robustness of system operating benefits, thereby enhancing the adaptability and engineering practical value of the shared energy storage system in complex and uncertain environments. Attached Figure Description

[0064] To more clearly illustrate the technical solutions of the exemplary embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:

[0065] Figure 1 A flowchart illustrating a robust optimization configuration method for shared energy storage based on genetic algorithms and reinforcement learning, provided for embodiments of this application;

[0066] Figure 2 A schematic diagram of a shared energy storage robust optimization configuration device based on genetic algorithm and reinforcement learning provided in an embodiment of this application;

[0067] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0068] Among them, 201-unified configuration module, 202-dual-layer optimization module, 203-optimization and enhancement module, 301-memory, 302-processor, and 303-bus. Detailed Implementation

[0069] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the embodiments and accompanying drawings. The illustrative embodiments and descriptions of this application are only for explaining this application and are not intended to limit this application.

[0070] like Figure 1 As shown in the embodiments of this application, a robust optimization configuration method for shared energy storage based on genetic algorithms and reinforcement learning is provided, including:

[0071] S101. Establish a unified equivalent model for centralized electrochemical energy storage and distributed generalized energy storage; the centralized electrochemical energy storage refers to the centralized electrochemical energy storage deployed by the shared energy storage operator, and the distributed generalized energy storage refers to the energy storage architecture composed of multiple generalized energy storage units provided by the user group; the generalized energy storage unit refers to the user-side adjustable resource coupled with the distributed generalized energy storage.

[0072] This application embodiment constructs a unified equivalent model for centralized electrochemical energy storage and distributed generalized energy storage, realizing a unified description and coordinated configuration of multiple types of energy storage resources and user-side adjustable resources, thus avoiding the problem of low resource utilization efficiency caused by traditional decentralized modeling.

[0073] S102. Based on the unified equivalent model, a two-layer collaborative optimization framework for shared energy storage operators and user groups is constructed. The upper layer of the two-layer collaborative optimization framework uses a genetic algorithm to determine the capacity and power configuration scheme of centralized electrochemical energy storage and the leasing ratio of distributed generalized energy storage. The lower layer models the user group as a multi-agent system and uses a multi-agent near-end strategy optimization algorithm to adaptively schedule distributed generalized energy storage and its corresponding user-side adjustable resources.

[0074] This application embodiment designs a two-layer collaborative optimization framework for shared energy storage operators and user groups, and adopts a combination of genetic algorithm and multi-agent proximal strategy optimization algorithm to achieve coordinated optimization of investment allocation decision and operation scheduling strategy, thereby improving the overall economy and operating efficiency of the shared energy storage system.

[0075] S103. Introduce time-aware adaptive affine decision rules into the two-layer optimization framework to achieve robust optimization configuration of the shared energy storage system.

[0076] This application's embodiments introduce time-aware adaptive affine decision rules, which effectively reduce the conservatism of traditional robust optimization methods while considering load uncertainty, and improve the stability and robustness of system operating benefits, thereby enhancing the adaptability and engineering practical value of shared energy storage systems in complex and uncertain environments.

[0077] In one possible implementation, a unified equivalent model for centralized electrochemical energy storage and distributed generalized energy storage is constructed, including:

[0078] This paper constructs system power balance constraints, state-of-charge constraints for centralized electrochemical energy storage, and capacity-power matching constraints. These constraints, along with the capacity-power matching constraints, limit the operational characteristics of the shared energy storage system, thereby constructing a unified equivalent model for centralized electrochemical energy storage and distributed generalized energy storage. The shared energy storage system refers to an energy storage system composed of centralized electrochemical energy storage and distributed generalized energy storage. The unified equivalent model describes the operational characteristics of the shared energy storage system.

[0079] In one possible implementation, the system power balance constraint is:

[0080] ;

[0081] In the formula, P buy (t) represents the power P that the shared energy storage operator purchases from the grid at time t. sell (t) represents the power P sold to the grid by the shared energy storage operator at time t. ES (t) represents the equivalent charge / discharge power of centralized electrochemical energy storage at time t, and P ES (t)>0 indicates that the centralized electrochemical energy storage is in a discharge state; P ES (t)<0 indicates that the centralized electrochemical energy storage is in a charging state; P load (t) represents the equivalent aggregated load power formed by distributed generalized energy storage and its corresponding user-side adjustable resources at time t.

[0082] In one possible implementation, the state-of-charge constraint of the centralized electrochemical energy storage is:

[0083] ;

[0084] ;

[0085] In the formula, SOC(t) represents the state of charge of the stored energy at time t. min State of Charge (SOC) represents the minimum permissible state of charge threshold for centralized electrochemical energy storage. maxP represents the maximum permissible state of charge threshold for centralized electrochemical energy storage. ch (t) represents the charging power, P dis (t) represents the discharge power, η c Indicates charging efficiency, η d Δt represents the discharge efficiency; Δt represents the scheduling cycle duration, which can be in hours.

[0086] In one possible implementation, the capacity and power matching constraint is:

[0087] ;

[0088] In the formula, E s P represents the rated capacity of the centralized electrochemical energy storage, φ represents the energy storage rate factor, and P represents the rated capacity of the centralized electrochemical energy storage. s This indicates the rated power of centralized electrochemical energy storage.

[0089] In one possible implementation, the upper layer of the two-layer collaborative optimization framework is configured as follows:

[0090] With the goal of maximizing the annual net revenue of shared energy storage operators, the revenue objective function is constructed as follows:

[0091] ;

[0092] In the formula, This represents the annual net revenue of shared energy storage operators. This refers to the benefits derived from the right to use distributed generalized energy storage. This represents the arbitrage profit obtained from the price difference between purchasing and selling electricity with the power grid; This represents the annualized investment cost of centralized electrochemical energy storage. This represents the operation and maintenance cost of centralized electrochemical energy storage. This represents the cost of distributed generalized energy storage leasing, with each cost item in millions of yuan per year.

[0093] The decision variables are coded as follows:

[0094] ;

[0095] In the formula, x represents the code of the decision variable. This indicates the rated capacity of centralized electrochemical energy storage. This indicates the rated power of centralized electrochemical energy storage. This represents the leasing ratio of the i-th generalized energy storage unit, where i = 1, 2, ..., N, and N represents the number of generalized energy storage units.

[0096] Using the unified equivalent model as a constraint and the objective function as the optimization direction, a genetic algorithm is used to optimize the encoding of decision variables and determine the optimal investment configuration scheme for shared energy storage operators.

[0097] For example, it can satisfy E s ≥γP s Under constraints, chromosome individuals are iteratively updated through selection, crossover, and mutation operations. The optimal solution is achieved when the change over 20 consecutive generations is less than 10. -4 Terminate the optimal configuration scheme for output [E] s *,P s *,λ*];where E s * represents the optimal configuration capacity of centralized electrochemical energy storage under the constraints and profit objective function, P s * represents the optimal rated power of centralized electrochemical energy storage that matches the optimal configuration capacity, and λ* represents the optimal storage leasing ratio of distributed generalized energy storage units determined under the optimal configuration scheme.

[0098] In one possible implementation, the lower layer of the two-layer collaborative optimization framework is configured as follows:

[0099] The user group in the shared energy storage system is modeled as a multi-agent system consisting of N agents; each agent in the agent system corresponds to a generalized energy storage unit.

[0100] The state vector of the agent at time t is defined as:

[0101] ;

[0102] In the formula, Let m represent the state vector of the m-th agent at time t, where m = 1, 2, ..., N. This represents the state of charge of the generalized energy storage unit corresponding to the m-th intelligent agent. This represents the current power demand of the generalized energy storage unit corresponding to the m-th intelligent agent. Indicates renewable energy output. Indicates the unit price of electricity. Indicates the unit price of electricity;

[0103] The action vector of the agent at time t is defined as:

[0104] ;

[0105] In the formula, This represents the action vector of the m-th agent at time t. This indicates the charging power selected by the agent at time t. This indicates the charging power selected by the agent at time t;

[0106] The reward function obtained by the agent executing the action vector under the state vector is defined as:

[0107] ;

[0108] In the formula, This represents the immediate reward function corresponding to the m-th agent. This refers to the economic benefits resulting from charging and discharging activities. Indicates carbon emission penalties, Constraints to ensure fairness in energy allocation among users; α represents the weighting coefficient of the carbon constraint, and β represents the weighting coefficient of the fairness constraint;

[0109] Based on the reward function, the cumulative discount reward maximizing over the entire scheduling period T is obtained as follows:

[0110] ;

[0111] In the formula, This indicates the cumulative discount return. This represents the mathematical expectation operator, and max represents finding the maximum value. This represents the discount factor between (0,1);

[0112] Based on the cumulative discount reward, a multi-agent proximal policy optimization algorithm is used to iterate the agents, and in the next selection of action vectors, the iterated agents are used for selection.

[0113] In one possible implementation, based on the cumulative discount reward, the agents are iteratively evaluated using a multi-agent proximal policy optimization algorithm, including:

[0114] Each agent is treated as an Actor network, and the same evaluation network is set for all agents to obtain the Critic network corresponding to all Actor networks, forming an optimized structure of centralized Critic and distributed Actor.

[0115] The joint state of all agents is processed using a Critic network to obtain an advantage function, and the objective function for policy optimization is obtained based on the advantage function:

[0116] ;

[0117] In the formula, This represents the policy optimization objective function. This indicates finding the minimum value. Represents the dominance function. For shearing function, The shearing threshold, The strategy ratio, and ; This indicates the agent's current policy. This represents the agent's strategy in the previous iteration. This represents the action vector of the agent at time t. This represents the local observation of the agent at time t. The local observation information includes the state of charge of the distributed generalized energy storage unit corresponding to the agent, renewable energy output information, and current electricity purchase price and electricity sales price information.

[0118] Based on the aforementioned strategy, the objective function is optimized, and the Actor network is updated using a gradient descent strategy to achieve iteration of the agent.

[0119] Based on the cumulative discount return, the loss function is obtained as follows:

[0120] ;

[0121] In the formula, Represents the loss function. The parameters representing the Critic network, This represents the fitted system state-value function. This represents the joint global state of all agents at time t. This indicates a cumulative discount report;

[0122] The Critic network is updated by minimizing the value loss function, and the updated Critic network is used to guide the update of the Actor network in the next iteration.

[0123] This application embodiment achieves collaborative learning among multiple agents and adaptive scheduling optimization of user groups by fitting the system state value function.

[0124] In step S103, considering load uncertainty, the system load is expressed as:

[0125] ;

[0126] Among them, P load (t) represents the actual load. Indicates the predicted load. This represents the load disturbance term.

[0127] In one possible implementation, the time-aware adaptive affine decision rule introduced into the two-layer optimization framework is as follows:

[0128] ;

[0129] In the formula, This represents an adaptive affine decision rule. Represents the decision-making benchmark coefficient. Represents the linear response coefficient. This represents the second-order correction coefficient. The uncertain disturbance variable at time t can be used to characterize time-series uncertainties such as renewable energy output deviation, load forecasting error, or electricity price fluctuation. Represents the uncertain disturbance variable at time t-1;

[0130] The upper-level optimization model aims at improving system economy and robustness by jointly optimizing the decision baseline coefficients, linear response coefficients, and quadratic correction coefficients (a, b, and c). The lower-level optimization model, after receiving the affine decision parameters determined by the upper-level optimization, optimizes based on the uncertain disturbance variables. The system generates specific time-series scheduling decisions according to the adaptive affine decision rules, thereby enabling the shared energy storage system to operate robustly in uncertain environments.

[0131] The joint optimization process includes: under the condition of a given set of uncertain operating scenarios, simulating and evaluating the operation of the shared energy storage system based on different parameter combinations, calculating the system operating benefits and benefit fluctuations corresponding to each parameter combination, and iteratively updating the decision benchmark coefficient, linear response coefficient, and quadratic correction coefficient with the constraints of improving the system benefit level and reducing benefit fluctuations, until the optimal affine decision parameters that meet the preset convergence conditions are obtained.

[0132] During system operation, the lower-level optimization model receives the affine decision parameters determined by the upper-level optimization model, and calculates the equivalent scheduling decision quantity corresponding to each scheduling moment according to the adaptive affine decision rule based on the uncertain disturbance variables obtained in real time or by prediction.

[0133] The equivalent scheduling decision quantity is further mapped to the charge and discharge power control command of centralized electrochemical energy storage and the user-side adjustable resource control command corresponding to distributed generalized energy storage. Under the conditions of satisfying system power balance constraints, energy storage state of charge constraints and charge and discharge power constraints, it is applied to the operation control of the shared energy storage system, thereby realizing the robust operation of the shared energy storage system in uncertain environments.

[0134] In one possible implementation, it also includes:

[0135] Within the robust optimization configuration framework, taking the optimal operational benefit of the two-layer collaborative optimization model under uncertain scenarios as the evaluation object, the outer robust optimization objective function is constructed as follows:

[0136] ;

[0137] In the formula, The summation objective function value represents the outer robust optimization, and δ represents the uncertainty control parameter in the robust optimization. This represents the optimal operating benefit of the two-layer collaborative model. This indicates the operating scenario of a shared energy storage system. Indicates expected return. This indicates the calculation of variance. The carbon emission cost is represented by α, the economic weight is represented by β, the robustness weight is represented by β, and the environmental weight is represented by γ. The uncertain scenario refers to a typical operating condition consisting of uncertain factors existing in the operation of the shared energy storage system. The uncertain factors include at least the fluctuation of renewable energy output and the change of load demand. The values ​​of each uncertain factor can be obtained based on historical data statistics, probability distribution models, prediction error ranges, or scenario generation methods.

[0138] Under various operating scenarios, a two-layer collaborative optimization model is executed to obtain the annual net income under each operating scenario. The optimal operating income under each operating scenario is evaluated through the outer robust optimization objective function, thereby achieving the collaborative optimization of the economy, robustness and low-carbon operation of the shared energy storage system under uncertain environments.

[0139] Based on the output of the outer robust optimization objective function, the optimal shared energy storage system operation scheme that satisfies the preset economic, robust, and low-carbon operation constraints is determined. The optimal operation scheme is then fed back to the two-layer collaborative optimization model to guide the operation control of the shared energy storage system under uncertain environments, thereby achieving the collaborative optimization of the economic, robust, and low-carbon operation of the shared energy storage system under uncertain environments.

[0140] like Figure 2 As shown, based on the same inventive concept, this application provides a robust optimization configuration device for shared energy storage based on genetic algorithms and reinforcement learning, comprising:

[0141] The unified configuration module 201 is used to construct a unified equivalent model for centralized electrochemical energy storage and distributed generalized energy storage; the centralized electrochemical energy storage refers to the centralized electrochemical energy storage deployed by the shared energy storage operator, and the distributed generalized energy storage refers to the energy storage architecture composed of multiple generalized energy storage units provided by the user group; the generalized energy storage unit refers to the user-side adjustable resources coupled with the distributed generalized energy storage.

[0142] The two-layer optimization module 202 is used to construct a two-layer collaborative optimization framework for shared energy storage operators and user groups based on the unified equivalent model. The upper layer of the two-layer collaborative optimization framework uses a genetic algorithm to determine the capacity and power configuration scheme of centralized electrochemical energy storage and the leasing ratio of distributed generalized energy storage. The lower layer models the user group as a multi-agent system and uses a multi-agent near-end strategy optimization algorithm to adaptively schedule distributed generalized energy storage and its corresponding user-side adjustable resources.

[0143] The optimization and enhancement module 203 is used to introduce time-aware adaptive affine decision rules into the two-layer optimization framework to achieve robust optimization configuration of the shared energy storage system.

[0144] The shared energy storage robust optimization configuration device based on genetic algorithm and reinforcement learning provided in this application embodiment can execute the above synchronization method. Its principle and beneficial effects are similar, and will not be repeated here.

[0145] like Figure 3 As shown, based on the same inventive concept, this application also provides an electronic device, including a processor 302 and a memory 301; the memory 301 and the processor 302 are interconnected via a bus 303.

[0146] The memory 301 stores computer-executed instructions;

[0147] The processor 302 executes the computer execution instructions stored in the memory 301, causing the processor 302 to execute a robust optimization configuration method for shared energy storage based on genetic algorithms and reinforcement learning as described in any embodiment of this application.

[0148] For specific examples, memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out (FIFO) memory, and / or first-in-last-out (FILO) memory, etc.; specifically, processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented using at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). Furthermore, the processor may include a main processor and coprocessors. The main processor, also known as the CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state.

[0149] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the robust optimization configuration method for shared energy storage based on genetic algorithms and reinforcement learning described in any of the above embodiments.

[0150] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the robust optimization configuration method for shared energy storage based on genetic algorithms and reinforcement learning described in any of the above embodiments.

[0151] All or part of the steps in the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof.

[0152] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0153] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0154] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0155] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.

[0156] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A robust optimization configuration method for shared energy storage based on genetic algorithms and reinforcement learning, characterized in that, include: A unified equivalent model is constructed for centralized electrochemical energy storage and distributed generalized energy storage. The unified equivalent model refers to a model that describes the operating characteristics of a shared energy storage system. The centralized electrochemical energy storage refers to the centralized electrochemical energy storage deployed by the shared energy storage operator, and the distributed generalized energy storage refers to an energy storage architecture composed of multiple generalized energy storage units provided by the user group. The generalized energy storage unit refers to the user-side adjustable resource coupled with the distributed generalized energy storage. Based on the unified equivalent model, a two-layer collaborative optimization framework for shared energy storage operators and user groups is constructed. The upper layer of the two-layer collaborative optimization framework uses a genetic algorithm to determine the capacity and power configuration scheme of centralized electrochemical energy storage and the leasing ratio of distributed generalized energy storage. The lower layer models the user group in the shared energy storage system as a multi-agent system composed of N agents. Each agent in the agent system corresponds to a generalized energy storage unit, and a multi-agent proximal strategy optimization algorithm is used to adaptively schedule the distributed generalized energy storage and its corresponding user-side adjustable resources. Introducing time-aware adaptive affine decision rules into a two-layer optimization framework enables robust optimization configuration of shared energy storage systems. The time-aware adaptive affine decision rule introduced into the two-level optimization framework is as follows: ; In the formula, This represents an adaptive affine decision rule. Represents the decision-making benchmark coefficient. Represents the linear response coefficient. This represents the second-order correction coefficient. The uncertain disturbance variable at time t is used to characterize renewable energy output deviation, load forecasting error, or electricity price fluctuation. This represents the uncertain disturbance variable at time t-1.

2. The robust optimization configuration method for shared energy storage based on genetic algorithm and reinforcement learning according to claim 1, characterized in that, A unified equivalent model for centralized electrochemical energy storage and distributed generalized energy storage is constructed, including: The system power balance constraint, the state of charge constraint of centralized electrochemical energy storage, and the capacity-power matching constraint are constructed. The operating characteristics of the shared energy storage system are restricted by the system power balance constraint, the state of charge constraint of centralized electrochemical energy storage, and the capacity-power matching constraint, thereby constructing a unified equivalent model of centralized electrochemical energy storage and distributed generalized energy storage. The shared energy storage system refers to an energy storage system composed of centralized electrochemical energy storage and distributed generalized energy storage.

3. The robust optimization configuration method for shared energy storage based on genetic algorithm and reinforcement learning according to claim 2, characterized in that, The system power balance constraint is: ; In the formula, P buy (t) represents the power P that the shared energy storage operator purchases from the grid at time t. sell (t) represents the power P sold to the grid by the shared energy storage operator at time t. ES (t) represents the equivalent charge / discharge power of centralized electrochemical energy storage at time t, and P ES (t)>0 indicates that the centralized electrochemical energy storage is in a discharge state; P ES (t)<0 indicates that the centralized electrochemical energy storage is in a charging state; P load (t) represents the equivalent aggregated load power formed by distributed generalized energy storage and its corresponding user-side adjustable resources at time t.

4. The robust optimization configuration method for shared energy storage based on genetic algorithm and reinforcement learning according to claim 2, characterized in that, The state-of-charge constraint of the centralized electrochemical energy storage is: ; ; In the formula, SOC(t) represents the state of charge of the stored energy at time t. min State of Charge (SOC) represents the minimum permissible state of charge threshold for centralized electrochemical energy storage. max P represents the maximum permissible state of charge threshold for centralized electrochemical energy storage. ch (t) represents the charging power, P dis (t) represents the discharge power, η c Indicates charging efficiency, η d Δt represents the discharge efficiency, and Δt represents the scheduling cycle duration.

5. The robust optimization configuration method for shared energy storage based on genetic algorithm and reinforcement learning according to claim 2, characterized in that, The capacity and power matching constraint is: ; In the formula, E s P represents the rated capacity of the centralized electrochemical energy storage, where φ represents the energy storage rate factor. s This indicates the rated power of centralized electrochemical energy storage.

6. The robust optimization configuration method for shared energy storage based on genetic algorithm and reinforcement learning according to claim 2, characterized in that, The upper layer of the two-layer collaborative optimization framework is configured as follows: With the goal of maximizing the annual net revenue of shared energy storage operators, the revenue objective function is constructed as follows: ; In the formula, This represents the annual net revenue of shared energy storage operators. This refers to the benefits derived from the right to use distributed generalized energy storage. This represents the arbitrage profit obtained from the price difference between purchasing and selling electricity with the power grid; This represents the annualized investment cost of centralized electrochemical energy storage. This represents the operation and maintenance cost of centralized electrochemical energy storage. This represents the cost of distributed generalized energy storage leasing; The decision variables are coded as follows: ; In the formula, x represents the code of the decision variable. This indicates the rated capacity of centralized electrochemical energy storage. This indicates the rated power of centralized electrochemical energy storage. This represents the leasing ratio of the i-th generalized energy storage unit, where i = 1, 2, ..., N, and N represents the number of generalized energy storage units. Using the unified equivalent model as a constraint and the objective function as the optimization direction, a genetic algorithm is used to optimize the encoding of decision variables and determine the optimal investment configuration scheme for shared energy storage operators.

7. The robust optimization configuration method for shared energy storage based on genetic algorithm and reinforcement learning according to claim 6, characterized in that, The lower layer of the two-layer collaborative optimization framework is configured as follows: The user group in the shared energy storage system is modeled as a multi-agent system consisting of N agents; each agent in the agent system corresponds to a generalized energy storage unit. The state vector of the agent at time t is defined as: ; In the formula, Let m represent the state vector of the m-th agent at time t, where m = 1, 2, ..., N. This represents the state of charge of the generalized energy storage unit corresponding to the m-th intelligent agent. This represents the current power demand of the generalized energy storage unit corresponding to the m-th intelligent agent. Indicates renewable energy output. Indicates the unit price of electricity. Indicates the unit price of electricity; The action vector of the agent at time t is defined as: ; In the formula, This represents the action vector of the m-th agent at time t. This indicates the charging power selected by the agent at time t. This represents the discharge power selected by the agent at time t; The reward function obtained by the agent executing the action vector under the state vector is defined as: ; In the formula, This represents the immediate reward function corresponding to the m-th agent. This refers to the economic benefits resulting from charging and discharging activities. Indicates carbon emission penalties. Constraints to ensure fairness in energy allocation among users; α represents the weighting coefficient of the carbon constraint, and β represents the weighting coefficient of the fairness constraint; Based on the reward function, the cumulative discount reward maximizing over the entire scheduling period T is obtained as follows: ; In the formula, This indicates the cumulative discount return. This represents the mathematical expectation operator, and max represents finding the maximum value. This represents the discount factor between (0,1); Based on the cumulative discount reward, a multi-agent proximal policy optimization algorithm is used to iterate the agents, and in the next selection of action vectors, the iterated agents are used for selection.

8. The robust optimization configuration method for shared energy storage based on genetic algorithm and reinforcement learning according to claim 7, characterized in that, Based on the cumulative discount reward, the agents are iteratively optimized using a multi-agent proximal policy optimization algorithm, including: Each agent is treated as an Actor network, and the same evaluation network is set for all agents to obtain the Critic network corresponding to all Actor networks, forming an optimized structure of centralized Critic and distributed Actor. The joint state of all agents is processed using a Critic network to obtain an advantage function, and the objective function for policy optimization is obtained based on the advantage function: ; In the formula, This represents the policy optimization objective function. This indicates finding the minimum value. Represents the dominance function. For shearing function, The shearing threshold, The strategy ratio, and ; This indicates the agent's current policy. This represents the agent's strategy in the previous iteration. This represents the action vector of the agent at time t. This represents the local observation of the agent at time t. The local observation information includes the state of charge of the distributed generalized energy storage unit corresponding to the agent, renewable energy output information, and current electricity purchase price and electricity sales price information. Based on the aforementioned strategy, the objective function is optimized, and the Actor network is updated using a gradient descent strategy to achieve iteration of the agent. Based on the cumulative discount return, the loss function is obtained as follows: ; In the formula, Represents the loss function. The parameters representing the Critic network, This represents the fitted system state-value function. This represents the joint global state of all agents at time t. Indicates cumulative discount return; The Critic network is updated by minimizing the value loss function, and the updated Critic network is used to guide the update of the Actor network in the next iteration.

9. The robust optimization configuration method for shared energy storage based on genetic algorithm and reinforcement learning according to claim 1, characterized in that, Also includes: Within the robust optimization configuration framework, taking the optimal operational benefit of the two-layer collaborative optimization framework under uncertain scenarios as the evaluation object, the outer robust optimization objective function is constructed as follows: ; In the formula, The summation objective function value represents the outer robust optimization, and δ represents the uncertainty control parameter in the robust optimization. This represents the optimal operational benefit of the two-layer collaborative optimization framework. This indicates the operating scenario of a shared energy storage system. Indicates expected return. This indicates the calculation of variance. The carbon emission cost is represented by α, the economic weight, β, and γ. The uncertain scenario refers to a typical operating condition consisting of uncertain factors existing in the operation of the shared energy storage system. These uncertain factors include at least fluctuations in renewable energy output and changes in load demand. The outer robust optimization objective function guides the operation and control of the shared energy storage system under uncertain environments.