A multi-point energy storage polymerization collaborative operation method, device, equipment and storage medium
By constructing a heterogeneous Markov decision process model and a multi-agent deep deterministic policy gradient algorithm, the problem of balancing dynamic response and economy in multi-point heterogeneous energy storage systems was solved, realizing the coordinated and optimized operation of lithium batteries and flywheels, and improving the stability and efficiency of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUADIAN ELECTRIC POWER SCI INST CO LTD
- Filing Date
- 2025-12-26
- Publication Date
- 2026-05-08
AI Technical Summary
In the collaborative operation of multi-point heterogeneous energy storage (lithium batteries and flywheels), existing technologies show poor adaptability of traditional algorithms and insufficient consideration of heterogeneous characteristics and multi-agent cooperation in reinforcement learning methods, making it difficult to balance dynamic response performance and full life cycle economy.
A heterogeneous Markov decision process model of lithium battery and flywheel energy storage is constructed. Combined with the heterogeneous multi-agent deep deterministic policy gradient algorithm framework, a centralized training and distributed execution mechanism is adopted. The trained algorithm model outputs a collaborative operation scheme with power allocation as the core under the objective constraint.
It achieves dynamic coordination and optimization of multi-point energy storage systems, taking into account dynamic response performance and full life cycle economy, improving operational stability and efficiency, and avoiding the limitations of traditional methods.
Smart Images

Figure CN121485036B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of integrated energy management technology, specifically to a method, apparatus, equipment, and storage medium for multi-point energy storage aggregation and collaborative operation. Background Technology
[0002] Driven by the transformation of the energy structure, the installed capacity of renewable energy (such as photovoltaic and wind power) continues to expand. However, renewable energy is highly intermittent and volatile, and its large-scale integration into the power grid poses serious challenges to the stable operation of the power system, the guarantee of power quality, and the regulation of supply and demand balance.
[0003] To mitigate fluctuations in renewable energy and enhance the flexibility of the power system, energy storage technology has become a key supporting means. Lithium-ion battery energy storage is widely used in the energy storage field due to its advantages such as high energy density and good cycle characteristics; flywheel energy storage, on the other hand, has the characteristics of fast response speed, high power density, long cycle life and is not affected by the depth of charge and discharge, and can quickly mitigate short-term power fluctuations.
[0004] In practical applications, a single type of energy storage often cannot simultaneously meet the system's multi-dimensional requirements for energy capacity, power response speed, and life-cycle economics. Therefore, integrating lithium batteries with heterogeneous energy storage such as flywheels to fully leverage the complementary advantages of different energy storage technologies has become a research hotspot in the energy storage field.
[0005] Currently, research on the collaborative operation of multi-point heterogeneous energy storage aggregation largely focuses on traditional optimization algorithms (such as linear programming and dynamic programming). However, traditional algorithms suffer from problems such as reliance on precise mathematical models, poor adaptability to complex nonlinear systems, and difficulty in handling real-time interactions among multiple agents. With the development of reinforcement learning technology, the Deep Deterministic Policy Gradient (DDPG) algorithm has shown great potential in control problems in continuous action spaces, providing a new technical path for the collaborative operation of heterogeneous energy storage aggregation. However, most existing reinforcement learning-based methods are designed for single-type energy storage or do not fully consider the heterogeneous characteristics and cooperation mechanisms among multiple agents. In multi-point heterogeneous energy storage aggregation scenarios, it is difficult to achieve optimal collaborative operation that balances dynamic response performance and full lifecycle economic efficiency. Summary of the Invention
[0006] To address the technical challenges of existing technologies in the collaborative operation of multi-point heterogeneous energy storage (lithium-ion batteries and flywheels), where traditional algorithms suffer from poor adaptability and reinforcement learning methods fail to adequately consider heterogeneous characteristics and multi-agent cooperation, making it difficult to balance dynamic response performance and lifecycle economics, this invention provides a method, apparatus, device, and storage medium for multi-point energy storage collaborative operation. By characterizing the physical properties of different energy storage media in lithium-ion battery energy storage and flywheel energy storage through heterogeneous intelligent agents, it achieves dynamic coordination and optimization of multi-point energy storage systems, balancing dynamic response performance and lifecycle economics.
[0007] In a first aspect, the present invention provides a method for multi-point energy storage aggregation and collaborative operation, comprising:
[0008] Based on the different physical characteristics of lithium battery energy storage and flywheel energy storage in terms of cycle life, energy density and operational constraints, Markov decision process models for lithium battery energy storage and flywheel energy storage are constructed respectively.
[0009] A heterogeneous multi-agent deep deterministic policy gradient algorithm framework is constructed. Based on two heterogeneous Markov decision process models, the lithium battery energy storage agent and the flywheel energy storage agent are modeled so that each agent is adapted to the physical characteristics of the corresponding energy storage.
[0010] The algorithm framework is trained using a centralized training and distributed execution mechanism. After training, the algorithm model, under the objective constraints, solves the optimal collaborative operation scheme of multi-point energy storage aggregation of lithium batteries and flywheels through a heterogeneous multi-agent collaborative optimization strategy. This scheme takes power allocation as the core and considers dynamic response performance and full life cycle economy.
[0011] The multi-point energy storage aggregation and collaborative operation method provided in this invention accurately captures the differences between the two types of energy storage in terms of cycle life, energy density, and operational constraints by constructing a heterogeneous Markov decision process model for lithium batteries and flywheel energy storage, laying the foundation for subsequent collaborative optimization. The heterogeneous multi-agent deep deterministic policy gradient algorithm framework built based on this model allows each agent to adapt to the corresponding energy storage physical characteristics, avoiding the waste of characteristics caused by a "one-size-fits-all" modeling approach. The "centralized training - distributed execution" mechanism balances global optimization and efficient local execution. The trained model outputs a power allocation-centric solution under target constraints, ensuring dynamic performance through rapid flywheel response and achieving full lifecycle economic efficiency through reasonable lithium battery scheduling and lifespan loss control. This effectively solves the problems of traditional methods being unable to adapt to multi-type energy storage collaboration and struggling to balance performance and economy, significantly improving the stability and efficiency of multi-point energy storage aggregation operation.
[0012] In one optional implementation, the components of the Markov decision process model for lithium battery energy storage include:
[0013] The state space of lithium battery energy storage includes electricity purchase price, photovoltaic power generation, wind power generation, user demand power, proton exchange membrane electrolyzer power consumption, and lithium battery state of charge.
[0014] The operating space of lithium battery energy storage: the charging and discharging power of the lithium battery, and respectively satisfying the maximum charging / discharging power constraints of the lithium battery;
[0015] The reward function for lithium-ion battery energy storage, considering power fluctuation penalties and system costs, is used to quantify the overall benefits of lithium-ion battery operation and is expressed as:
[0016]
[0017] Where B represents lithium battery, The instantaneous reward for lithium battery energy storage at time t; , These are weighting coefficients, determined using a grid search strategy, used to balance the priorities of different optimization objectives; For lithium battery and flywheel multi-point energy storage stations and the main grid, time The power of the exchange; The average switching power within a preset time period; For system transaction costs; The degradation cost of lithium batteries;
[0018] The state transition function of a lithium-ion battery describes the dynamic process of the state of charge changing with charging and discharging power, and is expressed as:
[0019]
[0020] in, This represents the state of charge of the lithium battery at time t; This indicates the state of charge of the lithium battery at time t+1; Improve lithium battery charging efficiency; For lithium battery discharge efficiency; For lithium battery capacity; The time interval between two adjacent decision moments; The charging power of the lithium battery; This refers to the discharge power of the lithium battery.
[0021] This invention clarifies the components of a Markov decision-making process model for lithium-ion battery energy storage. The state space encompasses key influencing factors such as electricity purchase price and wind and solar power, ensuring the model comprehensively reflects actual operating scenarios and provides sufficient information support for agent decision-making. The action space limits charging and discharging power constraints to avoid overcharging and discharging damage to equipment and extend lithium-ion battery life. The reward function determines weight coefficients through grid search, balancing power fluctuation penalties, transaction costs, and degradation costs, guiding the agent to control economic losses while smoothing out fluctuations. The state transition function accurately describes the dynamic changes in the state of charge, ensuring that decisions conform to the physical laws of lithium-ion battery charging and discharging. The overall model design makes the decision-making of the lithium-ion battery energy storage agent more aligned with actual needs, improving its own operational safety and economy, and laying a precise control foundation for collaboration with flywheel energy storage, contributing to the stable and efficient operation of the overall system.
[0022] In one optional implementation, the components of the Markov decision process model for flywheel energy storage are:
[0023] The state space of flywheel energy storage includes electricity purchase price, photovoltaic power generation, wind power generation, user power consumption, power consumption of proton exchange membrane electrolyzer, current speed of flywheel, and cumulative fatigue coefficient of flywheel.
[0024] The operating range of flywheel energy storage: flywheel charging and discharging power;
[0025] The reward function for flywheel energy storage, considering power fluctuation penalties, system costs, and speed deviation penalties, is used to quantify the overall benefits of flywheel operation and is expressed as:
[0026] ,
[0027] Where F represents the flywheel, The instant reward for the flywheel's energy storage at time t. FC t The cost of flywheel energy storage systems, For lithium battery and flywheel multi-point energy storage stations and the main grid, time The power of the exchange; The average switching power within a preset time period; Let be the cumulative fatigue coefficient of the flywheel at time t. Let t be the actual rotational speed of the flywheel at time t. To preset the optimal speed, , , The weighting coefficients are determined using a grid search strategy.
[0028] The state transition function for flywheel energy storage, based on rotational inertia and efficiency calculations of speed change and fatigue accumulation, is expressed as:
[0029]
[0030]
[0031] in, Let t be the rotational speed of flywheel F; Let F be the rotational speed of the flywheel F at time t+1; The time interval between two adjacent decision moments; The moment of inertia of the flywheel; The charging and discharging efficiency of the flywheel; The charging power of the flywheel at time t. Let be the discharge power of the flywheel at time t. This represents the discharge power of the flywheel at time t. The positive and negative values of the charging and discharging power are ensured by taking the max function. This is the loss coefficient; Let be the change in power of the flywheel at time t; The cumulative fatigue coefficient of the flywheel at time t+1; Let t be the cumulative fatigue coefficient of the flywheel at time t.
[0032] The flywheel energy storage Markov decision process model designed in this embodiment incorporates rotational speed and cumulative fatigue coefficient into its state space, accurately capturing the core characteristics of the flywheel—"energy storage and release through rotational speed" and "frequent start-stop cycles leading to fatigue"—providing a basis for targeted control. The action space is designed around charging and discharging power to adapt to the flywheel's power regulation needs. A new speed deviation penalty is added to the reward function to guide the flywheel to operate within its high-efficiency range. Combining power fluctuation penalties with cost considerations, multi-objective optimization is achieved. The state transition function calculates rotational speed and fatigue changes based on physical parameters such as moment of inertia and efficiency, ensuring the model conforms to the flywheel's operating mechanism. This model allows the flywheel energy storage agent to fully leverage its rapid response advantage to mitigate short-term fluctuations and extend its lifespan through fatigue control. It complements the lithium battery energy storage model, significantly improving the adaptability and overall benefits of heterogeneous energy storage collaborative operation, and avoiding the limitations of single energy storage control.
[0033] In one optional implementation, training the algorithm framework using a centralized training and distributed execution mechanism includes:
[0034] An evaluation network is constructed and trained centrally using global state information, which is a set of the state spaces of lithium battery energy storage and flywheel energy storage, to comprehensively evaluate the collaborative decision-making effect of heterogeneous intelligent agents.
[0035] An action network is constructed, and distributed execution is performed using local state information. The local state information of the lithium battery energy storage agent is the state space of the lithium battery energy storage, and the local state information of the flywheel energy storage agent is the state space of the flywheel energy storage, which is used to reduce the communication overhead between agents.
[0036] The training model iteratively updates the evaluation network and action network to enable the model output to meet the action space constraints of lithium battery energy storage charging and discharging power and flywheel energy storage charging and discharging power, thereby realizing the aggregated and coordinated operation of multi-point energy storage of lithium battery and flywheel.
[0037] This invention employs a "centralized training - distributed execution" mechanism. The evaluation network is trained using global state information, comprehensively considering the overall operating state of the lithium battery and flywheel energy storage. This avoids decision-making bias caused by local information and ensures the global optimality of the collaborative strategy. The action network executes based on local state information, reducing data interaction between agents, lowering communication overhead, improving real-time response speed, and adapting to the control latency requirements of the energy storage system. By iteratively updating the two types of networks, the model output satisfies the charging and discharging power constraints of the action space. This ensures that the lithium battery and flywheel energy storage operate within a safe range while achieving dynamic collaboration between the two. The flywheel quickly smooths out short-term fluctuations, while the lithium battery undertakes medium- and long-term energy regulation. This effectively improves the operational stability, response timeliness, and control efficiency of the multi-point energy storage aggregation system, solving the problems of heavy communication burden in traditional centralized control and poor optimization in distributed control.
[0038] In one alternative implementation, during model training, the evaluation network is updated by minimizing a loss function, which is defined as:
[0039]
[0040] in, The instant reward is the sum of the instant rewards for lithium battery energy storage and flywheel energy storage. This is a discount factor used to weigh the importance of current rewards against future rewards; In the state The cooperative strategy output by the next action network. In the state Take action Expected returns;
[0041] The Action Network The policy gradient is defined as follows: (Updated via policy gradient)
[0042]
[0043] in, These are the parameters of the action network. This is a performance indicator of the strategy. S represents the batch size. B S represents the energy storage state of a lithium battery. F This refers to the state of energy storage in the flywheel.
[0044] The evaluation network in this embodiment of the invention updates by minimizing a loss function that includes the total immediate reward and a discount factor. The total immediate reward comprehensively reflects the combined benefits of lithium battery and flywheel energy storage, thus reflecting the coordinated operation effect. The discount factor flexibly balances current and future benefits, avoiding short-sighted decision-making, and making the evaluation network's estimation of action value more accurate, providing a reliable basis for action optimization. The action network is based on policy gradient updates, clearly defining the relationship between action network parameters, policy performance indicators, and batch size. Batch training ensures the stability of parameter updates. Combined with optimization strategies for lithium battery and flywheel energy storage states, the action output better matches the physical characteristics and coordinated needs of the two types of energy storage. This update mechanism ensures that the algorithm framework can be continuously iterated and optimized, allowing the agent to gradually learn charging and discharging strategies that balance dynamic response and economy, improving the model's convergence speed and final control effect, and avoiding policy failure or inefficient optimization caused by unreasonable network update logic.
[0045] In one optional implementation, the target constraints include system power balance constraints, lithium battery energy storage operation constraints, and flywheel energy storage operation constraints; wherein:
[0046] The system power balance constraint requires that the sum of photovoltaic power generation, wind power generation, lithium battery charging and discharging power, and flywheel charging and discharging power match the sum of user demand power and proton exchange membrane electrolyzer power consumption.
[0047] The lithium battery energy storage operation constraints include state of charge constraints and charge / discharge power constraints.
[0048] The flywheel energy storage operation constraints include speed constraints and cumulative fatigue coefficient constraints.
[0049] The system power balance constraints provided by this invention ensure precise matching between the power of wind and solar power generation, the two types of energy storage, and the power consumption of the electrolyzer. This guarantees the stability of the power system's supply and demand from a global perspective, avoiding voltage fluctuations and frequency deviations caused by power imbalances, and improving power quality. Lithium battery energy storage operation constraints limit the state of charge and charging / discharging power, preventing overcharging and over-discharging damage to the battery and extending its lifespan. Simultaneously, it ensures that the lithium battery's output power remains within a safe range, avoiding impact on the system. Flywheel energy storage operation constraints control the rotational speed and fatigue coefficient, ensuring the flywheel operates within a high-efficiency and safe range, reducing mechanical losses and extending its cycle life. These three types of constraints form a comprehensive management and control system. They not only define safe boundaries for heterogeneous multi-agent collaborative optimization, preventing decisions from exceeding equipment capacity, but also guide the optimization direction through constraints, ensuring that the final collaborative operation scheme achieves a balance between economy and performance under safe and stable conditions, improving system reliability and overall efficiency.
[0050] Secondly, the present invention provides a multi-point energy storage aggregation and collaborative operation device, the device comprising:
[0051] The heterogeneous model building module is used to construct Markov decision process models for lithium battery energy storage and flywheel energy storage based on their different physical characteristics in terms of cycle life, energy density and operating constraints.
[0052] The agent modeling module is used to construct a heterogeneous multi-agent deep deterministic policy gradient algorithm framework. Based on two heterogeneous Markov decision process models, it completes the modeling of lithium battery energy storage agents and flywheel energy storage agents, so that each agent is adapted to the physical characteristics of the corresponding energy storage.
[0053] The multi-agent collaborative optimization strategy output module is used to train the algorithm framework using a centralized training and distributed execution mechanism. After training, the algorithm model, under the objective constraints, solves the optimal collaborative operation scheme of multi-point energy storage aggregation of lithium batteries and flywheels with power distribution as the core and dynamic response performance and full life cycle economy through the heterogeneous multi-agent collaborative optimization strategy.
[0054] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the multi-point energy storage aggregation and cooperative operation method described in the first aspect or any corresponding embodiment.
[0055] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the multi-point energy storage aggregation and collaborative operation method described in the first aspect or any of its corresponding embodiments.
[0056] Fifthly, the present invention provides a computer program product, including computer instructions, which are used to cause a computer to execute the multi-point energy storage aggregation and collaborative operation method described in the first aspect or any of its corresponding embodiments. Attached Figure Description
[0057] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0058] Figure 1 This is a schematic diagram of a multi-energy coupled regional energy system according to an embodiment of the present invention;
[0059] Figure 2This is a flowchart illustrating a multi-point energy storage aggregation and collaborative operation method according to an embodiment of the present invention;
[0060] Figure 3 This is a block diagram of the heterogeneous multi-agent deep deterministic policy gradient algorithm according to an embodiment of the present invention;
[0061] Figure 4 This is a schematic diagram comparing the exchange power of the lithium battery energy storage and flywheel energy storage coordinated operation strategy with that of the unoptimized operation when the multi-point energy storage aggregation and coordinated operation method of this invention is used in the wind-solar-storage combined system.
[0062] Figure 5 This is a comparison chart of the power fluctuation control effects of different reinforcement learning algorithms in a multi-point energy storage aggregation and collaboration scenario according to embodiments of the present invention.
[0063] Figure 6 This is a comparison chart of the comprehensive costs of different reinforcement learning algorithms in a multi-point energy storage aggregation and collaboration scenario according to embodiments of the present invention;
[0064] Figure 7 This is a structural block diagram of a multi-point energy storage aggregation and collaborative operation device according to an embodiment of the present invention;
[0065] Figure 8 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0067] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0068] In an application scenario, such as Figure 1 The diagram shows a multi-energy coupled regional energy system. Its core is an "energy management system" that coordinates the production, storage, conversion, and consumption of various energy sources, achieving multi-energy complementarity of electricity, hydrogen, and heat. Specifically:
[0069] 1. The energy production side includes:
[0070] Photovoltaic power generation: Converting solar energy into electrical energy through photovoltaic arrays is one of the renewable energy inputs.
[0071] Wind power generation: converting wind energy into electrical energy through wind farms, together with photovoltaic power generation, constitutes the main body of clean power generation.
[0072] 2. The energy storage and conversion side includes:
[0073] Multi-point energy storage stations: (1) Flywheel system: Utilizes the high-speed rotation of the flywheel to store kinetic energy, with the characteristics of fast response speed and high power density, used to smooth short-term power fluctuations. (2) Lithium battery: Stores electrical energy with chemical energy, with high energy density, used for medium and long-term energy regulation and storage.
[0074] Proton exchange membrane electrolyzer: converts electrical energy into hydrogen energy (green arrow), realizing the conversion of electrical energy into chemical energy, and supplying hydrogen to hydrogen energy users.
[0075] 3. The energy consumption side includes:
[0076] Residential areas consume electricity and are one of the main sources of electricity load.
[0077] Hydrogen energy users: consume hydrogen produced by electrolyzers to realize the end-use of hydrogen energy.
[0078] Heat pump systems consume electrical energy and output heat energy (red arrow) to meet the district heating needs.
[0079] 4. Energy Management and Interaction:
[0080] Energy Management System: This is the "brain" of the entire system. It collects data from various stages (such as wind and solar power output, load demand, energy storage status, etc.) through information flow (black dotted line) and outputs control commands to coordinate the flow and distribution of electrical energy, hydrogen energy, and thermal energy.
[0081] Main power grid: As an external power interaction interface, it enables bidirectional flow of electricity between the regional system and the main power grid (surplus electricity to the grid or electricity purchase when there is a shortage).
[0082] 5. Multi-current coupling logic:
[0083] Electrical energy flow (blue arrow): It flows between wind and solar power generation, energy storage (flywheel, lithium battery), main power grid, load (residential area), electrolyzer, and heat pump, and is the core energy carrier of the system.
[0084] Hydrogen energy flow (green arrow): The electrolyzer converts electrical energy into hydrogen energy, which then flows to hydrogen energy users, realizing cross-form storage and utilization of electrical energy into hydrogen energy.
[0085] Heat energy flow (red arrow): The heat pump system converts electrical energy into heat energy to meet the district heating demand and realize the conversion of electrical energy into heat energy.
[0086] The aforementioned multi-energy coupled regional energy system architecture fully leverages the potential of renewable energy and enhances system flexibility and economy through the synergy of multiple energy sources and energy storage.
[0087] According to an embodiment of the present invention, a method for multi-point energy storage aggregation and collaborative operation is provided, which can be applied to the energy management system described above. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here. Figure 2 This is a flowchart of a multi-point energy storage aggregation and collaborative operation method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps:
[0088] Step S1: Based on the different physical characteristics of lithium battery energy storage and flywheel energy storage in terms of cycle life, energy density and operating constraints, respectively, construct Markov decision process models for lithium battery energy storage and flywheel energy storage.
[0089] Specifically, the components of the Markov decision process model for lithium battery energy storage are:
[0090]
[0091] in, The state space of a lithium battery energy storage system (i.e., the set of all characteristic information of the system at a certain moment). The action space for lithium battery energy storage (i.e., the set of all operations that the system can perform). The reward function for lithium battery energy storage (i.e., the quantitative indicator of "gain / cost" obtained after the system performs an action, used to guide strategy optimization in reinforcement learning). The state transition function for lithium battery energy storage (i.e., the dynamic law of the system transitioning from the current state to the next state after performing an action).
[0092] The Markov decision process model for lithium battery energy storage comprehensively captures the external environment (electricity price, wind and solar power, load) and internal state (state of charge) through the state space; combined with the action space, reward function and state transition function, it enables the lithium battery agent to learn the optimal strategy of "when to charge, when to discharge, how much to charge, and how much to discharge", and finally achieves aggregated and coordinated operation with flywheel energy storage.
[0093] Furthermore, the state space of lithium battery energy storage encompasses key influencing factors such as electricity purchase price and wind and solar power, ensuring that the model can comprehensively reflect actual operating scenarios and provide sufficient information support for intelligent agent decision-making. Specifically, this is represented as follows:
[0094]
[0095] in, The electricity purchase price at time t (reflects the economic incentives on the grid side and affects the economic decisions regarding energy storage charging and discharging). Let be the power generated by photovoltaic power generation and wind power generation at time t, respectively. The user demand at time t (system load demand, which determines the charging and discharging direction of energy storage). Let be the power consumed by the proton exchange membrane electrolyzer at time t. The state of charge (SOC) of the lithium battery at time t (i.e., the proportion of remaining charge to rated capacity, which is the core constraint for the safe operation of lithium batteries).
[0096] Lithium-ion battery energy storage operates within a limited charging and discharging power constraint to prevent overcharging and discharging from damaging the equipment and to extend the lifespan of the lithium-ion battery. Specifically:
[0097]
[0098] in, Let t be the charging and discharging power of the lithium battery at time t.
[0099] The reward function for lithium-ion battery energy storage considers power fluctuation penalties and system costs, and is used to quantify the overall benefits of lithium-ion battery operation, expressed as:
[0100]
[0101] Where B represents lithium battery, The instantaneous reward for lithium battery energy storage at time t; , These are weighting coefficients, determined using a grid search strategy, used to balance the priorities of different optimization objectives; For lithium battery and flywheel multi-point energy storage stations and the main grid, time The power of the exchange; The average switching power within a preset time period; For system transaction costs; This refers to the degradation cost of lithium batteries.
[0102] The state transition function of a lithium-ion battery describes the dynamic process of the state of charge changing with charging and discharging power, ensuring that the decision-making conforms to the physical laws of lithium-ion battery charging and discharging, and is expressed as:
[0103]
[0104] in, This represents the state of charge of the lithium battery at time t; This indicates the state of charge of the lithium battery at time t+1; Improve lithium battery charging efficiency; For lithium battery discharge efficiency; For lithium battery capacity; The time interval between two adjacent decision moments; The charging power of the lithium battery; This refers to the discharge power of the lithium battery.
[0105] The components of the Markov decision process model for flywheel energy storage constructed in this embodiment of the invention are as follows:
[0106]
[0107] in, The state space for storing energy in the flywheel (i.e., the set of all characteristic information of the flywheel at a certain moment). The action space for storing energy in the flywheel (i.e., the set of all operations that the flywheel can perform). The reward function for energy storage in the flywheel (i.e., the quantitative indicator of "gain / cost" obtained after the flywheel performs an action, used to guide strategy optimization in reinforcement learning). The state transition function for energy storage in the flywheel (i.e., the dynamic law of the flywheel transitioning from the current state to the next state after performing an action).
[0108] Furthermore, the flywheel energy storage state space is:
[0109]
[0110] in, The electricity purchase price at time t (reflects the economic incentives on the grid side and affects the economic decisions regarding energy storage charging and discharging). Let be the power generated by photovoltaic power generation and wind power generation at time t, respectively. The user demand at time t (system load demand, which determines the charging and discharging direction of energy storage). Let be the power consumed by the proton exchange membrane electrolyzer at time t. Let t be the current rotational speed of the flywheel. The cumulative fatigue coefficient of the flywheel at time t reflects the losses from frequent start-stop cycles. By incorporating rotational speed and the cumulative fatigue coefficient into the state space, the core characteristics of the flywheel—its ability to store and release energy at rotational speed and its susceptibility to fatigue due to frequent start-stop cycles—are accurately captured, providing a basis for targeted control.
[0111] The motion space is: ,in, This represents the charging and discharging power of the flywheel at time t (positive numbers indicate discharging, negative numbers indicate charging; this is the direct interaction between the flywheel and the power grid). This action space is designed around the charging and discharging power to accommodate the flywheel's power regulation needs.
[0112] The reward function for flywheel energy storage, considering power fluctuation penalties, system costs, and speed deviation penalties, is used to quantify the overall benefits of flywheel operation and is expressed as:
[0113]
[0114] Where F represents the flywheel, The instant reward for the flywheel's energy storage at time t. FC t The cost of flywheel energy storage systems, For lithium battery and flywheel multi-point energy storage stations and the main grid, time The power of the exchange; The average switching power within a preset time period; Let be the cumulative fatigue coefficient of the flywheel at time t. Let t be the actual rotational speed of the flywheel at time t. To preset the optimal speed, , , The weighting coefficients, determined through a grid search strategy, aim to map different physical quantities and objectives of varying importance onto a single reward scalar. This scalar balances objectives related to power stability, economy, lifespan loss, and rotational efficiency. Its magnitude directly reflects the priority of the corresponding objective within the overall optimization objective, enabling the agent to distinguish the merits of different actions during exploration and learning, thus effectively learning the strategy. The reward function incorporates a rotational speed deviation penalty to guide the flywheel to operate within its efficient range. Combined with power fluctuation penalties and cost considerations, multi-objective optimization is achieved.
[0115] The state transition function for flywheel energy storage, based on rotational inertia and efficiency calculations of speed change and fatigue accumulation, is expressed as:
[0116]
[0117]
[0118] in, Let t be the rotational speed of the flywheel at time t; Let t+1 be the rotational speed of the flywheel. The time interval between two adjacent decision moments; The moment of inertia of the flywheel; The charging and discharging efficiency of the flywheel; The charging power of the flywheel at time t. Let be the discharge power of the flywheel at time t. This represents the discharge power of the flywheel at time t. The positive and negative values of the charging and discharging power are ensured by taking the max function. This is the loss coefficient; Let be the change in power of the flywheel at time t; The cumulative fatigue coefficient of the flywheel at time t+1; Let be the cumulative fatigue coefficient of the flywheel at time t. This state transition function calculates the changes in rotational speed and fatigue based on physical parameters such as moment of inertia and efficiency, ensuring the model conforms to the flywheel's operating mechanism. This model allows the flywheel energy storage agent to fully leverage its rapid response advantage to mitigate short-term fluctuations, while also extending its lifespan through fatigue control. It complements the lithium battery energy storage model, significantly improving the adaptability and overall benefits of heterogeneous energy storage collaborative operation, and avoiding the limitations of single energy storage control.
[0119] The Markov decision process model for flywheel energy storage in this invention comprehensively captures the external environment (electricity price, wind and solar power, load) and internal state (speed, fatigue coefficient) through the state space.
[0120] By combining action space, reward function, and state transition function, the flywheel agent can learn the optimal strategy of "when to charge, when to discharge, and how much power to charge or discharge," ultimately achieving coordinated operation with lithium battery energy storage, leveraging its characteristics of "fast response speed and high power density," while controlling fatigue loss to ensure economic efficiency throughout the entire life cycle.
[0121] Step S2: Construct a heterogeneous multi-agent deep deterministic policy gradient algorithm framework. Based on two heterogeneous Markov decision process models, complete the modeling of lithium battery energy storage agents and flywheel energy storage agents, so that each agent is adapted to the physical characteristics of the corresponding energy storage.
[0122] Specifically, based on the two heterogeneous Markov decision process models constructed in step S1, a heterogeneous multi-agent deep deterministic policy gradient algorithm framework is built to achieve the aggregated and collaborative operation of heterogeneous energy storage such as lithium batteries and flywheels. The lithium battery energy storage agent, based on its own model, focuses on the long-term storage and release of energy to match the energy fluctuations of renewable energy; the flywheel energy storage agent, relying on its own model, focuses on quickly responding to short-term power fluctuations while minimizing its own fatigue losses.
[0123] like Figure 3 As shown, the environment module includes external factors such as user load, photovoltaic power generation, and wind power generation, which are the input conditions affecting energy storage decisions. The environment module outputs the current state s. t The agent performs action a t The reward r t and the next state s t+1 Forming empirical tuples The experience replay pool stores the aforementioned experience tuples for random sampling during algorithm training, breaking data correlation and improving training stability. The distributed execution component handles the independent decision-making of each energy storage agent, while centralized training optimizes the global strategy. Based on the actions and states of each agent, the evaluation network calculates the value of actions and updates parameters. The architecture adopts a heterogeneous multi-agent paradigm of "distributed execution + centralized training." Distributed execution allows each energy storage agent to make rapid decisions based on local states (e.g., lithium batteries focus on long-term energy regulation, while flywheels focus on short-term power stabilization), reducing communication overhead. Centralized training allows the evaluation network to optimize strategies based on global information, ensuring the global optimality of multi-energy storage collaboration. Ultimately, it achieves complementary advantages between lithium batteries and flywheels, stabilizing wind and solar power fluctuations, ensuring power balance, and controlling equipment losses and operating costs. It is a highly efficient algorithm framework for multi-point heterogeneous energy storage aggregation and collaboration.
[0124] Step S3 involves training the algorithm framework using a centralized training and distributed execution mechanism. The trained algorithm model, under objective constraints, employs a heterogeneous multi-agent collaborative optimization strategy to find the optimal collaborative operation scheme for multi-point energy storage aggregation of lithium batteries and flywheels, centered on power allocation and considering both dynamic response performance and overall lifecycle economics. Specifically, this includes the following steps:
[0125] Step S31: Construct an evaluation network and perform centralized training using global state information. The global state information is a set of the state space of lithium battery energy storage and the state space of flywheel energy storage, which is used to comprehensively evaluate the collaborative decision-making effect of heterogeneous intelligent agents.
[0126] Step S32: Construct an action network and use local state information for distributed execution. The local state information of the lithium battery energy storage agent is the state space of the lithium battery energy storage, and the local state information of the flywheel energy storage agent is the state space of the flywheel energy storage, which is used to reduce the communication overhead between agents.
[0127] Step S33: Train the model. By iteratively updating the evaluation network and action network, the model outputs lithium battery energy storage charging and discharging power and flywheel energy storage charging and discharging power that satisfy the action space constraints, thereby realizing the aggregated and coordinated operation of multi-point energy storage of lithium battery and flywheel.
[0128] This invention employs a "centralized training - distributed execution" mechanism. The evaluation network is trained using global state information, comprehensively considering the overall operating state of the lithium battery and flywheel energy storage. This avoids decision-making bias caused by local information and ensures the global optimality of the collaborative strategy. The action network executes based on local state information, reducing data interaction between agents, lowering communication overhead, improving real-time response speed, and adapting to the control latency requirements of the energy storage system. By iteratively updating the two types of networks, the model output satisfies the charging and discharging power constraints of the action space. This ensures that the lithium battery and flywheel energy storage operate within a safe range while achieving dynamic collaboration between the two—the flywheel quickly smooths short-term fluctuations, and the lithium battery undertakes medium- and long-term energy regulation. This effectively improves the operational stability, response timeliness, and control efficiency of the multi-point energy storage aggregation system, solving the problems of heavy communication burden in traditional centralized control and poor optimization in distributed control.
[0129] Furthermore, the evaluation network is updated by minimizing the loss function, which is defined as:
[0130]
[0131] in, The instant reward is the sum of the instant rewards for lithium battery energy storage and flywheel energy storage. This is a discount factor used to weigh the importance of current rewards against future rewards; In the state The cooperative strategy output by the next action network. In the state Take action The expected return.
[0132] The Action Network The policy gradient is defined as follows: (Updated via policy gradient)
[0133]
[0134] in, These are the parameters of the action network. This is a performance indicator of the strategy. S represents the batch size. B S represents the energy storage state of a lithium battery. F This refers to the state of energy storage in the flywheel.
[0135] The evaluation network provided in this invention updates by minimizing a loss function that includes total immediate reward and a discount factor. The total immediate reward comprehensively reflects the combined benefits of lithium battery and flywheel energy storage, thus reflecting the coordinated operation effect. The discount factor flexibly balances current and future benefits, avoiding short-sighted decision-making and allowing the evaluation network to more accurately estimate the value of actions, providing a reliable basis for action optimization. The action network is based on policy gradient updates, clearly defining the relationship between action network parameters, policy performance indicators, and batch size. Batch training ensures the stability of parameter updates. Combined with optimization strategies for lithium battery and flywheel energy storage states, the action output better matches the physical characteristics and coordinated needs of the two types of energy storage. This update mechanism ensures that the algorithm framework can be continuously iterated and optimized, allowing the agent to gradually learn charging and discharging strategies that balance dynamic response and economy, improving model convergence speed and final control effect, and avoiding policy failure or inefficient optimization caused by unreasonable network update logic.
[0136] The trained algorithm model in this embodiment of the invention utilizes a heterogeneous multi-agent collaborative optimization strategy under objective constraints. These objective constraints include system power balance constraints, lithium battery energy storage operation constraints, and flywheel energy storage operation constraints; specifically:
[0137] 1. The system power balance constraint requires that the sum of photovoltaic power generation, wind power generation, lithium battery charging and discharging power, and flywheel charging and discharging power match the sum of the user's power demand and the proton exchange membrane electrolyzer's power consumption.
[0138] In one example, suppose at a certain moment the photovoltaic power generation is 80kW, the wind power generation is 30kW, the lithium battery discharge power is 20kW (discharge is positive, charging is negative), and the flywheel discharge power is 10kW; the user demand power is 85kW, and the proton exchange membrane electrolyzer consumes 5kW. Then the total power of power generation and energy storage discharge on the left is 80+30+20+10=110kW, and the total power of user demand and electrolyzer consumption on the right is 85+5=90kW. At this point, the power balance constraint is not met. The algorithm will adjust the charging and discharging power of the lithium battery and flywheel, for example, charging the lithium battery to 10kW (i.e., power of -10kW) and discharging the flywheel to 5kW. At this time, the total on the left is 80+30-10+5=75kW, which is still mismatched. The adjustment continues until the power on both sides is equal, ensuring the stability of the power system.
[0139] 2. Lithium-ion battery energy storage operation constraints include state-of-charge constraints and charge / discharge power constraints. In one example:
[0140] State of Charge (SOC) constraint: The SOC of the lithium battery must be maintained between 20% and 80%. If the current SOC is 15%, the algorithm will control the lithium battery to stop discharging and switch to charging, so that the SOC rises to a safe range; if the SOC is 85%, the algorithm will control the lithium battery to stop charging and switch to discharging or standby, so that the SOC drops to a safe range.
[0141] Charge / discharge power constraints: The maximum charging power of the lithium battery is 30kW, and the maximum discharging power is 40kW. When the algorithm calculates that the lithium battery needs to be charged at 80kW, the charging power will be limited to 30kW; if it needs to be discharged at 80kW, it will be limited to 40kW to prevent the lithium battery from being damaged by overcharging or over-discharging.
[0142] 3. Flywheel energy storage operation constraints include speed constraints and cumulative fatigue coefficient constraints. In one example:
[0143] Speed constraint: The flywheel speed must be between 10,000 r / min and 30,000 r / min. If the current speed is 8,000 r / min, the algorithm will control the flywheel to stop discharging and start charging, so that the speed rises to a safe range; if the speed is 32,000 r / min, the flywheel will be controlled to discharge, so that the speed drops to a safe range.
[0144] Cumulative fatigue coefficient constraint: The cumulative fatigue coefficient of the flywheel cannot exceed 1000 (assuming a threshold of 1000). When the cumulative fatigue coefficient approaches 1000, the algorithm will reduce the number of flywheel start-stop cycles and the magnitude of power changes. For example, if the flywheel originally needed to respond quickly to a 20kW power fluctuation, it will now be adjusted to allow the flywheel to respond to 15kW, while the lithium battery assists in responding to the remaining 5kW, in order to slow down the accumulation of flywheel fatigue.
[0145] This invention, through system power balance constraints, ensures precise matching between the power of wind and solar power generation, the two types of energy storage, and the power consumption of the electrolyzer. This globally guarantees the stability of the power system's supply and demand, avoiding voltage fluctuations and frequency deviations caused by power imbalances, and improving power quality. Lithium battery energy storage operation constraints limit the state of charge and charging / discharging power, preventing overcharging and over-discharging damage to the battery and extending its lifespan. Simultaneously, it ensures that the lithium battery's output power remains within a safe range, avoiding impact on the system. Flywheel energy storage operation constraints control the rotational speed and fatigue coefficient, ensuring the flywheel operates within a high-efficiency and safe range, reducing mechanical losses and extending its cycle life. These three types of constraints form a comprehensive management and control system. This system not only defines safety boundaries for heterogeneous multi-agent collaborative optimization, preventing decisions from exceeding equipment capacity, but also guides the optimization direction through constraints, ensuring that the final collaborative operation scheme achieves a balance between economy and performance under safe and stable conditions, improving system reliability and overall efficiency.
[0146] In one application scenario, a wind-solar-storage integrated system is operating. The synergistic control strategy for lithium battery energy storage and flywheel energy storage obtained using the method provided in this embodiment of the invention is as follows: Figure 4 As shown in the figure. The results show that the method proposed in this invention can achieve power balance at every time step: in Scheme 1, the flywheel energy storage performs rapid response charging or discharging, which is suitable for times when the supply and demand difference is minimal and fluctuations are rapid; Scheme 2 only activates lithium battery energy storage to achieve long-term energy balance when supply is insufficient; in Scheme 3, lithium battery and flywheel energy storage operate simultaneously, flexibly meeting the requirements of rapid response and long-term balance. By comparing the yellow curve of "exchange power without optimization", the power exchange fluctuation between the optimized system and the main grid is significantly reduced, reducing the frequency and scale of users purchasing electricity from the grid, which not only improves the utilization rate of renewable energy, but also enhances the stability and economy of the power system.
[0147] Figure 5 shows a comparison of the power fluctuation control effects of different reinforcement learning algorithms in a multi-point energy storage aggregation and collaboration scenario. This figure intuitively demonstrates the technical advantages of the algorithm of this invention in a multi-point energy storage aggregation and collaboration scenario: compared with mainstream reinforcement learning algorithms such as TD3, DDPG, and PPO, the algorithm of this invention can reduce system power fluctuations more quickly and thoroughly, ultimately achieving the lowest and most stable power fluctuation level. This result verifies the effectiveness of the design idea based on heterogeneous multi-agent deep deterministic policy gradient in adapting to the differentiated characteristics of lithium batteries and flywheel energy storage and achieving collaborative optimization, providing a better technical solution for smoothing renewable energy fluctuations and improving the stability of the power system.
[0148] Figure 6 The figure shown is a comparison of the overall costs of different reinforcement learning algorithms in a multi-point energy storage aggregation and collaboration scenario. This figure intuitively demonstrates the economic advantages of the algorithm of this invention in a multi-point energy storage aggregation and collaboration scenario: compared with mainstream reinforcement learning algorithms such as TD3, DDPG, and PPO, the algorithm of this invention can significantly reduce the overall system cost and has the best cost stability. This result verifies the effectiveness of the algorithm in balancing "dynamic response performance" and "full life cycle economy". It reduces equipment losses through precise power allocation and reduces transaction costs through optimized grid interaction, providing a technical solution that combines performance and economy for multi-point energy storage aggregation operation.
[0149] This embodiment also provides a multi-point energy storage aggregation and collaborative operation device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0150] This embodiment provides a multi-point energy storage aggregation and collaborative operation device, such as... Figure 7 As shown, it includes:
[0151] The heterogeneous model construction module 71 is used to construct Markov decision process models for lithium battery energy storage and flywheel energy storage based on the differentiated physical characteristics of lithium battery energy storage and flywheel energy storage in terms of cycle life, energy density and operating constraints, respectively.
[0152] Agent Modeling 72 is used to construct a heterogeneous multi-agent deep deterministic policy gradient algorithm framework. Based on two heterogeneous Markov decision process models, it completes the modeling of lithium battery energy storage agents and flywheel energy storage agents, so that each agent is adapted to the physical characteristics of the corresponding energy storage.
[0153] The multi-agent collaborative optimization strategy output module 73 is used to train the algorithm framework using a centralized training and distributed execution mechanism. After training, the algorithm model, under the objective constraint, solves the optimal collaborative operation scheme of lithium battery and flywheel multi-point energy storage aggregation with power distribution as the core and dynamic response performance and full life cycle economy through the heterogeneous multi-agent collaborative optimization strategy.
[0154] In some alternative implementations, the components of the Markov decision process model for lithium battery energy storage in heterogeneous model construction mode 71 include:
[0155] The state space of lithium battery energy storage includes electricity purchase price, photovoltaic power generation, wind power generation, user demand power, proton exchange membrane electrolyzer power consumption, and lithium battery state of charge.
[0156] The operating space of lithium battery energy storage: the charging and discharging power of the lithium battery, and respectively satisfying the maximum charging / discharging power constraints of the lithium battery;
[0157] The reward function for lithium-ion battery energy storage, considering power fluctuation penalties and system costs, is used to quantify the overall benefits of lithium-ion battery operation and is expressed as:
[0158]
[0159] Where B represents lithium battery, The instant reward for lithium battery energy storage at time t; , These are weighting coefficients, determined using a grid search strategy, used to balance the priorities of different optimization objectives; For lithium battery and flywheel multi-point energy storage stations and the main grid, time The power of the exchange; The average switching power within a preset time period; For system transaction costs; The degradation cost of lithium batteries;
[0160] The state transition function of a lithium-ion battery describes the dynamic process of the state of charge changing with charging and discharging power, and is expressed as:
[0161]
[0162] in, This represents the state of charge of the lithium battery at time t; This indicates the state of charge of the lithium battery at time t+1; Improve lithium battery charging efficiency; For lithium battery discharge efficiency; For lithium battery capacity; The time interval between two adjacent decision moments; The charging power of the lithium battery; This refers to the discharge power of the lithium battery.
[0163] In some optional implementations, the components of the Markov decision process model for flywheel energy storage in heterogeneous model construction module 71 are:
[0164] The state space of flywheel energy storage includes electricity purchase price, photovoltaic power generation, wind power generation, user power consumption, power consumption of proton exchange membrane electrolyzer, current speed of flywheel, and cumulative fatigue coefficient of flywheel.
[0165] The operating range of flywheel energy storage: flywheel charging and discharging power;
[0166] The reward function for flywheel energy storage, considering power fluctuation penalties, system costs, and speed deviation penalties, is used to quantify the overall benefits of flywheel operation and is expressed as:
[0167]
[0168] Where F represents the flywheel, The instant reward for the flywheel's energy storage at time t. FC t The cost of flywheel energy storage systems, For lithium battery and flywheel multi-point energy storage stations and the main grid, time The power of the exchange; The average switching power within a preset time period; Let be the cumulative fatigue coefficient of the flywheel at time t. Let t be the actual rotational speed of the flywheel at time t. To preset the optimal speed, , , The weighting coefficients are determined using a grid search strategy.
[0169] The state transition function for flywheel energy storage, based on rotational inertia and efficiency calculations of speed change and fatigue accumulation, is expressed as:
[0170]
[0171]
[0172] in, Let t be the rotational speed of the flywheel at time t; Let t+1 be the rotational speed of the flywheel. The time interval between two adjacent decision moments; The moment of inertia of the flywheel. The charging and discharging efficiency of the flywheel. The charging power of the flywheel at time t. Let be the discharge power of the flywheel at time t. This represents the discharge power of the flywheel at time t. The maximum value is used to ensure the reasonableness of the positive and negative signs of the charging and discharging power. The loss coefficient is... Let be the change in power of the flywheel at time t; The cumulative fatigue coefficient of the flywheel at time t+1; Let t be the cumulative fatigue coefficient of the flywheel at time t.
[0173] In some optional implementations, the multi-agent cooperative optimization strategy output module 73 employs a centralized training and distributed execution mechanism to train the algorithm framework, including:
[0174] An evaluation network is constructed and trained centrally using global state information, which is a set of the state spaces of lithium battery energy storage and flywheel energy storage, to comprehensively evaluate the collaborative decision-making effect of heterogeneous intelligent agents.
[0175] An action network is constructed, and distributed execution is performed using local state information. The local state information of the lithium battery energy storage agent is the state space of the lithium battery energy storage, and the local state information of the flywheel energy storage agent is the state space of the flywheel energy storage, which is used to reduce the communication overhead between agents.
[0176] The training model iteratively updates the evaluation network and action network to enable the model output to meet the action space constraints of lithium battery energy storage charging and discharging power and flywheel energy storage charging and discharging power, thereby realizing the aggregated and coordinated operation of multi-point energy storage of lithium battery and flywheel.
[0177] In some alternative implementations, during model training, the evaluation network is updated by minimizing a loss function, which is defined as:
[0178]
[0179] in, The instant reward is the sum of the instant rewards for lithium battery energy storage and flywheel energy storage. This is a discount factor used to weigh the importance of current rewards against future rewards; In the state The cooperative strategy output by the next action network. In the state Take action Expected returns;
[0180] The Action Network The policy gradient is defined as follows: (Updated via policy gradient)
[0181]
[0182] in, These are the parameters of the action network. This is a performance indicator of the strategy. S represents the batch size. B S represents the energy storage state of a lithium battery. F This refers to the state of energy storage in the flywheel.
[0183] In some optional implementations, the target constraints in the multi-agent cooperative optimization strategy output module 73 include system power balance constraints, lithium battery energy storage operation constraints, and flywheel energy storage operation constraints; wherein:
[0184] The system power balance constraint requires that the sum of photovoltaic power generation, wind power generation, lithium battery charging and discharging power, and flywheel charging and discharging power match the sum of user demand power and proton exchange membrane electrolyzer power consumption; the lithium battery energy storage operation constraint includes state of charge constraint and charging and discharging power constraint; the flywheel energy storage operation constraint includes speed constraint and cumulative fatigue coefficient constraint.
[0185] The multi-point energy storage aggregation and collaborative operation device provided in this embodiment of the invention can execute the multi-point energy storage aggregation and collaborative operation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0186] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0187] The following is a detailed reference. Figure 8This diagram illustrates a suitable structural schematic for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 801, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 802 or a program loaded from memory 808 into random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device. The processor 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0188] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0189] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a memory 808, or installed from a ROM 802. When the computer program is executed by the processor 801, it performs the functions defined in the multi-point energy storage aggregation and cooperative operation method of the embodiments of the present invention.
[0190] Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0191] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the multi-point energy storage aggregation and collaborative operation method shown in the above embodiments is implemented.
[0192] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0193] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for multi-point energy storage aggregation and collaborative operation, characterized in that, include: Based on the differentiated physical characteristics of lithium-ion battery energy storage and flywheel energy storage in terms of cycle life, energy density, and operational constraints, Markov decision process models for lithium-ion battery energy storage and flywheel energy storage are constructed respectively. The components of the Markov decision process model for lithium-ion battery energy storage include: The reward function for lithium-ion battery energy storage, considering power fluctuation penalties and system costs, is used to quantify the overall benefits of lithium-ion battery operation and is expressed as: Where B represents lithium battery, The instantaneous reward for lithium battery energy storage at time t; , These are weighting coefficients, determined using a grid search strategy, used to balance the priorities of different optimization objectives; For lithium battery and flywheel multi-point energy storage stations and the main grid, time The power of the exchange; The average switching power within a preset time period; For system transaction costs; The degradation cost of lithium batteries; The components of the Markov decision process model for flywheel energy storage include: The reward function for flywheel energy storage, considering power fluctuation penalties, system costs, and speed deviation penalties, is used to quantify the overall benefits of flywheel operation and is expressed as: Where F represents the flywheel, The instant reward for the flywheel's energy storage at time t. FC t The cost of flywheel energy storage systems, Let be the cumulative fatigue coefficient of the flywheel at time t. Let t be the actual rotational speed of the flywheel at time t. To preset the optimal speed, , , The weighting coefficients are determined using a grid search strategy. A heterogeneous multi-agent deep deterministic policy gradient algorithm framework is constructed. Based on two heterogeneous Markov decision process models, the lithium battery energy storage agent and the flywheel energy storage agent are modeled so that each agent is adapted to the physical characteristics of the corresponding energy storage. The algorithm framework is trained using a centralized training and distributed execution mechanism. After training, the algorithm model, under the objective constraints, solves the optimal collaborative operation scheme of multi-point energy storage aggregation of lithium batteries and flywheels through a heterogeneous multi-agent collaborative optimization strategy. This scheme takes power allocation as the core and considers dynamic response performance and full life cycle economy.
2. The method according to claim 1, characterized in that, The components of the Markov decision process model for lithium battery energy storage include: The state space of lithium battery energy storage includes electricity purchase price, photovoltaic power generation, wind power generation, user demand power, proton exchange membrane electrolyzer power consumption, and lithium battery state of charge. The operating range of lithium battery energy storage: the charging and discharging power of the lithium battery, and each of them meets the maximum charging and discharging power constraints of the lithium battery; The state transition function of a lithium-ion battery describes the dynamic process of the state of charge changing with charging and discharging power, and is expressed as: in, This represents the state of charge of the lithium battery at time t; This indicates the state of charge of the lithium battery at time t+1; Improve lithium battery charging efficiency; For lithium battery discharge efficiency; For lithium battery capacity; The time interval between two adjacent decision moments; The charging power of the lithium battery; This refers to the discharge power of the lithium battery.
3. The method according to claim 1 or 2, characterized in that, The components of the Markov decision process model for flywheel energy storage also include: The state space of flywheel energy storage includes electricity purchase price, photovoltaic power generation, wind power generation, user power consumption, power consumption of proton exchange membrane electrolyzer, current speed of flywheel, and cumulative fatigue coefficient of flywheel. The operating range of flywheel energy storage: flywheel charging and discharging power; The state transition function for flywheel energy storage, based on rotational inertia and efficiency calculations of speed change and fatigue accumulation, is expressed as: in, Let t be the rotational speed of the flywheel at time t; Let t+1 be the rotational speed of the flywheel. The time interval between two adjacent decision moments; The moment of inertia of the flywheel; The charging and discharging efficiency of the flywheel; The charging power of the flywheel at time t. Let be the discharge power of the flywheel at time t. For flywheel energy storage intelligent agents in The action values output at any time are used to ensure the reasonableness of the positive and negative values of charging and discharging power by taking the max function; This is the loss coefficient; Let be the change in power of the flywheel at time t; The cumulative fatigue coefficient of the flywheel at time t+1; Let t be the cumulative fatigue coefficient of the flywheel at time t.
4. The method according to claim 1, characterized in that, The training of the algorithm framework using a centralized training and distributed execution mechanism includes: An evaluation network is constructed and trained centrally using global state information, which is a set of the state spaces of lithium battery energy storage and flywheel energy storage, to comprehensively evaluate the collaborative decision-making effect of heterogeneous intelligent agents. An action network is constructed, and distributed execution is performed using local state information. The local state information of the lithium battery energy storage agent is the state space of the lithium battery energy storage, and the local state information of the flywheel energy storage agent is the state space of the flywheel energy storage, which is used to reduce the communication overhead between agents. The training model iteratively updates the evaluation network and action network to enable the model output to meet the action space constraints of lithium battery energy storage charging and discharging power and flywheel energy storage charging and discharging power, thereby realizing the aggregated and coordinated operation of multi-point energy storage of lithium battery and flywheel.
5. The method according to claim 4, characterized in that, During model training, the evaluation network is updated by minimizing a loss function, which is defined as: in, This is the sum of the instant rewards for lithium battery energy storage and flywheel energy storage. This is a discount factor used to weigh the importance of current rewards against future rewards; In the state The cooperative strategy output by the next action network. In the state Take action Expected returns; The Action Network The policy gradient is defined as follows: (Updated via policy gradient) in, These are the parameters of the action network. This is a performance indicator of the strategy. S represents the batch size. B S represents the energy storage state of a lithium battery. F The state of energy storage for the flywheel.
6. The method according to claim 3, characterized in that, The target constraints include system power balance constraints, lithium battery energy storage operation constraints, and flywheel energy storage operation constraints; wherein: The system power balance constraint requires that the sum of photovoltaic power generation, wind power generation, lithium battery charging and discharging power, and flywheel charging and discharging power match the sum of user demand power and proton exchange membrane electrolyzer power consumption. The lithium battery energy storage operation constraints include state of charge constraints and charge / discharge power constraints. The flywheel energy storage operation constraints include speed constraints and cumulative fatigue coefficient constraints.
7. A multi-point energy storage aggregation and collaborative operation device, characterized in that, The device includes: A heterogeneous model construction module is used to construct Markov decision process models for lithium battery energy storage and flywheel energy storage, respectively, based on the differentiated physical characteristics of lithium battery energy storage and flywheel energy storage in terms of cycle life, energy density, and operational constraints. The components of the Markov decision process model for lithium battery energy storage include: The reward function for lithium-ion battery energy storage, considering power fluctuation penalties and system costs, is used to quantify the overall benefits of lithium-ion battery operation and is expressed as: Where B represents lithium battery, The instantaneous reward for lithium battery energy storage at time t; , These are weighting coefficients, determined using a grid search strategy, used to balance the priorities of different optimization objectives; For lithium battery and flywheel multi-point energy storage stations and the main grid, time The power of the exchange; The average switching power within a preset time period; For system transaction costs; The degradation cost of lithium batteries; The components of the Markov decision process model for flywheel energy storage include: The reward function for flywheel energy storage, considering power fluctuation penalties, system costs, and speed deviation penalties, is used to quantify the overall benefits of flywheel operation and is expressed as: Where F represents the flywheel, The instant reward for the flywheel's energy storage at time t. FC t The cost of flywheel energy storage systems, Let be the cumulative fatigue coefficient of the flywheel at time t. Let t be the actual rotational speed of the flywheel at time t. To preset the optimal speed, , , The weighting coefficients are determined using a grid search strategy. The agent modeling module is used to construct a heterogeneous multi-agent deep deterministic policy gradient algorithm framework. Based on two heterogeneous Markov decision process models, it completes the modeling of lithium battery energy storage agents and flywheel energy storage agents, so that each agent is adapted to the physical characteristics of the corresponding energy storage. The multi-agent collaborative optimization strategy output module is used to train the algorithm framework using a centralized training and distributed execution mechanism. After training, the algorithm model, under the objective constraints, solves the optimal collaborative operation scheme of multi-point energy storage aggregation of lithium batteries and flywheels with power distribution as the core and dynamic response performance and full life cycle economy through the heterogeneous multi-agent collaborative optimization strategy.
8. An electronic device, characterized in that, include: The system includes a memory and a processor, which are interconnected and the memory stores computer instructions. The processor executes the computer instructions to perform the multi-point energy storage aggregation and collaborative operation method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the multi-point energy storage aggregation and collaborative operation method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes computer instructions for causing a computer to execute the multi-point energy storage aggregation and collaborative operation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Distribution type energy management method for hybrid power rubber-tyred container cranes
CN110294418A
Hybrid energy storage system capacity and power configuration method and system
CN116094000A
Distributed energy system cluster collaborative optimization method based on multi-agent reinforcement learning
CN117350423A
Energy storage coordination regulation and control method and device
CN121124140A