Receiving end urban power grid energy storage multi-time scale optimization regulation and control system and method
The energy storage multi-timescale optimization and control system, which utilizes a hierarchical architecture and multi-agent deep reinforcement learning, solves the control problem of energy storage in urban power grids under high uncertainty. It achieves multi-objective optimization of safety, economy and low carbon emissions, and ensures stable and efficient operation of the system under extreme disturbances.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
- Filing Date
- 2025-12-08
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies cannot effectively regulate energy storage in urban power grids under high uncertainty, and it is difficult to maintain energy storage safety while maximizing economic benefits and carbon emission reduction under real-time extreme disturbances.
The energy storage multi-timescale optimization and control system adopts a hierarchical architecture, including a daytime coordination layer, a real-time execution layer, and a multi-agent collaboration layer. It utilizes multi-agent deep reinforcement learning to control multiple energy storage devices, and combines multi-objective planning and real-time constraint optimization to achieve optimization and control driven by multiple objectives such as safety, economy, and low carbon.
It enables the optimization of energy storage system revenue and emission reduction in a carbon-electricity-green certificate coupled market, ensuring the safe and stable operation of the system under extreme disturbances, and maximizing economic benefits and carbon emission reduction.
Smart Images

Figure CN121984045A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of energy storage regulation technology, and in particular to a multi-timescale optimization regulation system and method for energy storage in urban power grids at the receiving end. Background Technology
[0002] In the process of low-carbon transformation of the new power system, the receiving-end urban power grid faces problems such as a high proportion of new energy and external power, rapid and fluctuating load growth, and limited adjustable load resources. Coupled with the risks of UHVDC blocking and the demand for new energy consumption, the system operation scenario is complex. The receiving-end urban power grid has prominent overvoltage and oscillation risks. Rotational inertia is gradually becoming a scarce resource, and frequency stability margin is constantly decreasing. The resulting voltage-frequency safety issues have become a new challenge for the low-carbon transformation of the receiving-end urban power grid.
[0003] In recent years, the rapid development of new energy sources with strong randomness, volatility, and intermittency has limited the output of traditional power units, represented by coal-fired power plants. The use of energy storage resources, which possess low-carbon attributes, a wide adjustment range, fast adjustment speed, and long duration, to participate in power system balancing has become an inevitable trend. Currently, there is a large body of research on the optimal allocation and control of energy storage resources. Scholars such as SCHICK C et al. used market clearing models to study the effect of energy storage on peak load and price reduction. Scholars such as LARSEN M et al. obtained the optimal capacity and optimal operating strategy for independent system operators (ISOs) in energy storage investment by solving a two-level "investment-operation" optimization model. Scholars such as KAZEMI M et al. considered the uncertainty of market prices and market energy deployment and proposed a joint bidding strategy for independent energy storage in the energy and reserve markets. Scholars such as GUSTAVO DV et al. introduced the state-of-charge constraint of energy storage into the market clearing model and studied the impact of bidding structure on the service provided by the energy storage system. Scholars such as Xiao Yunpeng et al. considered the coupling constraints when energy storage provides energy and ancillary services and proposed a sequential and joint clearing model for independent energy storage that takes opportunity cost into account. Scholars such as Li Yaowang proposed a low-carbon optimization method and demand response mechanism for power distribution networks that combines carbon emission flow theory to fully tap the emission reduction potential on the electricity consumption side. Hu Jingzhe rationally divides the responsibility for carbon emissions between generator sets and load nodes, and guides users to adjust their electricity consumption behavior according to the node carbon emission factors, thereby mobilizing users' enthusiasm for energy conservation and emission reduction.
[0004] However, the interactions among the carbon, electricity, and certificate markets are complex and volatile. In real-world engineering, intraday / real-time market signals, carbon prices, and power system security constraints differ significantly across time scales: daily planning needs to balance economic efficiency with long-term SOC (State of Charge), while real-time planning requires ensuring frequency and voltage security. Traditional MPC (Model Predictive Control) or rule-based methods lack robustness under high uncertainty, while pure RL (Reinforcement Learning) lacks long-term constraint guarantees and interpretability. This makes it impossible to comprehensively and accurately regulate energy storage in receiving-end urban power grids, and to maintain energy storage security while maximizing economic benefits and carbon emission reduction under real-time extreme disturbances. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a multi-timescale optimization and control system and method for energy storage in urban power grids at the receiving end. This system can achieve optimized control of energy storage in urban power grids at the receiving end, driven by multiple objectives such as safety, economy, and low carbon, and effectively optimize revenue and emission reduction in a carbon-electricity-green certificate coupled market.
[0006] The objective of this invention can be achieved through the following technical solution: a multi-timescale optimization and control system for energy storage in urban power grids at the receiving end, which adopts a hierarchical architecture, including a daytime coordination layer, a real-time execution layer and a multi-entity collaboration layer, wherein the daytime coordination layer is used to determine the daytime output target of energy storage; The real-time execution layer is used to optimize the allocation of the real-time available capacity of energy storage; The multi-agent collaborative layer employs multi-agent deep reinforcement learning (Multi-Agent DRL) to control multiple energy storage devices to perform corresponding actions.
[0007] Furthermore, the daytime coordination layer is equipped with a multi-objective integrated optimization model, including a comprehensive objective function that simultaneously minimizes system operating costs, operating risks, and carbon emissions, as well as multi-objective constraints. The comprehensive objective function includes economic sub-objectives, safety sub-objectives, and low-carbon sub-objectives.
[0008] Furthermore, the comprehensive objective function is specifically as follows: in, For the overall goal, These are respectively the economic sub-objectives, the safety sub-objectives, and the low-carbon sub-objectives. These are the weighting coefficients corresponding to the sub-objectives of economy, safety, and low carbon emissions, respectively, and satisfy the following conditions: ; The multi-objective constraints are specifically as follows: in, In the state of energy storage charge, These are charging power and discharging power, respectively. Contribute to renewable energy For load demand.
[0009] Furthermore, the economic sub-objective specifically includes: Where T is the number of time steps within the optimization period. Let be the electricity price at time t. The power purchased by the power grid at time t. For carbon prices, Let be the carbon emissions generated at time t; The specific security sub-objective is as follows: in, For the system's real-time voltage and frequency, For rated voltage and frequency, This is a weighting factor used to reflect security sensitivity; The specific low-carbon sub-target is as follows: in, This represents the carbon emissions of the power grid when energy storage is not used. This refers to the carbon emissions corresponding to energy storage participating in peak shaving.
[0010] Furthermore, the real-time execution layer is equipped with real-time constraint optimization functions and real-time safety constraints. The real-time execution layer employs a QP (Quadratic Programming) projection method to project actions... Project to the closest action that satisfies real-time safety constraints .
[0011] Furthermore, the real-time constraint optimization function is specifically as follows: in, The number of power grid nodes in the receiving city. The measured values of voltage and frequency at node i are given. The node reference voltage and the system rated frequency, This is an adjustment coefficient used to control the relative importance of each constraint in the optimization process.
[0012] Furthermore, the real-time security constraints are specifically as follows: in, These are the upper and lower limits of the allowable voltage for the node. For the allowable frequency deviation range, This is the upper limit of the rate of change of energy storage charging and discharging power, i.e., the power ramp-up constraint.
[0013] Furthermore, the multi-agent collaborative layer specifically achieves distributed optimization of the energy storage group by coupling local rewards with global performance, with the goal of maximizing global performance.
[0014] Furthermore, the global performance specifically refers to: in, For local rewards targeting energy storage unit i, The variance of the State of Charge (SOC) distribution is used to measure the balance of energy storage resource utilization. It is a balance factor used to penalize uneven energy distribution. The charging / discharging energy of energy storage unit i For node voltage deviation, The current carbon price, This is a reward factor used to balance economic benefits, security constraints, and carbon costs.
[0015] A method for multi-timescale optimization and control of energy storage in urban power grids at the receiving end includes the following steps: S1. Build a layered architecture in the simulation environment, including a daytime coordination layer, a real-time execution layer, and a multi-agent collaboration layer; S2, the daytime coordination layer solves the total amount and optimization model through multi-objective programming or Pareto-based multi-objective evolutionary algorithm, and outputs the power allocation and SOC reference trajectory of each energy storage unit in time period T; S3. The output of the daytime coordination layer is used as the scheduling benchmark. The real-time execution layer uses linearized predictive control (MPC) or fast constraint optimization algorithm to make second-level corrections to the energy storage output power using safety constraints. S4. Obtain global information and output cooperation strategies among multiple energy storage devices through multi-agent reinforcement learning to control the working status of each energy storage device.
[0016] Compared with the prior art, the present invention has the following advantages: This invention designs a layered architecture, including a daytime coordination layer, a real-time execution layer, and a multi-agent collaboration layer. The daytime coordination layer determines the daytime output target of energy storage; the real-time execution layer further optimizes the allocation of real-time available capacity of energy storage; and finally, the multi-agent collaboration layer uses a multi-agent deep reinforcement learning approach to control multiple energy storage devices to perform corresponding actions. This achieves a multi-timescale control scheme that enables distributed energy storage collaboration and optimizes revenue and emission reduction in a carbon-electricity-green certificate coupled market.
[0017] In this invention, the daytime coordination layer is responsible for achieving a "safety-economy-low-carbon multi-objective balance" between intraday regulation and real-time execution. By establishing a comprehensive optimization model, it realizes the short-to-medium-term scheduling of energy storage power and SOC. The optimization objective is to simultaneously minimize system operating costs, operating risks, and carbon emissions. This is solved through multi-objective programming or a Pareto-based multi-objective evolutionary algorithm, outputting the power allocation and SOC reference trajectory of each energy storage unit within the time period T. This is beneficial for subsequent implementation of energy storage optimization and regulation of the receiving-end urban power grid driven by safety-economy-low-carbon multi-objectives.
[0018] In this invention, the real-time execution layer is equipped with real-time constraint optimization functions and real-time safety constraints, and adopts the QP quadratic programming projection method to perform second-level correction on the energy storage power, ensuring system operation safety and frequency stability, so that the system can still maintain safety under real-time extreme disturbances (safety layer action projection guarantee), while maximizing economic benefits and carbon emission reduction.
[0019] In this invention, the multi-agent collaborative layer is based on multi-agent reinforcement learning. It achieves distributed optimization by coupling local rewards and global performance. Each energy storage device is treated as an agent, which performs P / Q actions and rapid support actions under the premise of ensuring safety constraints. At the same time, it can avoid "free-riding due to shared rewards and punishments", so that multiple energy storage devices can achieve near-global optimal scheduling and support under large-scale distributed energy storage deployment. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the system architecture of the present invention; Figure 2 This is a schematic diagram of the method flow of the present invention; The markings in the diagram are as follows: 1. Daytime coordination layer, 2. Real-time execution layer, 3. Multi-entity collaboration layer. Detailed Implementation
[0021] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0022] Example 1 like Figure 1As shown, a multi-timescale optimization and control system for energy storage in a receiving-end urban power grid adopts a hierarchical architecture, including a daytime coordination layer 1, a real-time execution layer 2, and a multi-entity collaboration layer 3. The daytime coordination layer 1 is used to determine the daytime output target of energy storage. The daytime coordination layer 1 is equipped with a multi-objective comprehensive optimization model, including a comprehensive objective function that simultaneously minimizes system operating costs, operating risks, and carbon emissions, as well as multi-objective constraints. The comprehensive objective function includes economic sub-objectives, safety sub-objectives, and low-carbon sub-objectives. The comprehensive objective function is as follows: In the formula, For the overall goal, These are respectively the economic sub-objectives, the safety sub-objectives, and the low-carbon sub-objectives. These are the weighting coefficients corresponding to the sub-objectives of economy, safety, and low carbon emissions, respectively, and satisfy the following conditions: ; The specific multi-objective constraints are as follows: In the formula, In the state of energy storage charge, These are charging power and discharging power, respectively. Contribute to renewable energy For load demand.
[0023] The specific economic sub-objectives are: In the formula, T represents the number of time steps within the optimization period. Let be the electricity price at time t. The power purchased by the power grid at time t. For carbon prices, Let be the carbon emissions generated at time t; The specific security sub-objectives are: In the formula, For the system's real-time voltage and frequency, For rated voltage and frequency, This is a weighting factor used to reflect security sensitivity; The specific low-carbon sub-targets are: In the formula, This represents the carbon emissions of the power grid when energy storage is not used. This refers to the carbon emissions corresponding to energy storage participating in peak shaving.
[0024] Real-time execution layer 2 is used to optimize the allocation of real-time available energy storage capacity. It includes real-time constraint optimization functions and real-time safety constraints. The real-time execution layer employs a QP (Quadratic Programming) projection method to project actions... Project to the closest action that satisfies real-time safety constraints ; The real-time constraint optimization function is as follows: In the formula, The number of power grid nodes in the receiving city. The measured values of voltage and frequency at node i are given. The node reference voltage and the system rated frequency, These are adjustment coefficients used to control the relative importance of each constraint in the optimization process; The specific real-time security constraints are as follows: In the formula, These are the upper and lower limits of the allowable voltage for the node. For the allowable frequency deviation range, This is the upper limit of the rate of change of energy storage charging and discharging power, i.e., the power ramp-up constraint.
[0025] The multi-agent collaboration layer 3 adopts a multi-agent deep reinforcement learning approach to control multiple energy storage devices to perform corresponding actions. Specifically, the multi-agent collaboration layer achieves distributed optimization of the energy storage group by coupling local rewards with global performance and taking the maximization of global performance as the goal. The overall performance is as follows: In the formula, For local rewards targeting energy storage unit i, The variance of the State of Charge (SOC) distribution is used to measure the balance of energy storage resource utilization. It is a balance factor used to penalize uneven energy distribution. The charging / discharging energy of energy storage unit i For node voltage deviation, The current carbon price, This is a reward factor used to balance economic benefits, security constraints, and carbon costs.
[0026] Based on the above system, a multi-timescale optimization and control method for energy storage in urban power grids at the receiving end is implemented, such as... Figure 2 As shown, it includes the following steps: S1. Build a layered architecture in the simulation environment, including a daytime coordination layer, a real-time execution layer, and a multi-agent collaboration layer; S2, the daytime coordination layer solves the total amount and optimization model through multi-objective programming or Pareto-based multi-objective evolutionary algorithm, and outputs the power allocation and SOC reference trajectory of each energy storage unit in time period T; S3. The output of the daytime coordination layer is used as the scheduling benchmark. The real-time execution layer uses linearized predictive control (MPC) or fast constraint optimization algorithm to make second-level corrections to the energy storage output power using safety constraints. S4. Obtain global information and output cooperation strategies among multiple energy storage devices through multi-agent reinforcement learning to control the working status of each energy storage device.
[0027] It should be noted that in the practical application of the above method, an electronic device including a central processing unit (CPU) can be used. This CPU can execute various appropriate actions and processes based on computer program instructions stored in read-only memory (ROM) or loaded from storage units into random access memory (RAM). RAM can also store various programs and data required for device operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0028] Multiple components in the device are connected to an I / O interface, including: input units such as a keyboard, mouse, etc.; output units such as various types of displays, speakers, etc.; storage units such as disks, optical disks, etc.; and communication units such as network interface cards, modems, wireless transceivers, etc. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks. The processing unit performs the various methods and processes described above, such as the method of the present invention. For example, in some embodiments, the method of the present invention may be implemented as a computer software program tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via ROM and / or the communication unit. When the computer program is loaded into RAM and executed by the CPU, one or more steps of the method of the present invention described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute the method of the present invention by any other suitable means (e.g., by means of firmware).
[0029] The functions described above in this invention can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0030] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0031] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0032] Example 2 This embodiment applies the technical solution in Embodiment 1, and the main processes are as follows: I. System Architecture Setup: Layer 1 (T=24h, step size 1h): Daily planning layer, planning the daytime energy storage output target, taking into account energy constraints, market prospects, carbon trading positions and green certificate targets.
[0033] Among them, Layer 1, as the upper layer (planning / intraday), uses a rolling MPC coupled with carbon cost and market or a two-stage optimization to determine the target curve.
[0034] Layer 2 (T=5–15min): Scheduling layer, handles short-term price / load adjustments, allocates reserve / frequency regulation capacity, and uses linear or quadratic programming for fast solution.
[0035] Layer 2, as the middle layer (minutes / seconds), uses rapid optimization / rules to ensure backup and market delivery.
[0036] Layer 3 (T=1s~1s / 100ms): Real-time control layer, introducing multi-agent deep reinforcement learning. Each energy storage device is an agent, which performs P / Q actions and rapid support actions (virtual inertia, low-frequency response) under the premise of ensuring safety constraints.
[0037] Layer 3, as the lower layer (real-time millisecond to second level), uses Safe Reinforcement Learning (Safe RL) / constrained action mapping to achieve fast frequency and voltage support, while ensuring that it does not violate the upper-layer SOC / market commitments.
[0038] II. Multi-objective optimization and reward structure (RL) The objectives are broken down into three categories: economic (E), carbon (C), and security (S), with a single-step reward defined as follows: in, High penalties are imposed for voltage / frequency out-of-range and large-scale oscillations. , , To adjust the weights, Pareto front analysis was used for selection.
[0039] III. Safety RL and Constraint Handling Using Lagrangian Constrained RL or Safety Layer (motion projection): RL proposes actions. The security layer performs a quadratic programming (QP) operation. Project to the closest action that satisfies real-time safety constraints (such as maximum output, avoiding voltage overruns, and frequency thresholds). .
[0040] Combining offline simulation replay with Imitation Learning Warm-start (using MPC strategy to generate demonstration trajectories) accelerates training and ensures initial safety.
[0041] IV. Multi-agent collaboration (communication / centralized critic) A training architecture using a centralized critic / distributed actor (such as the MADDPG or COMA style) is adopted: global information is used during training to learn collaborative strategies, and each agent uses only local observations and a small amount of neighbor information during deployment to reduce communication volume.
[0042] To avoid "free-riding" due to shared rewards and punishments, we define and optimize both local rewards and team-level performance indicators.
[0043] V. Algorithm Implementation and Training Process Establish a high-fidelity simulation environment (coupled with power grid simulator OpenDSS / GridLab-D + market / carbon simulator); The MPC generation strategy example is used as the initial strategy (behavioral cloning). Training phase: parallel environment sampling, centralized critic training, and periodic evaluation of safety constraint satisfaction rate; Deployment phase: Use the security layer of action projection to ensure zero default (or extremely low default probability) in the real system.
[0044] VI. Evaluation Indicators and Verification Evaluation metrics include: economic benefits (annualized net income), annualized carbon emission reduction, statistics on the number / frequency deviation of system safety incidents, and the impact on battery life due to charge-discharge cycles (degradation cost).
[0045] In summary, this solution combines the constraints of MPC (Model Predictive Control) with the real-time adaptability of RL. Its hierarchical design balances long-term commitments (market / carbon position) with short-term security, avoiding the long-term constraint failure of a single RL strategy. Multi-agent collaboration can achieve near-global optimal scheduling and support in large-scale distributed energy storage deployments. Compared with pure MPC / traditional rules, this solution can better adapt to the high volatility of the market / load, thereby increasing revenue and reducing carbon emissions.
Claims
1. A receiving-end urban power grid energy storage multi-timescale optimized control system, characterized in that, A layered architecture is adopted, including a daytime coordination layer, a real-time execution layer, and a multi-entity collaboration layer. The daytime coordination layer is used to determine the daytime output target of energy storage. The real-time execution layer is used to optimize the allocation of the real-time available capacity of energy storage; The multi-agent collaborative layer employs a multi-agent deep reinforcement learning approach to control multiple energy storage devices to perform corresponding actions.
2. The receiving-end urban power grid energy storage multi-timescale optimization and control system according to claim 1, characterized in that, The daytime coordination layer is equipped with a multi-objective integrated optimization model, including a comprehensive objective function that simultaneously minimizes system operating costs, operating risks, and carbon emissions, as well as multi-objective constraints. The comprehensive objective function includes economic sub-objectives, safety sub-objectives, and low-carbon sub-objectives.
3. The receiving-end urban power grid energy storage multi-timescale optimization and control system according to claim 2, characterized in that, The comprehensive objective function is specifically as follows: in, For the overall goal, These are respectively the economic sub-objectives, the safety sub-objectives, and the low-carbon sub-objectives. These are the weighting coefficients corresponding to the sub-objectives of economy, safety, and low carbon emissions, respectively, and satisfy the following conditions: ; The multi-objective constraints are specifically as follows: in, In the state of energy storage charge, These are charging power and discharging power, respectively. Contribute to renewable energy For load demand.
4. The receiving-end urban power grid energy storage multi-timescale optimization and control system according to claim 3, characterized in that, The specific economic sub-objective is as follows: Where T is the number of time steps within the optimization period. Let be the electricity price at time t. The power purchased by the power grid at time t. For carbon prices, Let be the carbon emissions generated at time t; The specific security sub-objective is as follows: in, For the system's real-time voltage and frequency, For rated voltage and frequency, This is a weighting factor used to reflect security sensitivity; The specific low-carbon sub-target is as follows: in, This represents the carbon emissions of the power grid when energy storage is not used. This refers to the carbon emissions corresponding to energy storage participating in peak shaving.
5. The receiving-end urban power grid energy storage multi-timescale optimization and control system according to claim 1, characterized in that, The real-time execution layer is equipped with real-time constraint optimization functions and real-time safety constraints. The real-time execution layer employs a QP quadratic programming projection method to process actions. Project to the closest action that satisfies real-time safety constraints .
6. The receiving-end urban power grid energy storage multi-timescale optimization and control system according to claim 5, characterized in that, The real-time constraint optimization function is specifically as follows: in, The number of power grid nodes in the receiving city. The measured values of voltage and frequency at node i are given. The node reference voltage and the system rated frequency, This is an adjustment coefficient used to control the relative importance of each constraint in the optimization process.
7. The receiving-end urban power grid energy storage multi-timescale optimization and control system according to claim 6, characterized in that, The specific real-time security constraints are as follows: in, These are the upper and lower limits of the allowable voltage for the node. For the allowable frequency deviation range, This is the upper limit of the rate of change of energy storage charging and discharging power, i.e., the power ramp-up constraint.
8. The receiving-end urban power grid energy storage multi-timescale optimization and control system according to claim 1, characterized in that, Specifically, the multi-agent collaboration layer achieves distributed optimization of the energy storage group by coupling local rewards with global performance, with the goal of maximizing global performance.
9. A receiving-end urban power grid energy storage multi-timescale optimization and control system according to claim 8, characterized in that, The overall performance specifically refers to: in, For local rewards targeting energy storage unit i, The variance of the State of Charge (SOC) distribution is used to measure the balance of energy storage resource utilization. It is a balance factor used to penalize uneven energy distribution. The charging / discharging energy of energy storage unit i For node voltage deviation, The current carbon price, This is a reward factor used to balance economic benefits, security constraints, and carbon costs.
10. A method for multi-timescale optimization and control of energy storage in receiving-end urban power grids, characterized in that, Includes the following steps: S1. Build a layered architecture in the simulation environment, including a daytime coordination layer, a real-time execution layer, and a multi-agent collaboration layer; S2, the daytime coordination layer solves the total amount and optimization model through multi-objective programming or Pareto-based multi-objective evolutionary algorithm, and outputs the power allocation and SOC reference trajectory of each energy storage unit in time period T; S3. The output of the daytime coordination layer is used as the scheduling benchmark. The real-time execution layer uses linear predictive control MPC or fast constraint optimization algorithm to make second-level corrections to the energy storage output power using safety constraints. S4. Obtain global information and output cooperation strategies among multiple energy storage devices through multi-agent reinforcement learning to control the working status of each energy storage device.