Partitioned virtual power plant optimization scheduling method and device based on reinforcement learning and medium

By constructing global and partition-based optimized scheduling models and combining near-end policy optimization algorithms and heterogeneous agent near-end policy optimization algorithms, the scheduling problem of virtual power plants under diverse resources is solved, achieving efficient and flexible resource allocation and cost optimization.

CN120728727BActive Publication Date: 2025-12-30EAST CHINA BRANCH OF STATE GRID CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510626896.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-12-30
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

Traditional virtual power plant architectures struggle to achieve real-time optimization and rapid response when faced with diverse resources. Centralized optimization models are ill-suited to diverse needs and widely distributed energy zones. Reinforcement learning methods in hierarchical architectures suffer from insufficient support for agent heterogeneity and imperfect inter-level coordination mechanisms, making it difficult to balance global objectives and autonomous decisions for each zone. Frequent policy conflicts also affect system stability.

Method used

By constructing global and regional optimization scheduling models, and combining near-end policy optimization algorithms and heterogeneous agent near-end policy optimization algorithms, cross-regional optimal scheduling strategies are generated, the target scheduling method for power zones is determined, the scheduling efficiency and flexibility of virtual power plants are improved, resource allocation is optimized, and energy consumption and costs are reduced.

Benefits of technology

It enables efficient scheduling of virtual power plants in multi-zone power systems, improves the flexibility and efficiency of resource allocation, and reduces energy consumption and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120728727B_ABST
    Figure CN120728727B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of virtual power plant, and provides a partitioned virtual power plant optimization scheduling method and device based on reinforcement learning and a medium, comprising: determining a power partition set based on resource information of a virtual power plant; constructing a global optimization scheduling model, establishing global constraint conditions about the global optimization scheduling model; constructing a partition optimization scheduling model, establishing a partition constraint condition set; determining a global optimization scheduling strategy based on a proximal policy optimization algorithm, the global optimization scheduling model and the global constraint conditions; determining a partition optimization scheduling strategy based on a heterogeneous agent proximal policy optimization algorithm, the partition optimization scheduling model and the partition constraint condition set; and performing joint training to generate a cross-zone optimal scheduling strategy and determine a target scheduling method of each partition. The embodiment improves the scheduling efficiency and flexibility of the virtual power plant in the multi-partition power system, optimizes resource allocation, and reduces energy consumption and cost by combining global and partition optimization scheduling strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of virtual power plant technology, and more specifically, to a method, apparatus, and medium for optimized scheduling of partitioned virtual power plants based on reinforcement learning. Background Technology

[0002] As renewable energy accounts for an increasing proportion of the global power system, the frequency stability of the power grid faces challenges. The volatility and randomness of new energy sources such as wind and solar power lead to a decrease in grid system inertia, putting pressure on frequency regulation capabilities. Traditional synchronous generator sets have strong frequency regulation capabilities, but with the increasing proportion of new energy sources, their response speed is insufficient to meet the demands of instantaneous frequency changes.

[0003] Virtual Power Plants (VPPs), as a technological means of aggregating and scheduling distributed flexible resources, can coordinate various resources such as distributed generation, energy storage, electric vehicles, and demand response to participate in grid frequency regulation ancillary services, improve grid regulation capabilities, and ensure the stability of new power systems. However, VPP architectures face problems such as complex information interaction, low decision-making efficiency, and difficulties in collaborative control. In particular, when dealing with diverse resources, centralized methods struggle to achieve real-time optimization and rapid response.

[0004] To address these issues, a partitioned hierarchical architecture has emerged, achieving a balance between "global optimization" and "local autonomy" through hierarchical collaboration mechanisms, becoming a key direction for improving VPP scheduling efficiency. However, while centralized optimization models can aggregate resources, they struggle to adapt to diverse needs and widely distributed energy partitions. Although reinforcement learning methods enhance distributed decision-making capabilities, problems such as insufficient support for agent heterogeneity and imperfect inter-level coordination mechanisms still exist in hierarchical architecture designs, making it difficult to balance global objectives and partitioned autonomous decisions, and frequent policy conflicts affect system stability. Furthermore, traditional multi-agent reinforcement learning algorithms suffer from inconsistent policy updates and convergence difficulties, failing to effectively meet the VPP's need for maximizing global returns. Summary of the Invention

[0005] This disclosure provides at least one method, apparatus, and medium for optimized scheduling of partitioned virtual power plants based on reinforcement learning. By combining global and partitioned optimized scheduling strategies, it improves the scheduling efficiency and flexibility of virtual power plants in multi-partition power systems, optimizes resource allocation, and reduces energy consumption and costs.

[0006] This disclosure provides a reinforcement learning-based method for optimizing the scheduling of partitioned virtual power plants, including:

[0007] The power zone set is determined based on the resource information of the virtual power plant; wherein, the power zone set includes wind-solar-storage coupling zone, energy storage-electric vehicle zone, adjustable load zone, and gas turbine zone;

[0008] A global optimization scheduling model is constructed, and global constraints on the global optimization scheduling model are established based on preset market capacity limits and preset capacity feasibility limits for each zone; wherein, the global optimization scheduling model takes maximizing the total profit of the virtual power plant participating in frequency regulation ancillary services as the global objective function, and takes the frequency regulation capacity of each zone and the total frequency regulation capacity declared to the market as decision variables;

[0009] A partitioned optimal scheduling model is constructed, and a set of partitioned constraints for the partitioned optimal scheduling model is established; wherein, the partitioned optimal scheduling model takes minimizing the resource operating cost and instruction tracking deviation of each partition as the partitioned objective function, and takes the frequency regulation resource output within each partition as the decision variable;

[0010] A global optimization scheduling strategy for the global optimization scheduling model is determined based on the near-end policy optimization algorithm, the global optimization scheduling model, and the global constraints; and a partition optimization scheduling strategy for the partition optimization scheduling model is determined based on the heterogeneous agent near-end policy optimization algorithm, the partition optimization scheduling model, and the set of partition constraints.

[0011] Based on the global optimization scheduling strategy and the partition optimization scheduling strategy, the near-end strategy optimization algorithm and the heterogeneous agent near-end strategy optimization algorithm are jointly trained to generate the cross-region optimal scheduling strategy, and the target scheduling method for each partition is determined based on the optimal scheduling strategy.

[0012] This disclosure provides a reinforcement learning-based partitioned virtual power plant optimization scheduling device, comprising:

[0013] The power zoning determination module is used to determine a set of power zoning zones based on the resource information of the virtual power plant; wherein, the set of power zoning zones includes wind-solar-storage coupling zones, energy storage-electric vehicle zones, adjustable load zones, and gas turbine zones;

[0014] A global model construction module is used to construct a global optimal scheduling model and establish global constraints on the global optimal scheduling model based on preset market capacity limits and preset capacity feasibility limits for each zone; wherein, the global optimal scheduling model takes maximizing the total profit of the virtual power plant participating in frequency regulation ancillary services as the global objective function, and takes the frequency regulation capacity of each zone and the total frequency regulation capacity declared to the market as decision variables;

[0015] The partition model construction module is used to construct a partition optimization scheduling model and establish a set of partition constraints for the partition optimization scheduling model; wherein, the partition optimization scheduling model takes minimizing the resource operating cost and instruction tracking deviation of each partition as the partition objective function, and takes the frequency regulation resource output within each partition as the decision variable;

[0016] The scheduling strategy determination module is used to determine a global optimization scheduling strategy for the global optimization scheduling model based on the near-end strategy optimization algorithm, the global optimization scheduling model, and the global constraints; and to determine a partition optimization scheduling strategy for the partition optimization scheduling model based on the heterogeneous agent near-end strategy optimization algorithm, the partition optimization scheduling model, and the set of partition constraints.

[0017] The scheduling method determination module is used to jointly train the near-end policy optimization algorithm and the heterogeneous agent near-end policy optimization algorithm based on the global optimization scheduling strategy and the partition optimization scheduling strategy, generate the cross-region optimal scheduling strategy, and determine the target scheduling method for each partition based on the optimal scheduling strategy.

[0018] This disclosure provides a computer device, including a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the reinforcement learning-based partitioned virtual power plant optimization scheduling method as described in any of the above possible embodiments.

[0019] This disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the reinforcement learning-based partitioned virtual power plant optimal scheduling method as described in any of the possible embodiments above.

[0020] The reinforcement learning-based method, apparatus, and medium for optimal scheduling of partitioned virtual power plants provided in this disclosure determine a set of power partitions, construct global and partitioned optimal scheduling models, and establish global and partitioned constraints respectively. Using a near-end policy optimization algorithm and a heterogeneous proxy near-end policy optimization algorithm, global and partitioned optimal scheduling strategies are determined and jointly trained to generate cross-regional optimal scheduling strategies. Thus, by combining global and partitioned optimal scheduling strategies, the scheduling efficiency and flexibility of virtual power plants in multi-partitioned power systems are improved, resource allocation is optimized, and energy consumption and costs are reduced.

[0021] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings referenced in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0023] Figure 1 A flowchart of a partitioned virtual power plant optimization scheduling method based on reinforcement learning provided in an embodiment of this disclosure is shown.

[0024] Figure 2 A flowchart of an algorithm joint training method provided by an embodiment of this disclosure is shown;

[0025] Figure 3 A flowchart illustrating a data processing method for training data provided in an embodiment of this disclosure is shown.

[0026] Figure 4 A flowchart of an algorithmic joint parameter update method provided in an embodiment of this disclosure is shown;

[0027] Figure 5 A schematic diagram of the structure of a partitioned virtual power plant optimization scheduling device based on reinforcement learning provided in an embodiment of this disclosure is shown.

[0028] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0030] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0031] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0032] To facilitate understanding of this embodiment, the execution subject of the reinforcement learning-based partitioned virtual power plant optimization scheduling method provided in this disclosure will first be described in detail. The execution subject of the reinforcement learning-based partitioned virtual power plant optimization scheduling method provided in this disclosure is a computer device. This computer device can be a server. Specifically, the server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms.

[0033] The reinforcement learning-based partitioned virtual power plant optimal scheduling method provided in this application embodiment will be described in detail below with reference to the accompanying drawings. See also Figure 1 The diagram shows a flowchart of a partitioned virtual power plant optimization scheduling method based on reinforcement learning provided in this embodiment of the present disclosure. The method includes the following steps S101 to S105:

[0034] S101, determine the set of power zones based on the resource information of the virtual power plant.

[0035] Understandingly, a virtual power plant refers to the use of information technology to combine different types of power resources (such as wind, solar, energy storage, and loads) distributed across various locations into a virtual whole for centralized management and optimized dispatch. Based on the characteristics, geographical location, and dispatch requirements of the resources within the virtual power plant, the resources can be divided into different management areas to obtain a set of power zones. These power zones include wind-solar-storage coupling zones, energy storage-electric vehicle zones, adjustable load zones, and gas turbine zones. Specifically, a wind-solar-storage coupling zone refers to an area that combines wind and solar power generation with energy storage systems (such as batteries). Wind and solar energy are volatile, and energy storage devices can smooth these fluctuations, improving system stability. An energy storage-electric vehicle zone combines energy storage devices with electric vehicles (EVs), using the charging and discharging of EVs to regulate grid load, thereby achieving optimized power management. An adjustable load zone can regulate power demand by adjusting user load demands. Adjustable loads can be equipment from industrial, commercial, or residential users, participating in dispatch through demand response mechanisms. Gas turbines in zonal gas turbine systems are one of the traditional frequency regulation devices, capable of quickly starting or adjusting output according to power demand to ensure the stability of the power grid frequency.

[0036] In this way, by accurately classifying and managing various types of power resources in a virtual power plant, this disclosure can achieve efficient scheduling and optimized allocation of power resources, further improving the flexibility and reliability of power supply.

[0037] Here, this disclosure also designs a hierarchical collaborative management and control mode for virtual power plants, that is, the VPP is designed as a three-layer structure (upper-middle-lower): the upper layer is the control and management system, which is responsible for integrating the real-time status and market information of all adjustable resources in the entire domain as the global decision-making core, generating the middle-layer collaborative scheduling strategy through optimization algorithms, and participating in the capacity declaration and clearing bidding in the frequency regulation ancillary service market; the middle layer is the frequency regulation unit aggregation operation system, which collects and summarizes the real-time operation data of the lower-layer frequency regulation units and feeds it back to the upper layer, and distributes the upper-layer instructions to the lower-layer frequency regulation units; the lower layer is the frequency regulation units, which adjust the real-time output of the units according to the instructions distributed by the middle layer.

[0038] For example, after implementing the three-layer structure of the virtual power plant, a framework for partitioned and hierarchical VPPs to participate in frequency regulation ancillary services can be further designed. This framework includes the following key steps: First, the intelligent interactive terminal of the frequency regulation unit collects the second-level operating data of the user's adjustable resources in real time and transmits the data to the middle-layer system through a standardized protocol; then, the middle-layer system aggregates heterogeneous resources into virtual machine groups based on this data and calculates the available capacity and response accuracy parameters of the virtual machine groups; next, the top-layer system integrates the parameters of these virtual machine groups and frequency regulation market information, and uses an optimization algorithm to generate partitioned frequency regulation capacity allocation ratio instructions; subsequently, the middle-layer system decomposes this frequency regulation capacity allocation ratio instruction into specific control parameters and sends them to the bottom-layer terminals; finally, the bottom-layer terminals execute the instructions and feed back the actual output data to the middle-layer system, completing a closed-loop verification process.

[0039] In this way, through this partitioned and layered structure, virtual power plants can participate in frequency regulation ancillary services more flexibly and efficiently, not only improving the system's response speed but also enhancing resource utilization efficiency. This model enables virtual power plants to achieve a high degree of synergy between different levels, thereby responding to frequently changing demands and dispatching tasks in the electricity market, ensuring the stable operation of the power grid and the reliability of power supply.

[0040] S102, construct a global optimization scheduling model, and establish global constraints on the global optimization scheduling model based on preset market capacity limits and preset capacity feasibility limits for each partition.

[0041] Here, the global optimal scheduling model aims to maximize the total profit of virtual power plants participating in frequency regulation ancillary services. It considers market capacity constraints (i.e., the maximum or minimum frequency regulation capacity the market can accept) and the capacity feasibility constraints of each zone (the range of frequency regulation capacity each zone can provide or absorb). The global optimal scheduling model uses maximizing the total profit of virtual power plants participating in frequency regulation ancillary services as the global objective function. Frequency regulation ancillary services refer to balancing power supply and demand by adjusting generation and load when grid frequency fluctuates, in order to maintain grid stability. The global optimal scheduling model uses the frequency regulation capacity of each zone and the total frequency regulation capacity declared to the market as decision variables. These variables determine how virtual power plants allocate and schedule resources in each zone to optimize overall scheduling.

[0042] Specifically, by comprehensively considering market revenue, operating costs, and penalty costs, a global objective function for maximizing the revenue of VPPs participating in frequency regulation ancillary services can be established, which can be expressed as:

[0043]

[0044] in, This represents the frequency modulation capacity allocated to partition i during time period t; λ represents the total frequency regulation capacity declared to the market by the virtual power plant during time period t; cap (t) represents the capacity price for time period t; λ mile (t) represents the mileage compensation price for time period t; i represents the zone number, where number 1 represents the wind-solar-storage coupling zone, number 2 represents the energy storage-electric vehicle zone, number 3 represents the adjustable load zone, and number 4 represents the gas turbine zone; η i This is represented as the frequency modulation efficiency coefficient of partition i, obtained by calibration using historical data. C represents the operating cost of partition i in time period t; pen (t) represents the capacity deviation penalty for time period t; η dev This is represented as the capacity deviation penalty coefficient.

[0045] For example, based on the aforementioned global objective function, global constraints for the global optimization scheduling model can be established based on preset market capacity limits and preset capacity feasibility limits for each partition. Here, the market capacity limit refers to the maximum amount of resources or scheduling that the entire market can withstand at a certain moment or within a certain period. This limit is usually set based on factors such as the market's capacity ceiling, the total supply of resources, and the market's load capacity. To ensure the sustainable operation of the market, the scheduling strategy needs to take into account these market capacity ceilings to avoid overload operation or resource waste. Meanwhile, the capacity feasibility limits for each partition are constraints on resource scheduling within different regions or partitions. The capacity limits for each partition need to take into account the resource supply capacity of each region, the carrying capacity of the infrastructure, and the operational limitations of the equipment. To ensure that each partition can operate effectively within its capacity, the scheduling scheme must comply with these local capacity constraints in the resource allocation of each partition.

[0046] Specifically, global constraints can be expressed as:

[0047]

[0048] in, This represents the maximum number of applications that are subject to the market capacity limit preset for time period t. This represents the maximum adjustable capacity of partition i within time period t, which is the preset capacity feasibility limit.

[0049] S103, construct a partition-optimized scheduling model and establish a set of partition constraints for the partition-optimized scheduling model.

[0050] Here, unlike the global optimal scheduling model, the regional optimal scheduling model uses minimizing the resource operating costs and command tracking deviation of each region as the objective function, and frequency regulation resource output within each region as the decision variable. The regional optimal scheduling model focuses more on the resource operating costs and command tracking deviation of each power region. Command tracking deviation refers to the difference between the actual frequency regulation output and the target frequency regulation output of a region; minimizing this deviation helps improve the accuracy of scheduling and the stability of the system. Using the frequency regulation resource output within each region as the decision variable means that each region needs to adjust its resource output according to the frequency regulation requirements of the power grid to ensure the stability of the power grid frequency.

[0051] Specifically, taking into account both resource operating costs and instruction tracking deviations, a partition objective function that minimizes partition frequency modulation costs can be established, as follows:

[0052]

[0053] in, This represents the frequency regulation cost of frequency regulation unit j within zone i; This represents the frequency regulation output of frequency regulation unit j within time period t and partition i.

[0054] Understandably, after constructing the zonal optimal scheduling model, a set of zonal constraints for the zonal optimal scheduling model can be established. This set of zonal constraints includes constraints corresponding to each power zonal, namely, wind-solar-storage coupling zonal constraints, energy storage-electric vehicle zonal constraints, adjustable load zonal constraints, and gas turbine zonal constraints.

[0055] Specifically, the zoning constraints for wind-solar-storage coupling can be expressed as:

[0056]

[0057] in, P represents the frequency regulation capacity allocated to the wind-solar-storage coupling zone; wind (t) represents the real-time output of wind power; P PV (t) represents the real-time output of the photovoltaic system; P bat (t) represents the real-time output of the energy storage battery; This represents the predicted wind power output. This is represented as the predicted output of photovoltaic power. This indicates the charging power limit of the energy storage battery; The expression represents the discharge power limit of the energy storage battery; SOC(t) represents the state of charge of the energy storage battery during time period t; η ch Indicated as the charging efficiency of the energy storage battery; η dis This is expressed as the discharge efficiency of the energy storage battery; The energy storage battery charging efficiency is expressed as the time period t. State of Charge (SOC) represents the battery discharge efficiency over time period t. min State of Charge (SOC) represents the minimum state of charge of an energy storage battery. max This represents the maximum state of charge of the energy storage battery.

[0058] Specifically, the energy storage-electric vehicle zoning constraint can be expressed as:

[0059]

[0060] in, This represents the frequency regulation capacity allocated to the energy storage-electric vehicle zone; P represents the charging and discharging power of the energy storage battery in the energy storage-electric vehicle zone. EV,k (t) represents the charging and discharging power of electric vehicle k; N EV This represents the number of electric vehicles participating in frequency modulation. This indicates the charging power limit for electric vehicles; This refers to the discharge power limit of electric vehicles; This represents the state of charge of the battery of electric vehicle k during the departure time t; This represents the minimum state of charge requirement that the user has for electric vehicle k at the moment of departure.

[0061] Specifically, the adjustable load zoning constraint can be expressed as:

[0062]

[0063] in, This represents the frequency regulation capacity allocated to the adjustable load zone; ΔP load,m (t) represents the power adjustment amount of the adjustable load during time period t; This represents the maximum allowable adjustable power for the adjustable load. This represents the total regulating energy limit of the adjustable load within a day.

[0064] Specifically, the gas turbine zoning constraints can be expressed as:

[0065]

[0066] in, P represents the frequency regulation capacity allocated to the gas turbine zone; gas (t) represents the real-time power of the gas turbine; This represents the minimum processing capacity of the gas turbine. This represents the maximum value that the gas turbine can handle. This represents the gas turbine ramp rate limit; T on(t) represents the continuous operating time of the gas turbine during time interval t; T off (t) represents the downtime of the gas turbine in time period t; This represents the minimum operating time of the gas turbine. This represents the minimum downtime of the gas turbine.

[0067] S104, determine a global optimization scheduling strategy for the global optimization scheduling model based on the near-end policy optimization algorithm, the global optimization scheduling model, and the global constraints; and determine a partition optimization scheduling strategy for the partition optimization scheduling model based on the heterogeneous agent near-end policy optimization algorithm, the partition optimization scheduling model, and the set of partition constraints.

[0068] Understandably, based on the Proximal Policy Optimization (PPO) algorithm, combined with a global optimal scheduling model and global constraints, a globally optimal scheduling policy can be determined. Here, PPO, as an efficient and stable reinforcement learning algorithm, gradually approaches the globally optimal scheduling scheme through trial and error and iteration in continuous interaction with the environment. In application, this algorithm not only considers the immediate benefits under the current policy but also ensures the stability and improvement direction of the policy by limiting the policy update magnitude, thus effectively addressing the complexity and uncertainty in the global scheduling problem.

[0069] Here, when determining the globally optimal scheduling policy based on the near-end policy optimization algorithm, the globally optimal scheduling model, and global constraints, it is necessary to first model the Markov Decision Process (MDP) based on the near-end policy optimization algorithm, the globally optimal scheduling model, and global constraints to obtain the global decision framework. The Markov Decision Process (MDP) is a fundamental framework in reinforcement learning used to describe decision problems. An MDP consists of a state space, an action space, transition probabilities, and a reward function. In this disclosure, MDP is used to model the globally optimal scheduling problem, treating each scheduling decision process as a state transition process, thereby learning the optimal policy through an algorithm. The global decision framework is a decision system framework constructed based on the modeling results of the Markov Decision Process. In this framework, various decision factors are comprehensively considered, and the optimal scheduling policy is solved under constraints. Then, after obtaining the global decision space and comprehensively considering various system constraints and objectives, a decision policy (i.e., the globally optimal scheduling policy) for scheduling can be determined based on reinforcement learning methods.

[0070] For example, the frequency regulation capacity price, frequency regulation mileage compensation price, maximum market-allowed application capacity, maximum adjustable capacity of each zone, historical frequency regulation efficiency coefficient of each zone, average state of charge of energy storage in the wind-solar-storage coupling zone, number of EVs connected in the energy storage-electric vehicle zone, average state of charge of electric vehicles, current output of gas turbines, current time period, predicted output of wind power, and predicted output of photovoltaic power are taken as the state space of the global optimization scheduling strategy, and are represented as follows:

[0071]

[0072] The frequency regulation capacity of each partition and the total frequency regulation capacity declared by VPPs in the frequency regulation ancillary service market are used as the action space of the global optimization scheduling strategy. The action space of the global optimization scheduling strategy is as follows:

[0073]

[0074] Based on the above global objective function, the reward function for the global optimization scheduling strategy can be designed as follows:

[0075]

[0076] in, This is represented as the state space of the global optimization strategy; N represents the average state of charge of the energy storage in the wind-solar-storage coupling zone; EV (t) represents the number of EVs connected to the energy storage-electric vehicle zone during time period t; This represents the average state of charge of electric vehicles during time period t. Represented as the action space of the global optimization strategy; r ppo (t) represents the reward function of the global optimization strategy.

[0077] Furthermore, to address the need for finer-grained optimization, this disclosure introduces the Heterogeneous-Agent Proximal Policy Optimization (HAPPO) algorithm. HAPPO is specifically designed for partition-based optimized scheduling models, aiming to handle the diverse challenges arising from the different characteristics of partitions. Unlike single-policy PPO, HAPPO can assign PPO agents with differentiated parameters to each partition for training based on the specific conditions of different partitions (such as resource distribution, demand patterns, etc.). This means that each partition can learn and converge to the scheduling policy most suitable for its environment based on its own characteristics, thereby maximizing resource utilization efficiency and response speed within each partition while ensuring global efficiency. Through HAPPO, it is possible to more flexibly adapt to the heterogeneity between partitions and achieve more refined and efficient partition-based optimized scheduling policies.

[0078] Specifically, similar to the method for determining the global optimal scheduling strategy, when solving the partitioned optimal scheduling strategy, a heterogeneous agent near-end strategy optimization algorithm, a partitioned optimal scheduling model, and a set of partitioned constraints can be combined to model the Markov decision process of each partition. In this way, the complex scheduling problem can be decomposed into multiple independent sub-problems, each corresponding to a partitioned decision framework, thereby achieving independent optimization of the scheduling of each partition. Specifically, the partitioned decision frameworks can include: a wind-solar-storage coupled partitioned decision framework, an energy storage-electric vehicle partitioned decision framework, an adjustable load partitioned decision framework, and a gas turbine partitioned decision framework. These frameworks model different types of energy resources and load characteristics, providing a refined reference for subsequent optimization decisions.

[0079] Next, based on these partition decision frameworks, and combining the heterogeneous agent near-end policy optimization algorithm with the partition optimization scheduling model and the set of partition constraints, the optimal scheduling strategy for each partition is further derived and determined. The goal of this stage is to determine, based on the specific scheduling model and constraints, a scheduling strategy that satisfies all constraints and effectively improves the overall system performance using the heterogeneous agent near-end policy optimization algorithm. These partition optimization scheduling strategies can include wind-solar-storage coupling partition optimization scheduling strategy, energy storage-electric vehicle partition optimization scheduling strategy, adjustable load partition optimization scheduling strategy, and gas turbine partition optimization scheduling strategy.

[0080] Here, the wind-solar-storage coupled zoning decision-making framework can be represented as:

[0081]

[0082]

[0083]

[0084] in, This represents the state space of the wind-solar-storage coupling zone optimization scheduling strategy; This represents the action space of the wind-solar-storage coupling zone optimization scheduling strategy; c represents the reward function of the wind-solar-storage coupled zonal optimization scheduling strategy; bat This is expressed as the frequency regulation cost of the energy storage battery.

[0085] Here, the energy storage-electric vehicle zoning decision-making framework can be represented as:

[0086]

[0087]

[0088]

[0089] in, This is represented as the state space of the energy storage-electric vehicle zone optimization scheduling strategy; This is represented as the state of charge of each electric vehicle; This represents the action space of the energy storage-electric vehicle zonal optimization scheduling strategy; This is represented as the reward function for the energy storage-electric vehicle zonal optimization scheduling strategy; The frequency regulation cost of energy storage batteries is represented as the energy storage-electric vehicle partition; c EV,k Let k be the frequency modulation cost of the electric vehicle.

[0090] Here, the adjustable load zoning optimization decision framework can be expressed as:

[0091]

[0092]

[0093]

[0094] in, This represents the state space of an adjustable load partitioning optimization scheduling strategy. This represents the action space of the adjustable load partitioning optimization scheduling strategy. c represents the reward function of the adjustable load partitioning optimization scheduling strategy; load,m This represents the adjustment cost of the adjustable load during time period t.

[0095] Here, the gas turbine zoning decision framework can be represented as:

[0096]

[0097]

[0098]

[0099] in, This represents the state space of the gas turbine partitioning optimization scheduling strategy. This represents the action space of the gas turbine zonal optimization scheduling strategy. c represents the reward function of the gas turbine partitioning optimization scheduling strategy; gas Expressed as the fuel cost of a gas turbine; Expressed as carbon emission cost.

[0100] In some possible embodiments, when implementing the HAPPO algorithm, the measurement data of each partition is set up as agents, such as Agent 1, Agent 2, etc., each representing a partition and trained and learned according to the specific model and constraints of that partition. In this way, the HAPPO algorithm can fully utilize the heterogeneity of each partition to achieve a more refined, efficient, and flexible partition optimization scheduling strategy. This highly customized and intelligent scheduling method can improve the overall performance of the energy system.

[0101] S105, based on the global optimization scheduling strategy and the partition optimization scheduling strategy, the near-end strategy optimization algorithm and the heterogeneous agent near-end strategy optimization algorithm are jointly trained to generate the cross-region optimal scheduling strategy, and the target scheduling method for each partition is determined based on the optimal scheduling strategy.

[0102] It is understandable that after obtaining the global optimal scheduling strategy and the partition optimal scheduling strategy, the optimal scheduling strategy across regions can be generated based on the global optimal scheduling strategy and the partition optimal scheduling strategy by jointly training the near-end policy optimization algorithm (PPO) and the heterogeneous agent near-end policy optimization algorithm (HAPPO), and the specific target scheduling method for each partition can be determined accordingly.

[0103] For example, refer to Figure 2 As shown, when jointly training the near-end policy optimization algorithm and the heterogeneous agent near-end policy optimization algorithm based on the global optimization scheduling strategy and the partition optimization scheduling strategy, the following steps S201 to S205 may be included:

[0104] S201: Obtain the training dataset.

[0105] The training dataset comprises multiple subsets, each covering several time periods within a preset time cycle. These subsets include both global training data and specific training data for each partition, reflecting the operational characteristics and demands of the power system under different times and conditions. For example, with a day as a time cycle and one hour as a time period, a training dataset subset would include 24 data points (global training data for each hour and training data corresponding to each partition).

[0106] In some other embodiments, the training dataset does not contain data generated by the PPO algorithm.

[0107] S202: Perform data processing operations on the global training data corresponding to each time period and the training data corresponding to each partition in each training data subset.

[0108] After acquiring the dataset, step S202 processes the data to generate optimized scheduling strategies for both global and partitioned scenarios. This aims to generate appropriate scheduling instructions based on the data and provide a basis for subsequent strategy optimization and reward calculation. (Refer to...) Figure 3 As shown, the data processing operation may include the following steps S2021 to S2025:

[0109] S2021, Generate a set of global action values ​​based on the near-end policy optimization algorithm and the global training data.

[0110] Here, based on the PPO algorithm and global training data, a global action value set can be generated, containing frequency regulation capacity allocation instructions for different zones and total frequency regulation capacity declaration values. These instructions are the concrete manifestation of the global scheduling strategy, guiding how each zone works collaboratively to meet the overall needs of the power grid.

[0111] S2022, the frequency modulation capacity allocation instruction of the corresponding partition in the global action value set is input to the corresponding partition optimization scheduling strategy in the partition optimization scheduling strategy.

[0112] Understandably, in order to ensure that the partition-optimized scheduling strategy can respond reasonably to global instructions according to the guiding principles of global scheduling, the frequency modulation capacity allocation instruction of the corresponding partition in the global action value set can be input into the corresponding partition-optimized scheduling strategy in the partition-optimized scheduling strategy.

[0113] S2023, Based on the heterogeneous agent near-end policy optimization algorithm, the optimized scheduling strategy of each partition and the training data corresponding to each partition, the specific output action value of the frequency modulation resource in each partition is generated.

[0114] Specifically, the frequency modulation capacity allocation instructions from the global action value set are input into the optimization scheduling strategy of each partition. Using the HAPPO algorithm and the training data of each partition, specific output action values ​​for frequency modulation resources within each partition are generated. Here, these action values ​​are the direct output of the partition scheduling strategy, reflecting how each partition makes the optimal response based on the global instructions and its own conditions.

[0115] S2024, calculate the partition reward value of the heterogeneous agent near-end strategy optimization algorithm based on the reward function of the partition optimization scheduling strategy and the specific output action value of the frequency modulation resources in each partition; and calculate the global reward value of the near-end strategy optimization algorithm based on the reward function of the global optimization scheduling strategy, the frequency modulation capacity allocation instructions corresponding to different partitions and the total frequency modulation capacity declaration value.

[0116] Understandably, the partitioned optimization scheduling strategy focuses more on the output action values ​​of specific frequency regulation resources within each partition, aiming to improve the stability and efficiency of the entire system by optimizing the local frequency regulation resource allocation. Under this strategy, the output action values ​​of frequency regulation resources in each partition are calculated by the Heterogeneous Proximity Policy Optimization (PPPO) algorithm. These partition reward values ​​are derived based on a comprehensive consideration of the specific output actions of each partition, accurately reflecting the degree of optimization in resource scheduling within the partition. In this way, the Heterogeneous Proximity Policy Optimization (PPPO) algorithm can optimize the scheduling strategy for each partition. Simultaneously, based on the reward function of the global optimization scheduling strategy, a global reward value can be calculated by comprehensively considering the frequency regulation capacity of each partition and the overall system performance. This reward value is calculated by considering the frequency regulation capacity allocation instructions for different partitions and the total frequency regulation capacity declaration value. In this process, there may be deviations in the execution of instructions by each unit; the role of the global reward function is to take into account the resource allocation situation of each partition, thereby evaluating the overall optimization effect.

[0117] This disclosure proposes a global optimization scheduling strategy and a partition optimization scheduling strategy, which, through the combination of a reward function and a near-end policy optimization algorithm, can perform resource scheduling and optimization at multiple levels. This not only improves the overall frequency regulation capacity management capability of the system but also refines the specific scheduling behavior of each partition, thereby achieving more efficient and flexible resource allocation.

[0118] S2025, based on the global training data corresponding to the current time period in the training data subset, the global action value set, the global reward value, and the global training data corresponding to the next time period in the training data subset, a global training sub-trajectory is generated, and the global training sub-trajectory is stored in the experience replay buffer of the near-end policy optimization algorithm; and based on the training data corresponding to each partition in the current time period in the training data subset, the specific output action value of the frequency modulation resource in each partition, and the training data corresponding to each partition in the next time period in the training data subset, a partition training sub-trajectory is generated, and the partition training sub-trajectory is stored in the experience replay buffer of the heterogeneous agent near-end policy optimization algorithm.

[0119] Here, based on the data from the current time period, action values, and data from the next time period, global and regional training sub-trajectories can be generated and stored in their respective experience replay buffers. These trajectories form the basis of the algorithm's learning and can be used for subsequent parameter updates.

[0120] For example, a training sub-trajectory in any set of training trajectories in the experience replay buffer of the near-end policy optimization algorithm can be represented as: A training sub-trajectory in any set of training trajectories in the experience replay buffer of the heterogeneous agent proximal policy optimization algorithm can be represented as:

[0121] S203, iteratively execute step S202 until multiple training trajectory sets are generated.

[0122] Specifically, by iteratively executing step S202, multiple sets of training trajectories can be generated. Each set of training trajectories corresponds to a subset of training data, and each set includes multiple training sub-trajectories, each corresponding to a time period within the subset of training data. In this way, through multiple iterations, a sufficient number of training trajectories are generated to cover various possible scheduling scenarios.

[0123] S204, based on the experience replay buffer of the near-end policy optimization algorithm and the sets of training trajectories in the experience replay buffer of the heterogeneous agent near-end policy optimization algorithm, perform joint parameter updates on the policy networks of the near-end policy optimization algorithm and the heterogeneous agent near-end policy optimization algorithm.

[0124] For example, the policy networks of the near-end policy optimization algorithm and the heterogeneous agent near-end policy optimization algorithm can be jointly updated using the training trajectory set in the experience replay buffer. This allows the algorithms to jointly optimize their scheduling policies at both the global and partition levels through multiple rounds of training. By jointly updating parameters, the global policy and the partition policy are ensured to coordinate with each other, thereby forming a cross-region optimal scheduling policy.

[0125] Specifically, refer to Figure 4 As shown, when jointly updating the policy networks of the near-end policy optimization algorithm and the heterogeneous agent near-end policy optimization algorithm, the following steps S2041 to S2042 may be included:

[0126] S2041, based on any training trajectory set in the experience replay buffer of the near-end policy optimization algorithm, calculate the advantage value through generalized advantage estimation, and update the parameters of the near-end policy optimization algorithm based on the advantage value and the near-end policy loss function calculation formula.

[0127] Here, the Generalized Advantage Estimation (GAE) method can be used to calculate the advantage value in each training trajectory. The advantage value reflects the additional reward obtained by taking a certain action relative to the average policy in a given state, and is an important indicator for measuring the quality of the action. Its calculation formula is as follows:

[0128]

[0129]

[0130] Where, γ ppo Let λ represent the reward discount factor of the PPO algorithm; λ represents the generalized advantage estimation parameter; δ t+l It is expressed as the time difference error at a time step of t+l.

[0131] For example, based on the calculated advantage value and the proximal policy optimization loss function (PPO) specific to the PPO algorithm, the policy network parameters of the PPO algorithm can be further updated. The proximal policy optimization loss function is designed to control the magnitude of policy updates to avoid learning instability caused by excessive policy changes. By minimizing the proximal policy optimization loss function, the policy network can be gradually optimized, allowing it to evolve towards a better direction. Here, the formula for calculating the proximal policy optimization loss function can be expressed as:

[0132]

[0133] Among them, L clip (θ) represents the near-end policy loss value, E t Represented as the expected value of all time steps in the current subset of training data; π θ (a t |s t ) represents the state s based on the current policy parameter θ. t Generate action a t The probability of; Represented as based on the old policy parameter θ old In state s t Generate action a t The probability; ε represents the hyperparameter; is represented by the dominance value for time period t; clip represents the clipping function.

[0134] S2042, based on any training trajectory set in the experience replay buffer of the heterogeneous agent near-end policy optimization algorithm, calculate the joint action advantage weight corresponding to each partition in the order of updating the partition policy, update the policy parameters of each partition policy based on the joint action advantage weight corresponding to each partition and the calculation formula of the heterogeneous agent near-end policy loss function, and simultaneously optimize the global value network.

[0135] Here, the policy network parameters of the HAPPO algorithm are updated, and the global value network is optimized simultaneously. This step is also based on any set of training trajectories in the HAPPO algorithm's experience replay buffer. Specifically: First, the joint action advantage weights corresponding to each partition are calculated sequentially according to the update order of the partition policies (which can be specified or randomly generated). These weights reflect the degree of influence of each partition policy on the overall scheduling effect within the global policy framework. By accurately calculating these weights, it can be ensured that the global benefit is fully considered when updating the partition policies.

[0136] For example, based on the global value network and generalized advantage estimation, the advantage function of joint actions can be calculated, and its calculation formula is expressed as:

[0137]

[0138] Where, γ HAPPO This represents the reward discount factor for the HAPPO algorithm.

[0139] Here, the formula for calculating the advantage weight of the first strategy can be expressed as:

[0140]

[0141] Here, the update of the advantage weights of the remaining strategies will take into account the impact of the updated strategies, and the update formula is as follows:

[0142]

[0143] Understandably, based on the calculated joint action advantage weights and the heterogeneous agent proximal policy optimization loss function of the HAPPO algorithm, the policy parameters of each partition policy can be updated. Similar to the PPO algorithm, the HAPPO algorithm also optimizes the policy network by minimizing the policy loss function. Simultaneously, to further improve the global consistency of the scheduling policy, the global value network also needs to be optimized synchronously. The global value network is responsible for estimating the global value in a given state and is a crucial basis for guiding global policy optimization. By synchronously optimizing the global value network and the partition policy network, it can be ensured that the global policy and the partition policy remain consistent throughout the training process.

[0144] Here, the formula for calculating the loss function of the heterogeneous agent's near-end policy can be expressed as:

[0145]

[0146] in, Let represent the heterogeneous agent near-end policy loss value for partition i; B represents the batch size of the training data subset; and T represents the number of time steps for training sub-trajectories for each partition. This represents the state of the i-th partition based on the current policy parameter θ. Generate Actions The probability of; Represented as the updated policy parameter θ in the i-th partition new In state Generate Actions The probability of; The optimal scheduling policy for partition i is represented in state i. and actions The potential for strategic improvement.

[0147] For example, the global value network parameter φ is optimized by minimizing the mean squared error between the predicted value and the actual return estimate, as expressed by the formula:

[0148]

[0149] in, V is represented as the reporting estimate for time period t; φ (s t ) represents the global value network pair of states s t Value prediction.

[0150] S205, repeat steps S202 to S204 until the algorithm reward value converges, and generate the cross-regional optimal scheduling strategy.

[0151] Here, steps S202 to S204 are repeated until the algorithm's reward value converges. The convergence of the reward value signifies that the model's scheduling strategy at both the global and partition levels has reached its optimal state. At this point, the generated cross-region optimal scheduling strategy can maximize the scheduling efficiency at both the global and partition levels, ensuring optimal resource allocation.

[0152] This disclosure achieves optimal cross-regional scheduling of power system frequency regulation resources at both the global and regional levels by jointly training the PPO and HAPPO algorithms. Through iterative processing of the training dataset and updating algorithm parameters, a cross-regional optimal scheduling strategy that maximizes both global and regional scheduling efficiency is ultimately generated. Based on this strategy, target scheduling methods for each region can be generated. These methods can include regional generation plans, load allocation, and energy storage device charging and discharging strategies. The target scheduling methods consider not only the local interests of each region but also the global interests, ensuring the stable operation and efficient scheduling of the entire power system. In practical applications, the target scheduling methods can be used as a decision-making reference for power system dispatchers or directly embedded into the power system's automated dispatching system to achieve intelligent dispatching and automated control.

[0153] The reinforcement learning-based partitioned virtual power plant optimization scheduling method, device, and medium provided in this disclosure improve the scheduling efficiency and flexibility of virtual power plants in multi-partition power systems by combining global and partitioned optimization scheduling strategies, optimize resource allocation, and reduce energy consumption and costs.

[0154] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0155] Based on the same inventive concept, this disclosure also provides a reinforcement learning-based partitioned virtual power plant optimization scheduling device corresponding to the reinforcement learning-based partitioned virtual power plant optimization scheduling method. Since the principle of the device in this disclosure for solving the problem is similar to the reinforcement learning-based partitioned virtual power plant optimization scheduling method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0156] Reference Figure 5 The diagram shown is a schematic of a reinforcement learning-based partitioned virtual power plant optimization scheduling device 500 provided in an embodiment of this disclosure. The device includes:

[0157] The power zoning determination module 501 is used to determine a set of power zoning based on the resource information of the virtual power plant; wherein, the set of power zoning includes wind-solar-storage coupling zoning, energy storage-electric vehicle zoning, adjustable load zoning, and gas turbine zoning;

[0158] The global model construction module 502 is used to construct a global optimization scheduling model and establish global constraints on the global optimization scheduling model based on preset market capacity limits and preset capacity feasibility limits for each zone; wherein, the global optimization scheduling model takes maximizing the total profit of the virtual power plant participating in frequency regulation ancillary services as the global objective function, and takes the frequency regulation capacity of each zone and the total frequency regulation capacity declared to the market as decision variables;

[0159] The partition model construction module 503 is used to construct a partition optimization scheduling model and establish a set of partition constraints for the partition optimization scheduling model; wherein, the partition optimization scheduling model takes minimizing the resource operating cost and instruction tracking deviation of each partition as the partition objective function, and takes the frequency regulation resource output within each partition as the decision variable;

[0160] The scheduling strategy determination module 504 is used to determine a global optimization scheduling strategy for the global optimization scheduling model based on the near-end strategy optimization algorithm, the global optimization scheduling model, and the global constraints; and to determine a partition optimization scheduling strategy for the partition optimization scheduling model based on the heterogeneous agent near-end strategy optimization algorithm, the partition optimization scheduling model, and the set of partition constraints.

[0161] The scheduling method determination module 505 is used to jointly train the near-end policy optimization algorithm and the heterogeneous agent near-end policy optimization algorithm based on the global optimization scheduling strategy and the partition optimization scheduling strategy, generate the cross-region optimal scheduling strategy, and determine the target scheduling method for each partition based on the optimal scheduling strategy.

[0162] In some possible embodiments, the global objective function is expressed as:

[0163]

[0164] in, This represents the frequency modulation capacity allocated to partition i during time period t; λ represents the total frequency regulation capacity declared to the market by the virtual power plant during time period t; cap (t) represents the capacity price for time period t; λ mile (t) represents the mileage compensation price for time period t; i represents the zone number, where number 1 represents the wind-solar-storage coupling zone, number 2 represents the energy storage-electric vehicle zone, number 3 represents the adjustable load zone, and number 4 represents the gas turbine zone; η i This is represented as the frequency modulation efficiency coefficient of partition i, obtained by calibration using historical data. C represents the operating cost of partition i in time period t; pen (t) represents the capacity deviation penalty for time period t; η dev This is expressed as a capacity deviation penalty coefficient;

[0165] The global constraint condition is expressed as follows:

[0166]

[0167] in, This represents the maximum number of applications that are subject to the market capacity limit preset for time period t. This represents the maximum adjustable capacity of partition i within time period t, which is the preset capacity feasibility limit.

[0168] In some possible embodiments, the partitioning objective function is expressed as:

[0169]

[0170] in, This represents the frequency regulation cost of frequency regulation unit j within zone i; This represents the frequency regulation output of frequency regulation unit j within time period t and partition i.

[0171] In some possible embodiments, the set of zoning constraints includes wind-solar-storage coupling zoning constraints, energy storage-electric vehicle zoning constraints, adjustable load zoning constraints, and gas turbine zoning constraints.

[0172] The zoning constraint condition for wind-solar-storage coupling is expressed as follows:

[0173]

[0174] in, P represents the frequency regulation capacity allocated to the wind-solar-storage coupling zone; wind (t) represents the real-time output of wind power; P PV (t) represents the real-time output of the photovoltaic system; P bat (t) represents the real-time output of the energy storage battery; This represents the predicted wind power output. This is represented as the predicted output of photovoltaic power. This indicates the charging power limit of the energy storage battery; The expression represents the discharge power limit of the energy storage battery; SOC(t) represents the state of charge of the energy storage battery during time period t; η ch Indicated as the charging efficiency of the energy storage battery; η dis This is expressed as the discharge efficiency of the energy storage battery; The energy storage battery charging efficiency is expressed as the time period t. State of Charge (SOC) represents the battery discharge efficiency over time period t. min State of Charge (SOC) represents the minimum state of charge of an energy storage battery. max This represents the maximum state of charge of the energy storage battery.

[0175] The energy storage-electric vehicle zoning constraint condition is expressed as follows:

[0176]

[0177] in, This represents the frequency regulation capacity allocated to the energy storage-electric vehicle zone; P represents the charging and discharging power of the energy storage battery in the energy storage-electric vehicle zone. EV,k (t) represents the charging and discharging power of electric vehicle k; N EV This represents the number of electric vehicles participating in frequency modulation. This indicates the charging power limit for electric vehicles; This refers to the discharge power limit of electric vehicles; This represents the state of charge of the battery of electric vehicle k during the departure time t; This represents the minimum state of charge requirement that the user has for electric vehicle k at the moment of departure;

[0178] The adjustable load zoning constraint is expressed as follows:

[0179]

[0180] in, This represents the frequency regulation capacity allocated to the adjustable load zone; ΔP load,m (t) represents the power adjustment amount of the adjustable load during time period t; This represents the maximum allowable adjustable power for the adjustable load. This represents the total regulating energy limit of the adjustable load over a day;

[0181] The gas turbine partitioning constraint condition is expressed as follows:

[0182]

[0183] in, P represents the frequency regulation capacity allocated to the gas turbine zone; gas (t) represents the real-time power of the gas turbine; This represents the minimum processing capacity of the gas turbine. This represents the maximum value that the gas turbine can handle. This represents the gas turbine ramp rate limit; T on (t) represents the continuous operating time of the gas turbine during time interval t; T off (t) represents the downtime of the gas turbine in time period t; This represents the minimum operating time of the gas turbine. This represents the minimum downtime of the gas turbine.

[0184] In some possible embodiments, the scheduling strategy determination module 504 is specifically used for:

[0185] Based on the near-end policy optimization algorithm, the global optimization scheduling model, and the global constraints, the Markov decision process is modeled to obtain a global decision framework.

[0186] Based on the near-end policy optimization algorithm, the global optimization scheduling model, the global constraints, and the global decision framework, a global optimization scheduling strategy is determined for the global optimization scheduling model.

[0187] The global decision-making framework is represented as follows:

[0188]

[0189]

[0190]

[0191] in, This is represented as the state space of the global optimization strategy; N represents the average state of charge of the energy storage in the wind-solar-storage coupling zone; EV (t) represents the number of EVs connected to the energy storage-electric vehicle zone during time period t; This represents the average state of charge of electric vehicles during time period t. Represented as the action space of the global optimization strategy; r ppo (t) represents the reward function of the global optimization strategy;

[0192] The scheduling strategy determination module 504 is specifically used for:

[0193] Based on the heterogeneous agent near-end strategy optimization algorithm, the partition optimization scheduling model, and the partition constraint set, the Markov decision process of each partition is modeled to obtain a partition decision framework set; wherein, the partition decision framework set includes a wind-solar-storage coupled partition decision framework, an energy storage-electric vehicle partition decision framework, an adjustable load partition decision framework, and a gas turbine partition decision framework.

[0194] Based on the heterogeneous agent near-end strategy optimization algorithm, the partitioned optimal scheduling model, the partitioned constraint set, and the partitioned decision framework set, a partitioned optimal scheduling strategy for the partitioned optimal scheduling model is determined; wherein, the partitioned optimal scheduling strategy includes a wind-solar-storage coupled partitioned optimal scheduling strategy, an energy storage-electric vehicle partitioned optimal scheduling strategy, an adjustable load partitioned optimal scheduling strategy, and a gas turbine partitioned optimal scheduling strategy.

[0195] The wind-solar-storage coupled zoning decision-making framework is represented as follows:

[0196]

[0197]

[0198]

[0199] in, This represents the state space of the wind-solar-storage coupling zone optimization scheduling strategy; This represents the action space of the wind-solar-storage coupling zone optimization scheduling strategy; c represents the reward function of the wind-solar-storage coupled zonal optimization scheduling strategy; bat Expressed as the frequency regulation cost of energy storage batteries;

[0200] The energy storage-electric vehicle zoning decision-making framework is represented as follows:

[0201]

[0202]

[0203]

[0204] in, This is represented as the state space of the energy storage-electric vehicle zone optimization scheduling strategy; This is represented as the state of charge of each electric vehicle; This represents the action space of the energy storage-electric vehicle zonal optimization scheduling strategy; This is represented as the reward function for the energy storage-electric vehicle zonal optimization scheduling strategy; The frequency regulation cost of energy storage batteries is represented as the energy storage-electric vehicle partition; c EV,k Let k be the frequency modulation cost of the electric vehicle.

[0205] The adjustable load zoning optimization decision framework is represented as follows:

[0206]

[0207]

[0208]

[0209] in, This represents the state space of an adjustable load partitioning optimization scheduling strategy. This represents the action space of the adjustable load partitioning optimization scheduling strategy. c represents the reward function of the adjustable load partitioning optimization scheduling strategy; load,m This represents the adjustment cost of the adjustable load over time period t.

[0210] The gas turbine zoning decision framework is represented as follows:

[0211]

[0212]

[0213]

[0214] in, This represents the state space of the gas turbine partitioning optimization scheduling strategy. This represents the action space of the gas turbine zonal optimization scheduling strategy. c represents the reward function of the gas turbine partitioning optimization scheduling strategy; gas Expressed as the fuel cost of a gas turbine; Expressed as carbon emission cost.

[0215] In some possible embodiments, the scheduling method determining module 505 is specifically used to perform:

[0216] Step 1: Obtain the training dataset; wherein the training dataset includes multiple training data subsets, and each training data subset includes global training data corresponding to multiple time periods within a preset time period and training data corresponding to each partition;

[0217] Step 2: For the global training data corresponding to each time period and the training data corresponding to each partition in each training data subset, perform the following operations:

[0218] (a) Generate a global action value set based on the near-end policy optimization algorithm and the global training data; wherein the global action value set includes frequency modulation capacity allocation instructions and total frequency modulation capacity declaration values ​​corresponding to different partitions;

[0219] (b) Input the frequency modulation capacity allocation instruction of the corresponding partition in the global action value set into the corresponding partition optimization scheduling strategy in the partition optimization scheduling strategy;

[0220] (c) Generate specific output action values ​​of frequency modulation resources in each partition based on the heterogeneous agent near-end policy optimization algorithm, the partition optimization scheduling strategy and the training data corresponding to each partition.

[0221] (d) Calculate the partition reward value of the heterogeneous agent near-end strategy optimization algorithm based on the reward function of the partition optimization scheduling strategy and the specific output action value of the frequency modulation resources in each partition; and calculate the global reward value of the near-end strategy optimization algorithm based on the reward function of the global optimization scheduling strategy, the frequency modulation capacity allocation instructions corresponding to different partitions and the total frequency modulation capacity declaration value.

[0222] (e) Generate a global training sub-trajectory based on the global training data corresponding to the current time period in the training data subset, the global action value set, the global reward value, and the global training data corresponding to the next time period in the training data subset, and store the global training sub-trajectory in the experience replay buffer of the near-end policy optimization algorithm; and generate a partition training sub-trajectory based on the training data corresponding to each partition in the current time period in the training data subset, the specific output action value of the frequency modulation resource in each partition, and the training data corresponding to each partition in the next time period in the training data subset, and store the partition training sub-trajectory in the experience replay buffer of the heterogeneous agent near-end policy optimization algorithm.

[0223] Step 3: Iteratively execute Step 2 until multiple training trajectory sets are generated; wherein each training trajectory set corresponds to a subset of training data, each training trajectory set includes multiple training sub-trajectories, and each training sub-trajectory corresponds to a time period in the subset of training data;

[0224] Step 4: Based on the sets of training trajectories in the experience replay buffers of the near-end policy optimization algorithm and the heterogeneous agent near-end policy optimization algorithm, perform joint parameter updates on the policy networks of the near-end policy optimization algorithm and the heterogeneous agent near-end policy optimization algorithm:

[0225] Step 5: Repeat steps 2 to 4 until the algorithm reward value converges, generating the cross-region optimal scheduling strategy.

[0226] In some possible embodiments, the scheduling method determining module 505 is specifically used to perform:

[0227] Based on any set of training trajectories in the experience replay buffer of the near-end policy optimization algorithm, the advantage value is calculated through generalized advantage estimation, and the parameters of the near-end policy optimization algorithm are updated based on the advantage value and the near-end policy loss function calculation formula.

[0228] Based on any training trajectory set in the experience replay buffer of the heterogeneous agent near-end policy optimization algorithm, the joint action advantage weights corresponding to each partition are calculated sequentially according to the update order of each partition policy. Based on the joint action advantage weights corresponding to each partition and the calculation formula of the heterogeneous agent near-end policy loss function, the policy parameters of each partition policy are updated, and the global value network is optimized simultaneously.

[0229] The formula for calculating the near-end strategy loss function is as follows:

[0230]

[0231] Among them, L clip(θ) represents the near-end policy loss value, E t Represented as the expected value of all time steps in the current subset of training data; π θ (a t |s t ) represents the state s based on the current policy parameter θ. t Generate action a t The probability of; Represented as based on the old policy parameter θ old In state s t Generate action a t The probability; ε represents the hyperparameter; The value of advantage is represented by the time interval t; clip represents the clipping function.

[0232] The formula for calculating the loss function of the heterogeneous proxy near-end strategy is as follows:

[0233]

[0234] in, Let represent the heterogeneous agent near-end policy loss value for partition i; B represents the batch size of the training data subset; and T represents the number of time steps for training sub-trajectories for each partition. This represents the state of the i-th partition based on the current policy parameter θ. Generate Actions The probability of; Represented as the updated policy parameter θ in the i-th partition new In state Generate Actions The probability of; The optimal scheduling policy for partition i is represented in state i. and actions The potential for strategic improvement.

[0235] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 6 The diagram shows the structure of a computer device 600 provided in this embodiment of the present disclosure, including a processor 601, a memory 602, and a bus 603. The memory 602 stores execution instructions and includes a main memory 6021 and an external memory 6022. The main memory 6021, also called internal memory, is used to temporarily store computational data in the processor 601 and data exchanged with external memory 6022 such as a hard disk. The processor 601 exchanges data with the external memory 6022 through the main memory 6021.

[0236] In this embodiment, the memory 602 is specifically used to store application code that executes the solution of this application, and its execution is controlled by the processor 601. That is, when the computer device 600 is running, the processor 601 communicates with the memory 602 through the bus 603, so that the processor 601 executes the application code stored in the memory 602, and then executes the method described in any of the foregoing embodiments.

[0237] The memory 602 may be, but is not limited to, random access memory, read-only memory, programmable read-only memory, erasable read-only memory, electrically erasable read-only memory, etc.

[0238] Processor 601 may be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a central processing unit (CPU), network processor, etc.; it can also be a digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.

[0239] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the computer device 600. In other embodiments of this application, the computer device 600 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0240] This disclosure also provides a computer-readable storage medium storing a computer program. When a processor executes the computer program, it performs the steps of the reinforcement learning-based partitioned virtual power plant optimal scheduling method described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0241] This disclosure also provides a computer program product carrying program code. The instructions included in the program code can be used to execute the steps of the reinforcement learning-based partitioned virtual power plant optimal scheduling method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here. The computer program product can be implemented specifically through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied as a computer storage medium; in another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK).

[0242] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0243] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional units in the various embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0244] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0245] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A partitioned virtual power plant optimization scheduling method based on reinforcement learning, characterized in that, The method comprises the following steps: determining a power partition set based on resource information of a virtual power plant; wherein the power partition set comprises a wind-solar-storage coupling partition, a storage-electric vehicle partition, an adjustable load partition, and a gas turbine partition; constructing a global optimization scheduling model, and establishing global constraint conditions of the global optimization scheduling model based on a preset market capacity limit and a preset capacity feasibility limit of each partition; wherein the global optimization scheduling model takes maximum total profit of the virtual power plant participating in frequency modulation auxiliary service as a global objective function, and takes frequency modulation capacity of each partition and total frequency modulation capacity declared to the market as decision variables; constructing a partition optimization scheduling model, and establishing a set of partition constraint conditions of the partition optimization scheduling model; wherein the partition optimization scheduling model takes a weighted sum of resource operation cost and instruction tracking deviation of each partition as a partition objective function, and takes output of frequency modulation resource in each partition as a decision variable; determining a global optimization scheduling strategy of the global optimization scheduling model based on a proximal policy optimization algorithm, the global optimization scheduling model, and the global constraint conditions; and determining a partition optimization scheduling strategy of the partition optimization scheduling model based on a heterogeneous agent proximal policy optimization algorithm, the partition optimization scheduling model, and the set of partition constraint conditions; training the proximal policy optimization algorithm and the heterogeneous agent proximal policy optimization algorithm based on the global optimization scheduling strategy and the partition optimization scheduling strategy, generating a cross-zone optimal scheduling strategy, and determining a target scheduling method of each partition based on the optimal scheduling strategy.

2. The method of claim 1, wherein the global objective function is represented as: the global constraint conditions are represented as: ; ; wherein, denotes the frequency regulation capacity allocated to the partition i at time period t; denotes the total frequency regulation capacity declared by the virtual power plant to the market at time period t; denotes the capacity price at time period t; denotes the mileage compensation price at time period t; i denotes the partition index, index 1 represents the wind-solar-storage coupled partition, index 2 represents the storage-electric vehicle partition, index 3 represents the adjustable load partition, and index 4 represents the gas turbine partition; denotes the frequency regulation efficiency coefficient of the partition i calibrated by historical data; denotes the operation cost of the partition i at time period t; denotes the capacity deviation penalty at time period t; denotes the capacity deviation penalty coefficient; 3. The method of claim 2, wherein the partition objective function is represented as: ; wherein, maximum bid quantity expressed as a market capacity limit for period t preset; maximum adjustable capacity expressed as a preset capacity feasibility limit for period t zone i.

4. The method of claim 3, wherein the set of partition constraint conditions comprises wind-solar-storage coupling partition constraint conditions, storage-electric vehicle partition constraint conditions, adjustable load partition constraint conditions, and gas turbine partition constraint conditions; the wind-solar-storage coupling partition constraint conditions are represented as: ; wherein, denotes the frequency regulation cost of frequency regulation unit j within partition i; denotes the frequency regulation output of frequency regulation unit j within partition i at time period t. the storage-electric vehicle partition constraint conditions are represented as: the adjustable load partition constraint conditions are represented as: the gas turbine partition constraint conditions are represented as: ; wherein, represents the frequency modulation capacity allocated to the wind-solar-storage coupling partition; represents the real-time output of wind power; represents the real-time output of photovoltaic power; represents the real-time output of energy storage battery; represents the predicted output of wind power; represents the predicted output of photovoltaic power; represents the charging power limit of energy storage battery; represents the discharging power limit of energy storage battery; represents the state of charge of energy storage battery at time period t; represents the charging efficiency of energy storage battery; represents the discharging efficiency of energy storage battery; represents the charging efficiency of energy storage battery at time period t; represents the discharging efficiency of energy storage battery at time period t; represents the minimum state of charge of energy storage battery; represents the maximum state of charge of energy storage battery; the determination of the global optimization scheduling strategy of the global optimization scheduling model based on the proximal policy optimization algorithm, the global optimization scheduling model, and the global constraint conditions comprises: ; wherein, represents the frequency modulation capacity allocated to the energy storage - electric vehicle partition; represents the energy storage battery charge-discharge power of the energy storage - electric vehicle partition; represents the charge-discharge power of the electric vehicle k; represents the number of electric vehicles participating in frequency modulation; represents the electric vehicle charging power limit; represents the electric vehicle discharging power limit; represents the battery state of charge of the electric vehicle k at the departure period t; represents the user's minimum state of charge requirement for the electric vehicle k at the departure time; represents the departure time of the electric vehicle k; modeling a Markov decision process based on the proximal policy optimization algorithm, the global optimization scheduling model, and the global constraint conditions to obtain a global decision framework; ; wherein, represents the frequency regulation capacity allocated to the adjustable load zone; represents the power regulation amount of the adjustable load at the time period t; represents the maximum allowable regulation power of the adjustable load; represents the total regulation energy limit of the adjustable load in a day; M represents the number of adjustable loads participating in frequency regulation; determining the global optimization scheduling strategy of the global optimization scheduling model based on the proximal policy optimization algorithm, the global optimization scheduling model, the global constraint conditions, and the global decision framework; ; wherein, represents the frequency modulation capacity allocated to the gas turbine partition; represents the real-time power of the gas turbine; represents the gas turbine handling minimum; represents the gas turbine handling maximum; represents the gas turbine ramp rate limit; represents the gas turbine continuous run time at period t; represents the gas turbine shutdown time at period t; represents the gas turbine minimum run time; represents the gas turbine minimum shutdown time.

5. The method of claim 4, wherein, the global decision framework is represented as: the determination of the partition optimization scheduling strategy of the partition optimization scheduling model based on the heterogeneous agent proximal policy optimization algorithm, the partition optimization scheduling model, and the set of partition constraint conditions comprises: ​ ​ ; ; ; wherein, state space of the global optimization strategy; average state of charge of the energy storage represented as the wind-solar-storage coupling zone; number of EVs accessing the energy storage represented as the energy storage-electric vehicle zone at time period t; average state of charge of the electric vehicle represented as the electric vehicle zone at time period t; action space of the global optimization strategy; reward function of the global optimization strategy; ​ modeling the Markov decision process of each partition based on the heterogeneous agent proximal policy optimization algorithm, the partition optimization scheduling model, and the set of partition constraint conditions, to obtain a set of partition decision frameworks; wherein the set of partition decision frameworks comprises a wind-solar-storage coupling partition decision framework, a storage-electric vehicle partition decision framework, a adjustable load partition decision framework, and a gas turbine partition decision framework; determining a partition optimization scheduling strategy for the partition optimization scheduling model based on the heterogeneous agent proximal policy optimization algorithm, the partition optimization scheduling model, the set of partition constraint conditions, and the set of partition decision frameworks; wherein the partition optimization scheduling strategy comprises a wind-solar-storage coupling partition optimization scheduling strategy, a storage-electric vehicle partition optimization scheduling strategy, an adjustable load partition optimization scheduling strategy, and a gas turbine partition optimization scheduling strategy; the wind-solar-storage coupling partition decision framework is represented as: ; ; ; wherein, a state space of the wind-solar-storage coupling partition optimization scheduling strategy; an action space of the wind-solar-storage coupling partition optimization scheduling strategy; a reward function of the wind-solar-storage coupling partition optimization scheduling strategy; a frequency modulation cost of the energy storage battery; the storage-electric vehicle partition decision framework is represented as: ; ; ; wherein, state space of the energy storage-electric vehicle partitioned optimization scheduling strategy; state of charge of the energy storage battery in time period t; state of charge of each electric vehicle; action space of the energy storage-electric vehicle partitioned optimization scheduling strategy; reward function of the energy storage-electric vehicle partitioned optimization scheduling strategy; frequency regulation cost of the energy storage battery of the energy storage-electric vehicle partitioned optimization scheduling strategy; frequency regulation cost of the electric vehicle k; the adjustable load partition optimization decision framework is represented as: ; ; ; wherein, a state space representing the adjustable load partitioning optimization scheduling strategy; an action space representing the adjustable load partitioning optimization scheduling strategy; a reward function representing the adjustable load partitioning optimization scheduling strategy; a regulation cost of the adjustable load at the time period t; the gas turbine partition decision framework is represented as: ; ; ; wherein, state space of the partitioned optimization scheduling strategy for the gas turbine; action space of the partitioned optimization scheduling strategy for the gas turbine; reward function of the partitioned optimization scheduling strategy for the gas turbine; fuel cost of the gas turbine; carbon emission cost.

6. The method of claim 5, wherein, the joint training of the proximal policy optimization algorithm and the heterogeneous agent proximal policy optimization algorithm based on the global optimization scheduling strategy and the partition optimization scheduling strategy comprises: Step 1: obtaining a training data set; wherein the training data set comprises a plurality of training data subsets, and each training data subset comprises global training data corresponding to a plurality of time periods in a preset time period and training data corresponding to each partition; Step 2: for the global training data corresponding to each time period and the training data corresponding to each partition in each training data subset, performing the following operations: (a) generating a global action value set based on the proximal policy optimization algorithm and the global training data; wherein the global action value set comprises frequency regulation capacity allocation instructions corresponding to different partitions and a total frequency regulation capacity declaration value; (b) inputting the frequency regulation capacity allocation instructions corresponding to the partition in the global action value set into each corresponding partition optimization scheduling strategy in the partition optimization scheduling strategy; (c) generating specific output action values of frequency regulation resources in each partition based on the heterogeneous agent proximal policy optimization algorithm, each partition optimization scheduling strategy, and the training data corresponding to each partition; (d) calculating a partition reward value of the heterogeneous agent proximal policy optimization algorithm according to a reward function of the partition optimization scheduling strategy and the specific output action values of the frequency regulation resources in each partition; and calculating a global reward value of the proximal policy optimization algorithm according to a reward function of the global optimization scheduling strategy, the frequency regulation capacity allocation instructions corresponding to different partitions, and the total frequency regulation capacity declaration value; (e) generating a global training sub-trajectory based on the global training data corresponding to the current time period in the training data subset, the global action value set, the global reward value, and the global training data corresponding to the next time period in the training data subset, and storing the global training sub-trajectory into the experience replay buffer of the proximal policy optimization algorithm; and generating a partition training sub-trajectory based on the training data corresponding to each partition corresponding to the current time period in the training data subset, the specific output action value of the frequency-adjustable resource in each partition, and the training data corresponding to each partition corresponding to the next time period in the training data subset, and storing the partition training sub-trajectory into the experience replay buffer of the heterogeneous agent proximal policy optimization algorithm; Step 3: iteratively performing step 2 until a plurality of training trajectory sets are generated; wherein each training trajectory set corresponds to a training data subset, and each training trajectory set comprises a plurality of training sub-trajectories, each of which corresponds to a time period in the training data subset; Step 4: performing joint parameter updating on the policy network of the proximal policy optimization algorithm and the policy network of the heterogeneous agent proximal policy optimization algorithm based on each training trajectory set in the experience replay buffer of the proximal policy optimization algorithm and the experience replay buffer of the heterogeneous agent proximal policy optimization algorithm: Step 5: repeating steps 2-4 until the algorithm reward value converges, to generate the cross-zone optimal scheduling strategy.

7. The method of claim 6, wherein, The joint parameter updating on the policy network of the proximal policy optimization algorithm and the policy network of the heterogeneous agent proximal policy optimization algorithm based on each training trajectory set in the experience replay buffer of the proximal policy optimization algorithm and the experience replay buffer of the heterogeneous agent proximal policy optimization algorithm comprises: based on any training trajectory set in the experience replay buffer of the proximal policy optimization algorithm, calculating an advantage value through generalized advantage estimation, and updating the parameters of the proximal policy optimization algorithm based on the advantage value and a proximal policy loss function calculation formula; based on any training trajectory set in the experience replay buffer of the heterogeneous agent proximal policy optimization algorithm, sequentially calculating the joint action advantage weight corresponding to each partition according to the update order of the policy of each partition, and updating the policy parameters of each partition policy based on the joint action advantage weight corresponding to each partition and a heterogeneous agent proximal policy loss function calculation formula, and synchronously optimizing the global value network; the proximal policy loss function calculation formula is represented as: ; wherein, denotes the proximal policy loss value, denotes the expected value over all time steps in the current training data subset; denotes the expected value based on the current policy parameters in state generates action with probability denotes the expected value based on the old policy parameters in state generates action with probability denotes the first hyperparameter; denotes the advantage value for time period t; denotes the clipping function; the heterogeneous agent proximal policy loss function calculation formula is represented as: ; wherein, represents the heterogeneous agent proximal policy loss value for partition i; B represents the data batch size for the training data subset; T represents the number of time steps for each partitioned training sub-trajectory; represents the probability of generating action in state under the current policy parameter ; represents the probability of generating action in state under the updated policy parameter ; represents the second hyperparameter; represents the policy improvement potential of the optimized scheduling policy for partition i in state and action .

8. A partitioned virtual power plant optimization scheduling device based on reinforcement learning, characterized in that, comprises: a power partition determination module configured to determine a power partition set based on resource information of a virtual power plant; wherein the power partition set comprises a wind-solar-storage coupling partition, a storage-electric vehicle partition, an adjustable load partition, and a gas turbine partition; The global model construction module is configured to construct a global optimization scheduling model and establish global constraint conditions for the global optimization scheduling model based on a preset market capacity limit and a preset partition capacity feasibility limit; wherein the global optimization scheduling model takes maximization of total profit of the virtual power plant participating in frequency modulation auxiliary services as a global objective function, and takes frequency modulation capacities of each partition and total frequency modulation capacities declared to the market as decision variables; The partition model construction module is configured to construct a partition optimization scheduling model and establish a set of partition constraint conditions for the partition optimization scheduling model; wherein the partition optimization scheduling model takes minimization of a weighted sum of resource operation costs and instruction tracking deviations of each partition as a partition objective function, and takes frequency modulation resource output in each partition as a decision variable; The scheduling strategy determination module is configured to determine a global optimization scheduling strategy for the global optimization scheduling model based on a proximal policy optimization algorithm, the global optimization scheduling model and the global constraint conditions, and determine a partition optimization scheduling strategy for the partition optimization scheduling model based on a heterogeneous agent proximal policy optimization algorithm, the partition optimization scheduling model and the set of partition constraint conditions; The scheduling method determination module is configured to jointly train the proximal policy optimization algorithm and the heterogeneous agent proximal policy optimization algorithm based on the global optimization scheduling strategy and the partition optimization scheduling strategy, generate a cross-zone optimal scheduling strategy, and determine a target scheduling method for each partition based on the optimal scheduling strategy.

9. A storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the method of any one of claims 1 to 7.

10. A computer device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, The processor executes the computer program to implement the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cooperative control method, device and system of micro-grid group and storage medium

    CN116706997A

  • Multi-type virtual power plant participation market regulation capacity optimization distribution method and system

    CN119904068A