Multi-station income-oriented aggregation regulation and control optimization method based on deep reinforcement learning
By using a multi-site control optimization method based on deep reinforcement learning, the problems of difficulty in quantifying the adjustable capacity of electric vehicle load dispatching and poor real-time performance of traditional dispatching methods are solved. This enables efficient and flexible control of electric vehicles and energy storage systems, improving the flexibility of the power grid and its market responsiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies cannot accurately characterize the adjustable capacity and response capability of electric vehicle loads. Traditional scheduling methods are computationally complex and have poor real-time performance, making them difficult to adapt to the dynamic changing environment of power systems. Traditional reinforcement learning methods lack the ability to explore in complex environments and are difficult to achieve efficient collaborative scheduling of multiple power stations.
A revenue-oriented aggregation regulation and optimization method based on deep reinforcement learning is adopted. By modeling the resources of photovoltaic, energy storage and charging stations, a GRU-PPO decision architecture is constructed, OU noise is introduced to enhance the exploration capability, controllable and uncontrollable resources are dynamically divided, and a real-time market response and optimization decision-making mechanism is established to achieve efficient regulation of flexible loads.
It significantly improves charging revenue and demand response revenue under the coordinated operation of multiple power stations, realizes the efficient participation of flexible loads in the optimized operation of the power grid, enhances strategy exploration capabilities and robustness, and enables rapid adaptation to electricity price fluctuations and power grid instructions.
Smart Images

Figure CN121791202A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power system dispatching, and more specifically, it relates to a multi-station revenue-oriented aggregated control optimization method based on deep reinforcement learning, and particularly to a method and system for optimizing multi-station dispatching command response and market revenue using deep reinforcement learning algorithms. Background Technology
[0002] With the advancement of the national new energy strategy and carbon neutrality goals, the electric vehicle industry has developed rapidly, and the number of electric vehicles in society has continued to rise. Driven by policies, market demand, and the maturity of business models, the number of public and commercial charging stations has shown a significant growth trend, gradually forming a multi-level charging infrastructure network covering urban transportation hubs, residential communities, commercial areas, and highway systems. Charging stations have become an important carrier connecting electric vehicle users and the power system, and their large-scale deployment is continuously enhancing their importance in the operation and regulation of the power system.
[0003] On the other hand, electric vehicles exhibit significant time redundancy during charging. Actual charging behavior shows that the time required to replenish a vehicle to the user's desired charge level at the rated charging power is typically much shorter than the physical connection time between the vehicle and the charging station. This means that users do not need to continuously charge at the rated power while their vehicles are parked to meet their travel needs, thus creating an adjustable window of opportunity in the time dimension. Based on this characteristic, load aggregators or control systems can dynamically adjust the charging power, charging time distribution, and charging curve shape within the user's permissible range, implementing operational strategies such as peak shaving and valley filling, responding to electricity price signals, participating in demand response, and ancillary services. Therefore, electric vehicle charging load possesses significant flexibility and controllability, and can serve as an important flexible load resource for participating in power system regulation and optimization.
[0004] In summary, as the scale of electric vehicles and charging stations continues to increase, they, as a new type of adjustable load resource, have important application value in improving grid flexibility, supporting the consumption of renewable energy, and improving the supply and demand balance of the electricity market.
[0005] However, the current aggregation and scheduling of flexible resources across multiple sites still faces the following challenges: 1. Insufficient understanding of resource adjustability: Electric vehicle loads exhibit significant randomness and discontinuity, while energy storage systems are limited by SOC, battery life, and operating strategies. Existing scheduling mechanisms cannot accurately characterize adjustable capacity and response capabilities.
[0006] 2. The market and operating environment are highly volatile: day-ahead prices, real-time electricity prices, dispatch instructions, weather forecasts and user charging behavior are all highly uncertain, and control methods based on static or rule-driven approaches are difficult to adapt to the rapidly changing environment.
[0007] 3. Insufficient intelligence in scheduling strategies: Traditional optimization methods are computationally complex and have poor real-time performance, making it impossible to adjust strategies in a timely manner for dynamic scenarios, resulting in large response deviations and low returns.
[0008] 4. Traditional reinforcement learning methods are based on Markov decision processes for modeling. Under this framework, the PPO algorithm has limited ability to express temporal correlations and is difficult to effectively capture long-term dependencies across states. In addition, PPO often uses simple Gaussian noise as an exploration mechanism, which is difficult to maintain stable and continuous exploration in complex environments with strong coupling and high uncertainty. It is prone to problems such as insufficient exploration or getting trapped in local optima.
[0009] Therefore, there is an urgent need for an intelligent decision-making method that can automatically quantify the flexibility of power stations, respond to grid dispatching needs, and improve aggregate response capabilities. Summary of the Invention
[0010] To address the problem that traditional charging station control strategies struggle to fully exploit the adjustable potential of electric vehicle charging loads, this invention proposes a revenue-oriented aggregation control optimization method for multiple charging stations based on deep reinforcement learning. This method quantifies and dynamically utilizes the adjustability of flexible load resources such as electric vehicles, and automatically learns market bidding strategies and station power allocation schemes through deep reinforcement learning algorithms, thereby forming an optimized control strategy that adapts to changes in electricity prices and demand response signals. This method significantly improves charging revenue and demand response revenue under multi-station collaborative operation, enabling the efficient participation of flexible loads in grid optimization.
[0011] To achieve the above objectives, the present invention adopts the following solution: A multi-site revenue-oriented aggregation regulation optimization method based on deep reinforcement learning includes the following steps: Step 1: Model the resources such as vehicles, energy storage and photovoltaic power output of photovoltaic and energy storage charging stations, and at the same time model and depict the clearing process and clearing results of the electricity spot market and demand response market; Step 2: Dynamically divide the internal resources of the photovoltaic-storage-charging station into controllable and uncontrollable resources according to their participation in regulation. Controllable resources include energy storage equipment and adjustable loads of electric vehicles, while uncontrollable resources include users' forced charging needs and photovoltaic resources with natural output. Then, quantify the adjustable range of the station's power. Step 3: Construct a GRU-based temporal policy network and train it using the PPO algorithm; introduce OU (Ornstein–Uhlenbeck) noise at the continuous action output to enhance the exploration capability, thereby forming a GRU-PPO decision architecture with time memory characteristics and a stable random perturbation mechanism.
[0012] Step 4: Based on the quantitative results of the internal resource flexibility of the photovoltaic-storage-charging station, construct a real-time market response and optimization decision-making mechanism. Participate in the real-time spot market and ancillary service market through load adjustment, energy storage charging and discharging strategies, or demand response behavior. Establish an intelligent agent optimization objective function with the goal of improving charging revenue and demand response revenue. Solve the function using an improved GRU-PPO algorithm to realize the decomposition and execution of real-time market declaration decisions and market clearing results.
[0013] In a further optimization of this technical solution, step 1 involves mathematical modeling of the photovoltaic-storage-charging station. S11, the quantity of energy stored is denoted as... Each energy storage moment status Recorded as , in, These represent the maximum capacity and rated charging power of the energy storage battery, respectively. They represent the times respectively. The state of charge and charging power of the energy storage battery; S12, finally, the optical storage charging station will be set up at the designated time. The average output of photovoltaic power is denoted as ; S13. Divide the vehicles within the photovoltaic-storage-charging station into plug-and-charge, orderly charging, and V2G types, with the following numbers of vehicles respectively. The real-time market clearing cycle is The day is divided into 288 time periods. ,in The arrival of the electric vehicle is taken as the triggering event, assuming that within a time period... The number of charging vehicles connected to the charging piles is , No. Vehicle arrival information Recorded as , Indicates the vehicle's rated charging power. Indicates the vehicle's rated discharge power. Indicates time Maximum power, Indicates time Minimum power; Indicates the type of vehicle (plug-and-charge, sequential charging, V2G). This indicates the initial state of charge (SOC) of the vehicle when it arrives at the depot. This represents the expected SOC (State of Charge) at the vehicle's arrival and departure times. Indicates the maximum capacity of the electric vehicle battery; Indicates the arrival time of the vehicle, when the vehicle is at If the vehicle arrives within the specified time period, then the arrival time will be... ; Indicates the expected departure time of the vehicle, when the vehicle is expected to... If the vehicle leaves within a certain time period, then the departure time is... It will be during the time period. Arrived inside The initial information of the vehicle is denoted as the joint state. ; S14. Because the power dispatching agency conducts real-time electricity market clearing 15 minutes before the actual system operation, and performs rolling optimization every five minutes, assuming that within the time period... The beginning moment, that is The real-time electricity spot market has completed clearing and been released. Time period clearing power and clearing electricity price Therefore , , Assuming participation in day-ahead demand response, the total demand response volume and response time for that day are known at midnight each day. Assuming a total of [number missing] demand responses are required throughout the day... Segment demand response, the first The start time, end time, and required response quantity of the segment are denoted as follows: Assuming ,So ; ,So ,when At that time, it was believed The time period demand response scheduling instruction is ,in , , , express The start and end times of the time period for responding to the demand response instruction. express The amount of demand response instructions to be fulfilled within a time period. express Baseline load for the time period.
[0014] The technical solution is further optimized, and step 2 is as follows: The resources within a charging station include vehicles currently charging, energy storage, and photovoltaic power generation systems, at any given time. At the start, the power station divides all internal resources into controllable and uncontrollable loads. Energy storage is always classified as a controllable load, while plug-and-charge vehicles are always classified as uncontrollable loads. To promote local consumption of photovoltaic power generation, the photovoltaic power generation system is always classified as an uncontrollable load. Orderly charging and V2G vehicles are classified as controllable and uncontrollable loads based on the relationship between their maximum rechargeable capacity and demanded power. When the maximum rechargeable capacity is greater than the demanded power, the vehicle is considered a controllable load; when the maximum rechargeable capacity is less than the demanded power, the vehicle is considered an uncontrollable load. Assume the current time is... Vehicle information is The information of the charging pile connected to it is The maximum rechargeable capacity is defined as follows: , The required electricity is defined as follows: , Assuming the vehicle is under an uncontrollable load, when The vehicle's maximum and minimum power are equal to its rated charging power, that is... , when The vehicle's maximum and minimum power are , Assuming the vehicle is under a controllable load, when The vehicle's maximum power is equal to its rated charging power, that is , , when The vehicle's maximum power is , , The power boundary of a vehicle is defined as ,in , , Energy storage is a controllable load, when The maximum power of energy storage is equal to the rated charging power, that is... , when The maximum power of energy storage is , The minimum power of energy storage is , The power boundary of a vehicle is defined as ,in , The upper and lower bounds of the power capacity of a power station are equal to the sum of the upper and lower bounds of the power capacity of all resources within the power station. The calculation formula is as follows: , .
[0015] In a further optimization of this technical solution, step 3 involves constructing a GRU-PPO decision architecture that possesses time memory characteristics and a stationary random perturbation mechanism. S31. Construct a temporal feature extraction module to extract the enhanced agent's ability to observe states related to time. The temporal feature extraction module consists of two stacked GRU neural networks. The input of the first GRU neural network, GRU_1, is the aggregate quotient state vector. , in, This indicates the minimum power combined state of the station complex. This indicates the combined maximum power state of the power station complex. Indicates the time period of the station consortium The lower limit of the power of vehicles connected to charging piles. Indicates the time period of the station consortium The upper limit of the power of vehicles connected to charging piles. This indicates the combined photovoltaic power output status of the power plant consortium, and the real-time market data within the specified time period. The clearing price is . Through the coordinated action of the gate control unit of GRU_1, local time-series modeling is performed on low-order features such as the aggregator's load at different times and real-time market node electricity prices. After nonlinear transformation by the Tanh activation function, the output of GRU_1 is... The input is fed into the second layer of the GRU neural network, GRU_2, to provide a higher-dimensional abstract representation of the aggregate quotient's temporal information state. The output of GRU_2 is denoted as... ; S32. Construct a state value assessment and decision-making module, wherein the value assessment module consists of a Critic network, which assesses the current aggregator state and outputs the state value. This allows for precise quantitative analysis of the advantages and disadvantages of different strategies; while the strategy network (Actor) utilizes the extracted feature vectors to optimize for maximizing total profit. Generates real-time market declaration capacity and total power output of charging stations, denoted as... The decision module output is the policy network Actor output superimposed with OU noise. , ,in Indicates the rate of mean regression. This represents the long-term mean. Indicates the amplitude of the noise. Indicates volatility. This represents a standard normal distribution.
[0016] This technical solution is further optimized, and step 4 is as follows: S41, Record the total load aggregator The first photovoltaic energy storage and charging station, the first The upper and lower limits of the power of a photovoltaic energy storage and charging station are denoted as follows: Time period The real-time market-clearing electricity price is within [the specified range]. The change in the clearing electricity price is , Real-time market clearing power is denoted as The demand response quantity is denoted as The start time of the demand response is recorded as The end time of the demand response is recorded as Time period The baseline load for demand response within the region is denoted as The input of the neural network GRU_1 include , , , , in, This indicates the minimum power combined state of the station complex. This indicates the combined maximum power state of the power station complex. Indicates the time period of the station consortium The lower limit of the power of vehicles connected to charging piles. Indicates the time period of the station consortium The upper limit of the power of vehicles connected to charging piles. This indicates the combined photovoltaic power output status of the power station consortium.
[0017] S42, Decision Network Module Output Recorded as ,in Indicates time period Real-time market clearing power express A solar-powered energy storage and charging station during a certain time period The net power between the internal substation and the power grid is positive when it indicates that the power grid is discharging to the substation, and negative when it indicates that the substation is discharging to the power grid. S44. Record the state of the intelligent agent. Take action After reaching the state The single-step transfer reward is , in, Indicates time period The profit from charging services within the area Indicates time period Arbitrage opportunities arising from the price difference between the day-ahead market and the real-time market. Indicates time period Profits obtained from responding to internal demand and scheduling instructions; The charging service price is as follows: Real-time market in time period The clearing price is The clearing power is The formula for calculating the profit of charging services is as follows: , Record the market time period of the day The clearing price is The clearing power is Price arbitrage over time period The calculation formula is as follows: , S45. The agent's optimization objective is to maximize the cumulative reward within a finite time. Due to the randomness of the vehicle's arrival time, departure time, and expected SOC, it naturally also possesses a certain degree of randomness. Ignoring the randomness of the initial state, starting from a fixed initial state... Start Passing The cumulative benefits of the step transfer are , If we consider the randomness of the initial state, then the agent goes through... The cumulative benefits of the step transfer are Optimal strategy To maximize The strategy at the time, that is The agent uses the PPO algorithm to solve the problem.
[0018] Further optimizations to this technical solution include S5, which, based on a trained GRU-PPO decision-making architecture, performs online rolling optimization and control across multiple sites. Specifically, this includes the following steps: S51, Assuming we are currently in a time period The beginning time, that is station To perform ultra-short-term photovoltaic (PV) processing forecasting, the predicted average PV output for the next time period is denoted as... The previous time period The forecast of the average photovoltaic output for the current time period at the start time is as follows: Meanwhile, load aggregators predict the time period for each station. Initial arrival information of vehicles connected to the charging station from start to end time The load aggregator calculates the change in the upper and lower limits of the power for each station based on S21, denoted as... ; S52. Load aggregators collect real-time market clearing information from the previous time period, including real-time market clearing prices. Changes in clearing electricity prices Real-time market clearing power is denoted as Demand response clearing information, including the quantity of demand response offers, is recorded as follows: The start time of the demand response is recorded as follows: The end time of the demand response is recorded as Time period The baseline load for demand response within the region is denoted as Record the market clearing status as ; S53, The load aggregator calculates the upper and lower bounds of the power for each substation based on S21, and the load aggregator status is... ; S54. The time-series feature extraction module reads the upper and lower bounds of the power output of the power plant consortium, the photovoltaic output prediction status, and the market clearing status. It abstracts and represents the time-series state information of the agent and inputs it into the Actor network module. The Actor network module then applies the current strategy... Output Action With OU noise superimposed, the actual actions performed by the aggregator are as follows: ,in Indicates real-time market reporting power. express A solar-powered energy storage and charging station during a certain time period Net power between the internal substation and the power grid; S55. When entering the start of the next time period, the load aggregator repeats steps S51 to S54, and with the goal of maximizing the revenue of the power station consortium, it continues to participate in real-time market reporting and power station control decisions to achieve rolling optimization operation.
[0019] Compared with the prior art, the present invention has the following beneficial effects: 1. Maximizing profitability: By coordinating electricity spot market prices and demand response signals and through multi-site collaborative dispatch, the overall economic benefits of charging operation and ancillary services have been significantly improved.
[0020] 2. Superior algorithm performance: The GRU-PPO algorithm, which introduces OU noise, effectively captures the time-dependent characteristics of the optical storage and charging system, enhances the strategy exploration capability and robustness, and avoids getting trapped in local optima.
[0021] 3. Unlocking Load Potential: By dynamically dividing controllable and uncontrollable resources and accurately quantifying the station's adjustment capacity, the flexibility of electric vehicles and energy storage is fully released while ensuring users' rigid needs are met.
[0022] 4. Real-time and precise response: A closed-loop mechanism has been established from market clearing to station execution, which can quickly adapt to electricity price fluctuations and grid instructions, and realize efficient interaction between flexible loads and the grid. Attached Figure Description
[0023] Figure 1 The flowchart shows a multi-site revenue-oriented aggregation regulation optimization method based on deep reinforcement learning. Detailed Implementation
[0024] To explain in detail the technical content, structural features, objectives, and effects of the technical solution, the following description is provided in conjunction with specific embodiments and accompanying drawings.
[0025] Please see Figure 1 The diagram shown illustrates the process of a multi-site revenue-oriented aggregated regulation and optimization method based on deep reinforcement learning. This invention proposes a multi-site revenue-oriented aggregated regulation and optimization method based on deep reinforcement learning, which includes the following steps: S1: Model the resources of photovoltaic-storage-charging stations, such as vehicles, energy storage, and photovoltaic output. This involves mathematical modeling of the photovoltaic-storage-charging stations, as well as modeling and depicting the clearing process and results of the electricity spot market and demand response market.
[0026] S11, the quantity of energy stored is denoted as... Each energy storage moment status Recorded as
[0027] in, These represent the maximum capacity and rated charging power of the energy storage battery, respectively. They represent the times respectively. The state of charge and charging power of the energy storage battery.
[0028] S12, finally, the optical storage charging station will be set up at the designated time. The average output of photovoltaic power is denoted as .
[0029] S13. Divide the vehicles within the photovoltaic-storage-charging station into plug-and-charge, orderly charging, and V2G types, with the following numbers of vehicles respectively. The real-time market clearing cycle is The day is divided into 288 time periods. ,in The arrival of the electric vehicle is taken as the triggering event, assuming a time period. The number of charging vehicles connected to the charging piles is , No. Vehicle arrival information Recorded as
[0030] Indicates the vehicle's rated charging power. Indicates the vehicle's rated discharge power. Indicates time Maximum power, Indicates time Minimum power; Indicates the type of vehicle (plug-and-charge, sequential charging, V2G). This indicates the initial state of charge (SOC) of the vehicle when it arrives at the depot. This represents the expected SOC (State of Charge) at the vehicle's arrival and departure times. Indicates the maximum capacity of the electric vehicle battery; Indicates the arrival time of the vehicle, when the vehicle is at If the vehicle arrives within the specified time period, then the arrival time will be... ; Indicates the expected departure time of the vehicle, when the vehicle is expected to... If the vehicle leaves within a certain time period, then the departure time is... It will be during the time period. Arrived inside The initial information of the vehicle is denoted as the joint state. .
[0031] S14. Because the power dispatching agency conducts real-time electricity market clearing 15 minutes before the actual system operation, and performs rolling optimization every five minutes. Assuming that during the time period... The beginning moment, that is The real-time electricity spot market has completed clearing and been released. Time period clearing power and clearing electricity price Therefore
[0032]
[0033] Assuming participation in day-ahead demand response, the total demand response volume and response time for that day are known at midnight each day. Assuming a total of [number missing] demand responses are required throughout the day... Segment demand response, the first The start time, end time, and required response quantity of the segment are denoted as follows: Assuming ,So ; ,So .when At that time, it was believed The time period demand response scheduling instruction is ,in
[0034]
[0035]
[0036] express The start and end times of the time period for responding to the demand response instruction. express The amount of demand response instructions to be fulfilled within a time period. express Baseline load for the time period.
[0037] S2: Dynamically divide the internal resources of the photovoltaic-storage-charging station into controllable and uncontrollable resources according to their participation in regulation. Controllable resources include energy storage equipment and adjustable loads of electric vehicles, while uncontrollable resources include users' forced charging needs and naturally generated photovoltaic resources. Then, quantify the adjustable range of the station's power.
[0038] S21, Assuming that during the time period Inner Arrival information for vehicles connected to charging stations:
[0039] S22. Resources within the charging station include vehicles currently charging, energy storage, and photovoltaic power generation systems, in each time period. At the start, the power station classifies all internal resources into controllable and uncontrollable loads. Energy storage is always classified as a controllable load, while plug-and-charge vehicles are always classified as uncontrollable loads. Plug-and-charge vehicles are considered uncontrollable loads. To promote local consumption of photovoltaic power generation, the photovoltaic power generation system is always classified as an uncontrollable load. Ordered charging and V2G vehicles are classified as controllable and uncontrollable loads based on the relationship between their maximum rechargeable capacity and demanded power. When the maximum rechargeable capacity is greater than the demanded power, the vehicle is a controllable load; when the maximum rechargeable capacity is less than the demanded power, the vehicle is an uncontrollable load. Assume the current time is... Vehicle information is The information of the charging pile connected to it is The maximum rechargeable capacity is defined as follows:
[0040] The required electricity is defined as follows:
[0041] Assuming the vehicle is under an uncontrollable load, when The vehicle's maximum and minimum power are equal to its rated charging power, that is...
[0042] when The vehicle's maximum and minimum power are
[0043] Assuming the vehicle is under a controllable load, when The vehicle's maximum power is equal to its rated charging power, that is
[0044]
[0045] when The vehicle's maximum power is
[0046]
[0047] The power boundary of a vehicle is defined as ,in
[0048]
[0049] Energy storage is a controllable load, when The maximum power of energy storage is equal to the rated charging power, that is...
[0050] when The maximum power of energy storage is
[0051] The minimum power of energy storage is
[0052] The power boundary of a vehicle is defined as ,in
[0053] The upper and lower bounds of the power capacity of a power station are equal to the sum of the upper and lower bounds of the power capacity of all resources within the power station (electric vehicles, energy storage, and photovoltaics). The calculation formula is as follows:
[0054]
[0055] in, This indicates the assembly of all vehicles within the station. This refers to the collection of all energy storage facilities within the site.
[0056] S3: Construct the GRU-PPO decision architecture. The aggregator's state is abstracted and extracted by the time-series feature extraction module and then input into the decision output module. The Actor network output is superimposed with OU (Ornstein–Uhlenbeck) noise and outputs the real-time market declaration capacity and total power of the station. The Critic network outputs the state value.
[0057] A GRU-based temporal policy network is constructed and trained using the PPO algorithm. OU noise is introduced at the continuous action output to enhance the exploration capability, thus forming a GRU-PPO decision architecture with time memory characteristics and a stable random perturbation mechanism.
[0058] S31. Construct a temporal feature extraction module to extract the enhanced agent's ability to observe states related to time. The temporal feature extraction module consists of two stacked GRU neural networks. The input of the first GRU neural network, GRU_1, is the aggregate quotient state vector.
[0059] in, This indicates the minimum power combined state of the station complex. This indicates the combined maximum power state of the power station complex. Indicates the time period of the station consortium The lower limit of the power of vehicles connected to charging piles. Indicates the time period of the station consortium The upper limit of the power of vehicles connected to charging piles. This indicates the combined photovoltaic power output status of the power plant consortium, and the real-time market data within the specified time period. The clearing price is . Through the coordinated action of the gate control unit of GRU_1, local time-series modeling is performed on low-order characteristics such as the aggregator's load at different times and real-time market node electricity prices. After nonlinear transformation by the Tanh activation function, the output of GRU_1 is... The input is fed into the second layer of the GRU neural network, GRU_2, to provide a higher-dimensional abstract representation of the aggregate quotient's temporal information state. The output of GRU_2 is denoted as... .
[0060] S32. Construct a state value assessment and decision-making module, wherein the value assessment module consists of a Critic network, which assesses the current aggregator state and outputs the state value. This allows for precise quantitative analysis of the advantages and disadvantages of different strategies; while the strategy network (Actor) utilizes the extracted feature vectors to optimize for maximizing total profit. Generates real-time market declaration capacity and total power output of charging stations, denoted as... The decision module output is the policy network Actor output superimposed with OU noise.
[0061]
[0062] in Indicates the rate of mean regression. This represents the long-term mean. Indicates the amplitude of the noise. Indicates volatility. This represents a standard normal distribution.
[0063] S4: Based on the quantitative results of the internal resource flexibility of photovoltaic, energy storage and charging stations, a real-time market response and optimization decision-making mechanism is constructed. It participates in the real-time spot market and ancillary service market through load adjustment, energy storage charging and discharging strategies or demand response behavior. An intelligent agent optimization objective function is established with the goal of improving charging revenue and demand response revenue. The improved GRU-PPO algorithm is used to solve the problem, so as to realize the decomposition and execution of real-time market declaration decision and market clearing results.
[0064] Based on the quantitative results of the load flexibility assessment of the power station, with the optimization objective of maximizing the total revenue of the power station (including charging service and demand response revenue), the power station participates in real-time market bidding to realize the decomposition of real-time market clearing results and demand response dispatch instructions.
[0065] S41, Record the total load aggregator The first photovoltaic energy storage and charging station, the first The upper and lower limits of the power of a photovoltaic energy storage and charging station are denoted as follows: Time period The real-time market-clearing electricity price is within [the specified range]. The change in the clearing electricity price is
[0066] Real-time market clearing power is denoted as The demand response quantity is denoted as The start time of the demand response is recorded as The end time of the demand response is recorded as Time period The baseline load for demand response within the region is denoted as Input to the GRU_1 neural network include
[0067]
[0068]
[0069]
[0070] in, This indicates the minimum power combined state of the station complex. This indicates the combined maximum power state of the power station complex. Indicates the time period of the station consortium The lower limit of the power of vehicles connected to charging piles. Indicates the time period of the station consortium The upper limit of the power of vehicles connected to charging piles. This indicates the combined photovoltaic power output status of the power station consortium.
[0071] S42, Decision Network Module Output Recorded as
[0072] in Indicates time period Real-time market clearing power express A solar-powered energy storage and charging station during a certain time period The net power between the internal substation and the power grid is positive when it indicates that the power grid is discharging to the substation, and negative when it indicates that the substation is discharging to the power grid.
[0073] S44. Record the state of the intelligent agent. Take action After reaching the state The single-step transfer reward is
[0074] in, Indicates time period The profit from charging services within the area Indicates time period Arbitrage opportunities arising from the price difference between the day-ahead market and the real-time market. Indicates time period Profits obtained from internal response demand scheduling instructions.
[0075] The charging service price is as follows: Real-time market in time period The clearing price is The clearing power is The formula for calculating the profit of charging services is as follows:
[0076] Record the market time period of the day The clearing price is The clearing power is Price arbitrage over time period The calculation formula is as follows:
[0077] S45. The agent's optimization objective is to maximize the cumulative reward within a finite time. Due to the randomness of the vehicle's arrival time, departure time, and expected SOC, it naturally also possesses a certain degree of randomness. Ignoring the randomness of the initial state, starting from a fixed initial state... Start Passing The cumulative benefits of the step transfer are
[0078] If we consider the randomness of the initial state, then the agent goes through... The cumulative benefits of the step transfer are
[0079] Optimal Strategy To maximize The strategy at the time, that is The agent uses the PPO algorithm to solve the problem.
[0080] S5: Based on the trained GRU-PPO decision architecture, perform online rolling optimization and control of multiple sites.
[0081] S51, Assuming we are currently in a time period The beginning time, that is station To perform ultra-short-term photovoltaic (PV) processing forecasting, the predicted average PV output for the next time period is denoted as... The previous time period The forecast of the average photovoltaic output for the current time period at the start time is as follows: Meanwhile, load aggregators predict the time period for each station. Initial arrival information of vehicles connected to the charging station from start to end time The load aggregator calculates the change in the upper and lower limits of the power for each station based on S21, denoted as... .
[0082] S52. Load aggregators collect real-time market clearing information from the previous time period, including real-time market clearing prices. Changes in clearing electricity prices Real-time market clearing power is denoted as Demand response clearing information, including the quantity of demand response offers, is recorded as follows: The start time of the demand response is recorded as follows: The end time of the demand response is recorded as Time period The baseline load for demand response within the region is denoted as Record the market clearing status as .
[0083] S53. The load aggregator calculates the upper and lower bounds of power for each power station based on S21. The load aggregator status (including the upper and lower bounds of power for the power station consortium and the photovoltaic output forecast) is as follows: .
[0084] S54. The time-series feature extraction module reads the upper and lower bounds of the power output of the power plant consortium, the photovoltaic output prediction status, and the market clearing status. It abstracts and represents the time-series state information of the agent and inputs it into the Actor network module. The Actor network module then applies the current strategy... Output Action With OU noise superimposed, the actual actions performed by the aggregator are as follows:
[0085] in Indicates real-time market reporting power. express A solar-powered energy storage and charging station during a certain time period Net power between the internal substation and the power grid.
[0086] S55. When entering the start of the next time period, the load aggregator repeats steps S51 to S54, and with the goal of maximizing the revenue of the power station consortium, it continues to participate in real-time market reporting and power station control decisions to achieve rolling optimization operation.
Claims
1. A multi-site revenue-oriented aggregation regulation and optimization method based on deep reinforcement learning, characterized in that, Including the following methods: Step 1: Model the vehicles, energy storage and photovoltaic power output resources of the photovoltaic-storage-charging station, and at the same time model and depict the clearing process and results of the electricity spot market and demand response market; Step 2: Dynamically divide the internal resources of the photovoltaic-storage-charging station into controllable and uncontrollable resources according to their participation in regulation. Controllable resources include energy storage equipment and adjustable loads of electric vehicles, while uncontrollable resources include users' forced charging demand and photovoltaic resources with natural output. Then, quantify the adjustable range of the station's power. Step 3: Construct a GRU-based temporal policy network and train it using the PPO algorithm; introduce OU noise at the continuous action output to enhance the exploration capability, thereby forming a GRU-PPO decision architecture with time memory characteristics and a stable random perturbation mechanism. Step 4: Based on the quantitative results of the internal resource flexibility of the photovoltaic-storage-charging station, construct a real-time market response and optimization decision-making mechanism. Participate in the real-time spot market and ancillary service market through load adjustment, energy storage charging and discharging strategies, or demand response behavior. Establish an intelligent agent optimization objective function with the goal of improving charging revenue and demand response revenue. Solve the function using an improved GRU-PPO algorithm to realize the decomposition and execution of real-time market declaration decisions and market clearing results.
2. The multi-site revenue-oriented aggregation regulation and optimization method based on deep reinforcement learning as described in claim 1, characterized in that: In step 1, mathematical modeling of the photovoltaic storage and charging station is performed. S11, the quantity of energy stored is denoted as... Each energy storage moment status Recorded as , in, These represent the maximum capacity and rated charging power of the energy storage battery, respectively. They represent the times respectively. The state of charge and charging power of the energy storage battery; S12, Finally, the optical storage charging station will be set up at the designated time. The average output of photovoltaic power is denoted as ; S13. Divide the vehicles within the photovoltaic-storage-charging station into plug-and-charge, orderly charging, and V2G types, with the following numbers of vehicles respectively. The real-time market clearing cycle is The day is divided into 288 time periods. ,in The arrival of the electric vehicle is taken as the triggering event, assuming that within a time period The number of charging vehicles connected to the charging piles is , No. Vehicle arrival information Recorded as , Indicates the vehicle's rated charging power. Indicates the vehicle's rated discharge power. Indicates time Maximum power, Indicates time Minimum power; Indicates the type of vehicle (plug-and-charge, sequential charging, V2G). This indicates the initial state of charge of the vehicle when it arrives at the depot. This indicates the expected state of charge of the vehicle at the arrival and departure times. Indicates the maximum capacity of the electric vehicle battery; Indicates the arrival time of the vehicle, when the vehicle is at If the vehicle arrives within the specified time period, then the arrival time will be... ; Indicates the expected departure time of the vehicle, when the vehicle is expected to... If the vehicle leaves within a certain time period, then the departure time is... During the time period Arrived inside The initial information of the vehicle is denoted as the joint state. ; S14. Because the power dispatching agency conducts real-time electricity market clearing 15 minutes before the actual system operation, and performs rolling optimization every five minutes, assuming that within the time period... The beginning moment, that is The real-time electricity spot market has completed clearing and been released. Time period clearing power and clearing electricity price Therefore , , Assuming participation in day-ahead demand response, the total demand response volume and response time for that day are known at midnight each day. Assuming a total of [number missing] demand responses are required throughout the day... Segment demand response, the first The start time, end time, and required response quantity of the segment are denoted as follows: Assuming ,So ; ,So ,when At that time, it was believed The time period demand response scheduling instruction is ,in , , , express The start and end times of the time period for responding to the demand response instruction. express The amount of demand response instructions to be fulfilled within a time period. express Baseline load for the time period.
3. The multi-site revenue-oriented aggregation regulation and optimization method based on deep reinforcement learning as described in claim 2, characterized in that: Step 2 is described in detail below: The resources within a charging station include vehicles currently charging, energy storage, and photovoltaic power generation systems, at any given time. At the start, the power station divides all internal resources into controllable and uncontrollable loads. Energy storage is always classified as a controllable load, while plug-and-charge vehicles are always classified as uncontrollable loads. To promote local consumption of photovoltaic power generation, the photovoltaic power generation system is always classified as an uncontrollable load. Orderly charging and V2G vehicles are classified as controllable and uncontrollable loads based on the relationship between their maximum rechargeable capacity and demanded power. When the maximum rechargeable capacity is greater than the demanded power, the vehicle is considered a controllable load; when the maximum rechargeable capacity is less than the demanded power, the vehicle is considered an uncontrollable load. Assume the current time is... Vehicle information is The information of the charging pile connected to it is The maximum rechargeable capacity is defined as follows: The required electricity is defined as follows: , Assuming the vehicle is under an uncontrollable load, when The vehicle's maximum and minimum power are equal to its rated charging power, that is... ,when The vehicle's maximum and minimum power are , Assuming the vehicle is under a controllable load, when The vehicle's maximum power is equal to its rated charging power, that is , , when The vehicle's maximum power is , , The power boundary of a vehicle is defined as ,in , , Energy storage is a controllable load, when The maximum power of energy storage is equal to the rated charging power, that is... , when The maximum power of energy storage is , The minimum power of energy storage is , The power boundary of a vehicle is defined as ,in , The upper and lower bounds of the power capacity of a power station are equal to the sum of the upper and lower bounds of the power capacity of all resources within the power station. The calculation formula is as follows: , .
4. The multi-site revenue-oriented aggregation regulation and optimization method based on deep reinforcement learning as described in claim 3, characterized in that: In step 3, a GRU-PPO decision architecture with time memory characteristics and a stationary random perturbation mechanism is constructed. S31. Construct a temporal feature extraction module to extract the enhanced agent's ability to observe states related to time. The temporal feature extraction module consists of two stacked GRU neural networks. The input of the first GRU neural network, GRU_1, is the aggregate quotient state vector. , in, This indicates the minimum power combined state of the station complex. This indicates the combined maximum power state of the power station complex. Indicates the time period of the station consortium The lower limit of the power of vehicles connected to charging piles. Indicates the time period of the station consortium The upper limit of the power of vehicles connected to charging piles. This indicates the combined photovoltaic power output status of the power plant consortium, and the real-time market data within the specified time period. The clearing price is , Through the coordinated action of the gate control unit of GRU_1, local time-series modeling is performed on the low-order characteristics of aggregator load at different times and real-time market node electricity prices. After nonlinear transformation by the Tanh activation function, the output of GRU_1 is... The input is fed into the second layer of the GRU neural network, GRU_2, to provide a higher-dimensional abstract representation of the aggregate quotient's temporal information state. The output of GRU_2 is denoted as... ; S32. Construct a state value assessment and decision-making module, wherein the value assessment module consists of a Critic network, which assesses the current aggregator state and outputs the state value. This allows for precise quantitative analysis of the advantages and disadvantages of different strategies; while the strategy network (Actor) utilizes the extracted feature vectors to optimize for maximizing total profit. Generates real-time market declaration capacity and total power output of charging stations, denoted as... The decision module output is the policy network Actor output superimposed with OU noise. , ,in Indicates the rate of mean regression. This represents the long-term mean. Indicates the amplitude of the noise. Indicates volatility. This represents a standard normal distribution.
5. The multi-site revenue-oriented aggregation regulation and optimization method based on deep reinforcement learning as described in claim 4, characterized in that: Step 4 is described in detail below: S41, Record the total load aggregator The first photovoltaic energy storage and charging station, the first The upper and lower limits of the power of a photovoltaic energy storage and charging station are denoted as follows: Time period The real-time market-clearing electricity price is [price missing]. The change in the clearing electricity price is Real-time market clearing power is denoted as The demand response quantity is denoted as The start time of the demand response is recorded as The end time of the demand response is recorded as Time period The baseline load for demand response within the region is denoted as The input of the neural network GRU_1 include , , , , in, This indicates the minimum power combined state of the station complex. This indicates the combined maximum power state of the power station complex. Indicates the time period of the station consortium The lower limit of the power of vehicles connected to charging piles. Indicates the time period of the station consortium The upper limit of the power of vehicles connected to charging piles. This indicates the combined photovoltaic power output status of the power plant consortium; S42, Decision Network Module Output Recorded as ,in Indicates time period Real-time market clearing power express A solar-powered energy storage and charging station during a certain time period The net power between the internal substation and the power grid is positive when it indicates that the power grid is discharging to the substation, and negative when it indicates that the substation is discharging to the power grid. S44. Record the state of the intelligent agent. Take action After reaching the state The single-step transfer reward is ,in, Indicates time period The profit from charging services within the area Indicates time period Arbitrage opportunities arising from the price difference between the day-ahead market and the real-time market. Indicates time period Profits obtained from responding to internal demand and scheduling instructions; The charging service price is as follows: Real-time market in time period The clearing price is The clearing power is The formula for calculating the profit of charging services is as follows: , Record the market time period of the day The clearing price is The clearing power is Price arbitrage over time period The calculation formula is as follows: , S45. The agent's optimization objective is to maximize the cumulative reward within a finite time. Due to the randomness of the vehicle's arrival time, departure time, and expected SOC, it naturally also possesses a certain degree of randomness. Ignoring the randomness of the initial state, starting from a fixed initial state... Start Passing The cumulative benefits of the step transfer are , If we consider the randomness of the initial state, then the agent goes through... The cumulative benefits of the step transfer are Optimal strategy To maximize The strategy at the time, that is The agent uses the PPO algorithm to solve the problem.
6. The multi-site revenue-oriented aggregation regulation and optimization method based on deep reinforcement learning according to claim 5, characterized in that: It also includes S5, which, based on a pre-trained GRU-PPO decision architecture, performs online rolling optimization and control across multiple sites, specifically including the following steps. S51, Assuming we are currently in a time period The beginning time, i.e. station To perform ultra-short-term photovoltaic (PV) processing forecasting, the predicted average PV output for the next time period is denoted as... The previous time period The forecast of the average photovoltaic output for the current time period at the start time is as follows: Meanwhile, load aggregators predict the time period for each station. Initial arrival information of vehicles connected to the charging station from start to end time The load aggregator calculates the change in the upper and lower limits of the power for each station based on S21, denoted as... ; S52. Load aggregators collect real-time market clearing information from the previous time period, including real-time market clearing prices. Changes in clearing electricity prices Real-time market clearing power is denoted as Demand response clearing information, including the quantity of demand response offers, is recorded as follows: The start time of the demand response is recorded as follows: The end time of the demand response is recorded as Time period The baseline load for demand response within the region is denoted as Record the market clearing status as ; S53, The load aggregator calculates the upper and lower bounds of the power for each substation based on S21, and the load aggregator status is... ; S54. The time-series feature extraction module reads the upper and lower bounds of the power output of the power plant consortium, the photovoltaic output prediction status, and the market clearing status. It abstracts and represents the time-series state information of the agent and inputs it into the Actor network module. The Actor network module then applies the current strategy... Output Action With OU noise superimposed, the actual actions performed by the aggregator are as follows: ,in Indicates real-time market reporting power. express A solar-powered energy storage and charging station during a certain time period Net power between the internal substation and the power grid; S55. When entering the start of the next time period, the load aggregator repeats steps S51 to S54, and with the goal of maximizing the revenue of the power station consortium, it continues to participate in real-time market reporting and power station control decisions to achieve rolling optimization operation.