Electric vehicle aggregator electric energy market optimization decision-making system and method based on reinforcement learning

Through the electric vehicle aggregator electricity energy market optimization decision-making system based on reinforcement learning, the problems of uncertainty modeling, multi-time scale coupling and low computational efficiency of electric vehicle aggregators in the electricity energy market are solved, and the maximum profit of electric vehicle aggregators and the stable operation of the power system are achieved.

CN120689088AInactive Publication Date: 2025-09-23NORTH CHINA ELECTRIC POWER UNIV +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510742674.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies have problems in insufficient uncertainty modeling, insufficient multi-time scale coupling, insufficient dynamic adaptability and low computational efficiency in optimizing decision-making for electric vehicle aggregators participating in the electricity energy market, resulting in poor robustness and profitability of the optimization results.

Method used

An electric vehicle aggregator electricity energy market optimization decision-making system based on reinforcement learning is adopted, including a data acquisition module, a multi-time scale modeling module, a reinforcement learning optimization module, a market clearing feedback module and a control execution module. Through the deep deterministic policy gradient algorithm and the Lagrange multiplier coupling mechanism, the energy bid amount, the capacity bid amount and the charging and discharging power allocation strategy are dynamically adjusted, and a two-layer optimization model is constructed to optimize the bidding strategy and power allocation.

Benefits of technology

It significantly improves the profits and dynamic adaptability of electric vehicle aggregators, enhances the robustness and computational efficiency of multi-time scale coupling, meets the rapid response requirements of the real-time market, and achieves maximum profits and stable operation of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689088A_ABST
    Figure CN120689088A_ABST
Patent Text Reader

Abstract

The invention relates to an electric vehicle aggregator electric energy market optimization decision-making system and method based on reinforcement learning. The system comprises a data acquisition module, a multi-time scale modeling module, a reinforcement learning optimization module, a market clearing feedback module and a control execution module. The data acquisition module collects the state and market information of the electric vehicle; the multi-time scale modeling module constructs a double-layer optimization model of a day-ahead market and a real-time market, and realizes time scale coupling through a Lagrange multiplier; the reinforcement learning optimization module adopts a depth deterministic strategy gradient algorithm to dynamically adjust the energy bidding amount, the capacity bidding amount and the charging and discharging power distribution strategy; the market clearing feedback module calculates bid winning electric quantity and node marginal electricity price; and the control execution module issues a charging and discharging instruction. According to the method, the bidding strategy is optimized through reinforcement learning, electricity price uncertainty and multi-time scale coupling are comprehensively considered, the income of an electric vehicle aggregator is remarkably improved, and meanwhile, the adjustment requirement of a power system is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of battery management technology, and in particular to an electric energy market optimization decision-making system and method for electric vehicle aggregators based on reinforcement learning. Background Art

[0002] With the transformation of the global energy mix and the rapid development of renewable energy, the operational model of the power system is undergoing profound changes. Electric vehicles (EVs), as a key distributed energy resource, can not only participate in the power market as loads but also provide regulation services to the grid through vehicle-to-grid (V2G) technology. Electric vehicle aggregators (EVAs), acting as intermediaries between EVs and the power market, can participate in the energy and ancillary services markets by aggregating the charging and discharging capacity of large numbers of EVs, thereby maximizing revenue while providing flexibility to the power system.

[0003] In existing technologies, the optimization decisions of electric vehicle aggregators participating in the electricity energy market are typically based on traditional optimization methods, such as linear programming (LP), mixed integer linear programming (MILP), or stochastic programming (SP). These methods optimize energy bids and capacity bids by constructing mathematical models, combining electricity price forecasts and electric vehicle status information. For example, existing technologies often use a two-layer optimization framework of the day-ahead market and the real-time market to determine the bidding strategy of electric vehicle aggregators, and allocate the winning electricity through a market clearing algorithm. However, these methods have the following problems:

[0004] Inadequate uncertainty modeling: Existing technologies typically rely on simple scenario generation methods, such as probability distributions based on historical data, to address the uncertainties of electricity prices and regulation signals. However, fluctuations in electricity prices and regulation signals are highly random and nonlinear, making it difficult for traditional methods to accurately capture these uncertainties, resulting in insufficient robustness in optimization results. For example, when electricity prices fluctuate significantly, traditional methods may result in reduced profits or an inability to meet regulation requirements.

[0005] Inadequate multi-timescale coupling: EV aggregators need to make decisions in both the day-ahead and real-time markets, requiring multi-timescale optimization at both hourly and minute levels. Existing technologies typically separate day-ahead and real-time market decision-making, lacking an effective coupling mechanism, resulting in poor overall optimization. For example, bidding strategies in the day-ahead market may not adapt to rapid changes in the real-time market, impacting EV charging and discharging scheduling and revenue.

[0006] Insufficient dynamic adaptability: Traditional optimization methods are typically based on static models and struggle to adapt to the dynamic changes in the electricity market and electric vehicle status. For example, the state of charge (SOC) of electric vehicles, user charging demand, and market prices change in real time. Traditional methods are unable to adjust bidding and power allocation strategies in real time, resulting in optimization results that do not match actual demand.

[0007] Low computational efficiency: When faced with large-scale EV aggregation scenarios, traditional optimization methods (such as MILP) require solving complex mathematical models, resulting in high computational complexity and long solution times, making it difficult to meet the rapid response requirements of the real-time market. For example, when an aggregator manages hundreds of EVs, traditional methods can take several minutes to solve, failing to meet the four-minute scheduling requirements of the real-time market.

[0008] To address these issues, reinforcement learning (RL) technology has garnered widespread attention in recent years in the field of electricity market optimization. RL learns optimal policies through interaction with the environment and is capable of handling high-dimensional state spaces and dynamically changing decision-making problems. For example, the Deep Deterministic Policy Gradient (DDPG) algorithm has been applied to dynamic bidding and power scheduling in electricity markets. However, existing RL methods still have the following shortcomings in optimizing the energy market for electric vehicle aggregators:

[0009] Imperfect state and action space design: Existing methods generally do not fully consider the diverse states of electric vehicles (such as SOC, charging demand, battery degradation cost) and the complexity of market information (such as electricity prices, regulation signals), resulting in overly simple state space and action space designs that are difficult to capture the dynamic characteristics of the system.

[0010] Incomplete reward function design: Existing reinforcement learning methods often only consider energy revenue and capacity adjustment revenue when designing reward functions, ignoring factors such as mileage adjustment revenue, deployment adjustment revenue, and battery degradation costs. This leads to incomplete optimization objectives and difficulty in maximizing benefits.

[0011] Multi-time-scale optimization is not effectively combined: Existing methods lack an effective coupling mechanism for the day-ahead market and the real-time market in multi-time-scale optimization, making it difficult to achieve global optimization.

[0012] To this end, we proposed an electric vehicle aggregator electricity energy market optimization decision-making system based on reinforcement learning to solve the above problems. Summary of the Invention

[0013] The purpose of the present invention is to solve the existing technical problems raised in the above background technology and provide an electric energy market optimization decision-making system for electric vehicle aggregators based on reinforcement learning.

[0014] The above-mentioned object of the present invention is achieved as follows: an electric vehicle aggregator electric energy market optimization decision system based on reinforcement learning, comprising:

[0015] A data acquisition module is used to collect status information of electric vehicles and market information of the electric energy market, wherein the status information of the electric vehicles includes the remaining battery power, the charge and discharge time, and the charge and discharge power limit; and the market information includes the real-time electricity price, the adjustment capacity price, and the adjustment signal;

[0016] A multi-timescale modeling module is used to construct a two-tier optimization model for the multi-timescale electricity energy market. The upper-tier model uses reinforcement learning to optimize the bidding strategies of electric vehicle aggregators in the day-ahead and real-time markets, while the lower-tier model allocates winning bids for electric vehicles based on a market-clearing algorithm. The two-tier optimization model achieves timescale coupling through Lagrange multipliers.

[0017] a reinforcement learning optimization module configured to dynamically adjust energy bids, capacity bids, and charging and discharging power allocation strategies based on the electric vehicle status information and market information using a deep deterministic policy gradient algorithm;

[0018] A market clearing feedback module is used to calculate the market clearing results, including the winning electricity volume and the node marginal electricity price, through mixed integer linear programming and Karush-Kuhn-Tucker conditions, and input the results as environmental feedback to the reinforcement learning optimization module to update the strategy parameters;

[0019] The control execution module is used to issue charging and discharging instructions to electric vehicles based on the optimized bidding strategy and power allocation strategy.

[0020] As a preferred technical solution of the present invention, the multi-time scale modeling module includes:

[0021] The day-ahead time scale submodule is used to build a bidding decision model based on hours, predict electricity prices and regulation signal scenarios, and determine the energy bid amount and regulation capacity bid amount of electric vehicle aggregators. The day-ahead time scale is expressed as: T = {t begin ,t begin +

[0022] Δt,t begin +2Δt,...,t final},Δt=0.5h / 1h / 2h...;

[0023] The real-time time scale submodule is used to respond to the regulation signal in real time in units of minutes and adjust the charging and discharging power distribution of the electric vehicle. The real-time time scale is expressed as: Δt = 2min / 3min / 4min...;

[0024] The time-scale coupling unit couples the bidding decision of the day-ahead market with the power allocation of the real-time market by transferring the dual variables through Lagrange multipliers to optimize the overall decision-making process.

[0025] As a preferred technical solution of the present invention, the state space of the reinforcement learning optimization module includes: the current time, which is used to distinguish different time steps; market information, including the real-time price of the electric energy market, the price of the regulation capacity market, and the regulation signal issued by the power system operator; the electric vehicle state, including the current state of charge of each electric vehicle, charging demand, and battery degradation cost; wherein, the state space is expressed as:

[0026]

[0027] is time, λ e n is the electricity price, λ r eg is the adjustment capacity price, ξ is the adjustment signal, (SOC(t)) is the state of charge, (Demand(t)) is the charging demand, C de grade(t) is the battery degradation cost, and (N) is the number of electric vehicles.

[0028] As a preferred technical solution of the present invention, the action space of the reinforcement learning optimization module includes: bidding strategies, including the energy bid amount submitted by the electric vehicle aggregator to the electric energy market and the regulation capacity bid amount submitted to the regulation capacity market at time t; power allocation strategies, including the charge and discharge power of each electric vehicle at time t;

[0029] The action space is expressed as:

[0030]

[0031] Where (P) is the energy bid amount, (R) is the regulation capacity bid amount, and (p(t)) is the charging and discharging power of the i-th vehicle;

[0032] As a preferred technical solution of the present invention, the reward function of the reinforcement learning optimization module includes: energy revenue, calculated based on the real-time electricity price and the energy bid amount; adjustment capacity revenue, calculated based on the adjustment capacity price and the adjustment capacity bid amount; adjustment mileage revenue, calculated based on the adjustment mileage price and the flexibility of the electric vehicle in responding to the adjustment signal; adjustment deployment revenue, calculated based on the power of the electric vehicle in responding to the adjustment signal; battery degradation cost, calculated based on the discharge power and unit degradation cost. The reward function is expressed as:

[0033]

[0034] Among them, energy income is:

[0035] The revenue from regulating capacity is:

[0036] Adjusted mileage earnings are:

[0037] Adjusted deployment income is:

[0038] The battery degradation cost is:

[0039] Among them, π s is the scene probability, P s ,t en is the energy bid amount, λ t is the electricity price, P s ,t cap To adjust the capacity bidding amount, R s To adjust the capacity, is the performance score, P s ,t mil To adjust the mileage bid amount, To adjust the mileage coefficient, δ t To regulate the signal, p i ,s,k deg is the degradation cost coefficient, P i ,s,k dis is the discharge power.

[0040] As a preferred technical solution of the present invention, the multi-time scale modeling module further includes an uncertainty modeling unit for generating electricity price scenarios and regulation signal scenarios through stochastic programming and calculating the joint scenario probability, wherein the joint scenario probability is expressed as: π s,t =π s π t,s Among them, π s is the probability of scene s, π t,s is the conditional probability of time t under scenario s;

[0041] Among them, the scene set is represented as: s∈S={low,medium,high},π s =0.3,0.5,0.2.

[0042] As a preferred technical solution of the present invention, the control execution module satisfies the following constraints when issuing charge and discharge instructions:

[0043] State of charge constraints:

[0044] Charge and discharge power limit: almost electric rate to the end.

[0045] Power balance constraints:

[0046] Adjust capacity maintenance time constraint:

[0047] Among them, the initial and final state of charge constraints are:

[0048]

[0049] Among them, SOC i , t is the power state, and are the charging and discharging efficiencies, p i ,t ch and p i ,t dis are charging and discharging power respectively, Δt is the time interval, u i , t is the state variable, p i ,t dis(ch) is the upper power limit, δ s To adjust the signal.

[0050] As a preferred technical solution of the present invention, the market clearing feedback module optimizes the winning bid volume and node marginal electricity price with the goal of maximizing social welfare, where the objective function is expressed as:

[0051]

[0052] The market clearing constraint is:

[0053] The optimization results are fed back to the reinforcement learning optimization module through an iterative coupling mechanism, wherein the iterative coupling mechanism includes: the upper-level reinforcement learning outputs the bidding strategy; the lower-level market clearing calculates the winning electricity volume and node marginal electricity price; and the winning electricity volume and node marginal electricity price are used as environmental feedback to update the reinforcement learning strategy.

[0054] As a preferred technical solution of the present invention, the training process of the reinforcement learning optimization module includes: initializing the Actor network and the Critic network; observing the current market state and generating a bidding strategy; calculating the market clearing results and storing the experience samples; updating the network parameters to minimize the temporal difference error, wherein the temporal difference error is: TDerror=E[(Q target -Q(s,a|θ Q )) 2 ] and soft-update the target network, where the target network is updated as: θ Q′ ←τθQ +(1-τ)θ Q′ θ μ′ ←τθ μ +(1-τ)θ μ′ .

[0055] Determine whether the cumulative reward fluctuation is less than the set threshold to determine training convergence, where the cumulative reward is expressed as: Cumulative reward = ∑ t =1 T γ t-1 r t Where γ is the discount factor, r t is the reward at time t, and τ is the soft update coefficient.

[0056] As a preferred technical solution of the present invention, the system also includes a performance analysis module for evaluating the convergence, computational efficiency and robustness of the system, wherein: Convergence is analyzed by cumulative reward curve, and cumulative reward is expressed as: cumulative reward = ∑ t =1 T γ t-1 r t The computational efficiency is evaluated by comparing the solution time of the reinforcement learning algorithm and the traditional stochastic programming. The robustness is analyzed by testing the revenue fluctuation amplitude in the perturbed electricity price scenario.

[0057] A reinforcement learning-based electric vehicle aggregator energy market optimization decision-making method is characterized by comprising the following steps:

[0058] Collect information on the remaining battery capacity, charge and discharge power limits, and real-time market electricity prices and capacity adjustment prices of electric vehicles;

[0059] A two-layer optimization model is constructed, in which the upper layer model dynamically optimizes the bidding strategies in the day-ahead and real-time markets through reinforcement learning; the lower layer model allocates the winning electricity volume through a market clearing algorithm; and Lagrange multipliers are used to couple the time-scale interaction of the two models.

[0060] Through a deep deterministic policy gradient algorithm, the energy bid amount, capacity bid amount, and charging and discharging power allocation strategy are dynamically adjusted based on real-time power consumption and node marginal electricity price feedback;

[0061] The market clearing results are solved based on mixed integer linear programming, and the winning electricity volume and node marginal electricity price are fed back into the reinforcement learning strategy update;

[0062] According to the optimized strategy, charging and discharging instructions are issued to electric vehicles to meet the state of charge constraints and power balance constraints.

[0063] As a preferred technical solution of the present invention, the multi-time scale modeling includes:

[0064] Day-ahead timescale modeling: forecasting electricity prices and regulation signal scenarios on an hourly basis, and determining the energy and regulation capacity bids of EV aggregators;

[0065] Real-time timescale modeling: responding to regulation signals in real time on a minute-by-minute basis to adjust the charging and discharging power distribution of electric vehicles;

[0066] Time scale coupling: The dual variables are transferred via Lagrange multipliers to couple the bidding decision of the day-ahead market with the power allocation in the real-time market to optimize the overall decision-making process.

[0067] As a preferred technical solution of the present invention, in the reinforcement learning optimization, the state space includes: current time, market information and electric vehicle status, the action space includes: bidding strategy and power allocation strategy, and the reward function includes: energy income, capacity adjustment income, mileage adjustment income, deployment adjustment income and battery degradation cost.

[0068] As a preferred technical solution of the present invention, it also includes uncertainty modeling: generating electricity price scenarios and regulation signal scenarios through random programming, and calculating the probability of the joint scenario.

[0069] As a preferred technical solution of the present invention, the control execution satisfies the following constraints: power state constraint; charge and discharge power constraint; power balance constraint; regulation capacity maintenance time constraint; initial and final power state constraint.

[0070] As a preferred technical solution of the present invention, the market clearing feedback aims to maximize social welfare, optimizes the winning electricity volume and node marginal electricity price, and feeds back the optimization results to the reinforcement learning optimization module through an iterative coupling mechanism.

[0071] As a preferred technical solution of the present invention, the training process of the reinforcement learning optimization module includes: initializing the Actor network and the Critic network; observing the current market state and generating a bidding strategy; calculating the market clearing results and storing experience samples; updating the network parameters to minimize the temporal difference error and soft-updating the target network; judging whether the cumulative reward fluctuation is less than a set threshold to determine the training convergence.

[0072] As a preferred technical solution of the present invention, it also includes performance analysis: evaluating the convergence, computational efficiency and robustness of the system.

[0073] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the electric energy market optimization decision-making method for electric vehicle aggregators based on reinforcement learning.

[0074] Compared with the prior art, the present invention has the following beneficial effects:

[0075] The present invention provides an electric energy market optimization decision-making system for electric vehicle aggregators based on reinforcement learning, which first has significant advantages in terms of revenue optimization and dynamic adaptability. By introducing the deep deterministic policy gradient (DDPG) algorithm, the system can dynamically adjust the energy bid amount, the adjustment capacity bid amount, and the charging and discharging power allocation strategy according to the electric vehicle status information (such as SOC, charging demand, battery degradation cost) and market information (such as real-time electricity price, regulation signal). Compared with traditional optimization methods, the present invention can adapt to the dynamic changes of market prices and electric vehicle status in real time, significantly improving the revenue of electric vehicle aggregators. For example, in an embodiment, the revenue after system optimization is 15% higher than that of the traditional method, while comprehensively considering energy income, adjustment capacity income, adjustment mileage income, adjustment deployment income and battery degradation cost, ensuring the comprehensiveness and sustainability of revenue maximization.

[0076] Secondly, the present invention demonstrates superior performance in multi-time-scale coupling and uncertainty modeling. Through the multi-time-scale modeling module, the system constructs a two-layer optimization model of the day-ahead market and the real-time market, and realizes time-scale coupling by transferring dual variables through Lagrange multipliers, thereby optimizing the overall decision-making process. At the same time, the uncertainty modeling unit generates electricity prices and regulation signal scenarios through stochastic programming, and calculates the probability of joint scenarios to effectively cope with market uncertainty fluctuations. Compared with the shortcomings of the existing technology that lack multi-time-scale coupling and uncertainty modeling, the present invention can maintain robustness in complex market environments. For example, in a scenario where the electricity price fluctuates by ±10%, the fluctuation range of system revenue is less than 5%, which significantly improves the stability and reliability of decision-making.

[0077] Finally, the present invention has important application value in terms of computing efficiency and power system support. The system optimizes bidding strategies and power allocation through reinforcement learning algorithms, which significantly reduces the computational complexity compared to traditional methods (such as MILP) and meets the rapid response requirements of the real-time market. For example, the average solution time of the reinforcement learning algorithm in the embodiment is 10 seconds, while the traditional method requires 30 seconds. In addition, by optimizing the charging and discharging scheduling of electric vehicles, the system provides flexible adjustment capabilities for the power system and supports the stable operation of the power system. The SOC of electric vehicles is always kept within a safe range during the optimization process, ensuring user needs and equipment life, while providing effective support for renewable energy consumption and grid balance, and has broad practical application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0079] Figure 1 : Cumulative reward curve during reinforcement learning training;

[0080] Figure 2 :Analysis of revenue fluctuations under electricity price disturbances;

[0081] Figure 3 : Electric vehicle SOC and charge and discharge power change over time;

[0082] Figure 4 : Energy state after polymerization changes with time;

[0083] Figure 5 : The graph of capacity bidding after aggregation changing with time;

[0084] Figure 6 : Comparison chart of day-ahead market regulation capacity bidding;

[0085] Figure 7 : Comparison chart of benefits between optimized learning model and traditional model. DETAILED DESCRIPTION

[0086] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0087] The following is combined with Figure 1-Figure 7 , the specific implementation methods of the present invention are described in detail.

[0088] The present invention provides a reinforcement learning-based electric vehicle aggregator energy market optimization decision-making system. The technical solution of the present invention is described in detail below with reference to specific embodiments. This embodiment is intended to illustrate the implementation of the present invention and does not limit the scope of protection of the present invention in any way.

[0089] This embodiment provides a reinforcement learning-based electric vehicle aggregator energy market optimization decision-making system, specifically comprising the following modules: a data acquisition module, a multi-timescale modeling module, a reinforcement learning optimization module, a market clearing feedback module, and a control execution module. The system aims to optimize the bidding and power allocation strategies of electric vehicle aggregators in the energy market through reinforcement learning algorithms to maximize revenue and achieve stable operation of the power system.

[0090] The data acquisition module is used to collect the status information of electric vehicles and market information of the electric energy market. The status information of electric vehicles includes the remaining battery capacity (SOC), the charge and discharge time, and the charge and discharge power limit. The market information includes the real-time electricity price, the adjustment capacity price, and the adjustment signal. For example, in this embodiment, the system obtains the SOC data of each vehicle through the communication interface with the electric vehicle, and obtains the real-time electricity price λ through the power market trading platform. t , adjust capacity prices and the regulation signal r t . These data provide the basis for subsequent optimization decisions.

[0091] The multi-timescale modeling module is used to construct a two-layer optimization model for the multi-timescale electricity energy market. The upper layer model optimizes the bidding strategy of electric vehicle aggregators in the day-ahead market and the real-time market based on reinforcement learning, and the lower layer model allocates the winning bid electricity of electric vehicles based on the market clearing algorithm. In this embodiment, the day-ahead timescale submodule predicts electricity prices and regulation signal scenarios in hourly units. The time scale is expressed as: T = {t begin ,t begin +Δt,t begin +2Δt,...,t final},Δt=0.5h / 1h / 2h...;

[0092] For example, if Δt = 1 hour and the time range is 24 hours, that is, T = {1, 2, ..., 24}. The real-time time scale submodule responds to the real-time regulation signal in minutes, and the time scale is expressed as: Δt = 2 minutes / 3 minutes / 4 minutes...; in this embodiment, Δt = 4 minutes, which means that each hour contains 15 real-time time steps.

[0093] The time scale coupling unit couples the bidding decision of the day-ahead market with the power allocation of the real-time market by transferring the dual variable through the Lagrange multiplier. For example, the energy bidding amount of the day-ahead market is realized by the Lagrange multiplier λ and the dual variable μ. With real-time market power allocation p i,t Coordinated optimization.

[0094] The reinforcement learning optimization module uses the deep deterministic policy gradient (DDPG) algorithm to dynamically adjust the energy bid amount, adjust the capacity bid amount, and the charging and discharging power allocation strategy of electric vehicles based on electric vehicle status information and market information to maximize the profits of electric vehicle aggregators.

[0095] The market clearing feedback module calculates the market clearing results, including the winning electricity volume and node marginal electricity price, through mixed integer linear programming (MILP) and Karush-Kuhn-Tucker (KKT) conditions, and feeds the results back to the reinforcement learning optimization module to update the bidding strategy.

[0096] The control execution module issues charging and discharging instructions to the electric vehicles according to the optimized bidding strategy and power allocation strategy. For example, the charging and discharging power instruction p is sent to each electric vehicle through the communication interface with the electric vehicle charging pile. i,t , in order to achieve the optimal allocation of power resources and the stable operation of the power system.

[0097] The day-ahead timescale submodule predicts electricity prices and regulation signal scenarios on an hourly basis to determine the energy bids of EV aggregators. and adjust capacity bids In this embodiment, the system generates electricity price scenarios and regulation signal scenarios based on historical data. The scenario set is represented as:

[0098] s∈S={low,medium,high},π s =0.3,0.5,0.2

[0099] Among them, π s is the scenario probability, for example, the probability of a low electricity price scenario is 0.3, the probability of a medium electricity price scenario is 0.5, and the probability of a high electricity price scenario is 0.2.

[0100] The real-time time scale submodule responds to the regulation signal in real time in minutes and adjusts the charging and discharging power distribution of electric vehicles. For example, in the real-time market, the system responds to the regulation signal r issued by the power system operator in real time. t , adjust the charging and discharging power p of each electric vehicle i,t , to meet the regulation needs.

[0101] The uncertainty modeling unit generates electricity price scenarios and regulation signal scenarios through random programming and calculates the joint scenario probability: π s,t =π s π t,s ;

[0102] For example, for scenario s = medium, time t = 1, the conditional probability π t , s=0.6, then the joint scene probability is π s,t=0.5·0.6=0.3.

[0103] The state space of the reinforcement learning optimization module includes the current time, market information, and electric vehicle status, which can be expressed as:

[0104]

[0105] In this embodiment, it is assumed that there are 100 electric vehicles (N=100), time t=1, and electricity price λ e n=0.5 yuan / kWh, the adjustment capacity price is:

[0106] λ reg =0.2 yuan / kW, the regulation signal ξ=0.8, and the SOC, charging demand and battery degradation cost of each vehicle are obtained in real time through the data acquisition module.

[0107] The action space includes bidding strategies and power allocation strategies, which can be expressed as follows:

[0108]

[0109] For example, at time t = 1, the system output energy bid amount P = 500kW, the regulation capacity bid amount R = 200kW, and the charging and discharging power of each electric vehicle p i,t Calculated by DDPG algorithm.

[0110] The reward function includes energy revenue, capacity adjustment revenue, mileage adjustment revenue, deployment adjustment revenue, and battery degradation cost, which can be expressed as:

[0111]

[0112] Of which: Energy income:

[0113] For example, P s ,t en =500kW,λ t = 0.5 yuan / kWh, then INCOME s ,t en =500·0.5=250 yuan.

[0114] Adjustment capacity income:

[0115] For example, P s ,t cap =200kW, λ t = 0.2 yuan / kW, then Yuan.

[0116] Adjusting mileage earnings:

[0117] Reconciling deployment revenue:

[0118] Battery degradation costs:

[0119] For example, the degradation cost coefficient p i ,s,k deg =0.01 yuan / kWh, discharge power P i ,s,k dis =10kW, time interval Δt=0.1h, then the degradation cost of a single vehicle is 0.01·10·0.1=0.01 yuan.

[0120] The training process of the reinforcement learning optimization module includes the following steps: initializing the Actor network and the Critic network, setting the initial parameters θ Q and θ μ ;

[0121] Observe the current market status t , generate action a through the Actor network t =μ(s t |θ μ )+N t ;

[0122] Calculate market clearing results and store experience samples;

[0123] New network parameters to minimize the temporal difference error: TDerror = E[(Q target -Q(s,a|θ Q )) 2 ];

[0124] Among them, the target value Q target =r t +γQ(s t+1 ,μ(s t+1 |θ μ′ )|θ Q′ );

[0125] Soft update target network: θ Q′ ←τθ Q +(1-τ)θ Q′ θ μ′ ←τθ μ +(1-τ)θ μ′ ;

[0126] In this embodiment, the soft update coefficient τ is set to 0.001, and the discount factor γ is set to 0.99.

[0127] Determine whether the cumulative reward fluctuation is less than the set threshold (for example, the cumulative reward fluctuation is less than 1%) to determine training convergence:

[0128] The market clearing feedback module aims to maximize social welfare and optimize the winning electricity volume and node marginal electricity price. The objective function is:

[0129] Constraints include:

[0130] In this embodiment, the system calculates the winning electricity volume and node marginal electricity price through a MILP solver (such as Gurobi), and feeds the results back to the reinforcement learning optimization module through an iterative coupling mechanism.

[0131] The control execution module issues charging and discharging instructions to electric vehicles based on the optimized bidding strategy and power allocation strategy, meeting the following constraints:

[0132] State of charge constraints:

[0133] For example, charging efficiency Discharge efficiency The time interval Δt=0.1h, the SOC range is ([0.2,0.8]).

[0134] Charge and discharge power constraints:

[0135] Power balance constraints:

[0136] For example, set P t =500kW,R t =200kW,δ s =0.8, the system optimizes and calculates the charging and discharging power of each vehicle and Ensure power balance.

[0137] Adjust capacity maintenance time constraint:

[0138]

[0139] The initial and final state of charge constraints are:

[0140]

[0141] For example, the initial charge E i ,t init =20kWh, time interval Δt=0.1h, charging efficiency ηch =0.95, discharge efficiency η dis =0.9, the system optimizes and calculates the charging and discharging power of each vehicle to ensure that the power state meets the constraints.

[0142] The performance analysis module is used to evaluate the convergence, computational efficiency, and robustness of the system:

[0143] Convergence analysis through cumulative reward curve

[0144] For example, the cumulative reward stabilizes after 1,000 iterations, with a fluctuation of less than 1%.

[0145] Computational efficiency is evaluated by comparing the solution time of the reinforcement learning algorithm with that of the traditional randomized programming algorithm. For example, the average solution time of the reinforcement learning algorithm is 10 seconds, while the traditional randomized programming algorithm takes 30 seconds.

[0146] Robustness is tested by analyzing the revenue fluctuation range in a perturbation price scenario. For example, in a scenario where the price fluctuates by ±10%, the system revenue fluctuation range is less than 5%.

[0147] Through the system of this embodiment, electric vehicle aggregators can maximize their profits in the electric energy market while meeting the regulation needs of the power system. For example, in one experiment, the optimized system increased profits by 15% compared to traditional methods, while maintaining the SOC of electric vehicles within a safe range, ensuring the stability of the power system.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An electric vehicle aggregator energy market optimization decision-making system based on reinforcement learning, characterized by: include: A data acquisition module is used to collect status information of electric vehicles and market information of the electric energy market, wherein the status information of the electric vehicles includes the remaining battery power, the charge and discharge time, and the charge and discharge power limit; and the market information includes the real-time electricity price, the adjustment capacity price, and the adjustment signal; A multi-timescale modeling module is used to construct a two-tier optimization model for the multi-timescale electricity energy market. The upper-tier model uses reinforcement learning to optimize the bidding strategies of electric vehicle aggregators in the day-ahead and real-time markets, while the lower-tier model allocates winning bids for electric vehicles based on a market-clearing algorithm. The two-tier optimization model achieves timescale coupling through Lagrange multipliers. a reinforcement learning optimization module configured to dynamically adjust energy bids, capacity bids, and charging and discharging power allocation strategies based on the electric vehicle status information and market information using a deep deterministic policy gradient algorithm; A market clearing feedback module is used to calculate the market clearing results, including the winning electricity volume and the node marginal electricity price, through mixed integer linear programming and Karush-Kuhn-Tucker conditions, and input the results as environmental feedback to the reinforcement learning optimization module to update the strategy parameters; The control execution module is used to issue charging and discharging instructions to electric vehicles based on the optimized bidding strategy and power allocation strategy.

2. The electric vehicle aggregator electric energy market optimization decision system based on reinforcement learning according to claim 1 is characterized in that: The multi-time scale modeling module includes: The day-ahead time scale submodule is used to build a bidding decision model based on hours, predict electricity prices and regulation signal scenarios, and determine the energy bid amount and regulation capacity bid amount of electric vehicle aggregators. The day-ahead time scale is expressed as: T = {t begin ,t begin + Δt,t begin +2Δt,...,t final },Δt=0.5h / 1h / 2h...; The real-time time scale submodule is used to respond to the regulation signal in real time in units of minutes and adjust the charging and discharging power distribution of the electric vehicle. The real-time time scale is expressed as: Δt = 2min / 3min / 4min...; The time-scale coupling unit couples the bidding decision of the day-ahead market with the power allocation of the real-time market by transferring the dual variables through Lagrange multipliers to optimize the overall decision-making process.

3. The electric vehicle aggregator electric energy market optimization decision system based on reinforcement learning according to claim 1 is characterized in that: The state space of the reinforcement learning optimization module includes: the current time, which is used to distinguish different time steps; market information, including the real-time price of the electric energy market, the price of the regulation capacity market, and the regulation signal issued by the power system operator; and the electric vehicle state, including the current state of charge of each electric vehicle, charging demand, and battery degradation cost. The state space is represented as follows: Where (t) is time, λ e n is the electricity price, λ r eg is the adjustment capacity price, ξ is the adjustment signal, (SOC(t)) is the state of charge, (Demand(t)) is the charging demand, C degrade (t) is the battery degradation cost, and (N) is the number of electric vehicles.

4. The electric vehicle aggregator power energy market optimization decision system based on reinforcement learning according to claim 1 is characterized in that: The action space of the reinforcement learning optimization module includes: bidding strategies, including the energy bid amount submitted by the electric vehicle aggregator to the electric energy market and the regulation capacity bid amount submitted to the regulation capacity market at time t; power allocation strategies, including the charge and discharge power of each electric vehicle at time t; The action space is expressed as: Where (P) is the energy bid amount, (R) is the regulation capacity bid amount, and (p(t)) is the charging and discharging power of the i-th vehicle.

5. The electric vehicle aggregator energy market optimization decision system based on reinforcement learning according to claim 1 is characterized in that: The reward function of the reinforcement learning optimization module includes: energy revenue, calculated based on the real-time electricity price and the energy bid amount; regulation capacity revenue, calculated based on the regulation capacity price and the regulation capacity bid amount; regulation mileage revenue, calculated based on the regulation mileage price and the flexibility of the electric vehicle in responding to the regulation signal; regulation deployment revenue, calculated based on the actual power of the electric vehicle in responding to the regulation signal; battery degradation cost, calculated based on the discharge power and unit degradation cost. The reward function is expressed as: Among them, energy income is: The revenue from regulating capacity is: Adjusted mileage earnings are: Adjusted deployment income is: The battery degradation cost is: Among them, π s is the scene probability, P s ,t en is the energy bid amount, λ t is the electricity price, P s ,t cap To adjust the capacity bidding amount, R s To adjust the capacity, is the performance score, P s ,t mil To adjust the mileage bid amount, To adjust the mileage coefficient, δ t To regulate the signal, p i ,s,k deg is the degradation cost coefficient, P i ,s,k dis is the discharge power.

6. The electric vehicle aggregator energy market optimization decision system based on reinforcement learning according to claim 1 is characterized in that: The multi-time scale modeling module further includes an uncertainty modeling unit for generating electricity price scenarios and regulation signal scenarios through stochastic programming and calculating the joint scenario probability, wherein the joint scenario probability is expressed as: s,t =π s π t,s Among them, π s is the probability of scene s, π t,s is the conditional probability of time t under scenario s; Among them, the scene set is represented as: s∈S={low,medium,high},π s =0.3,0.5,0.

2.

7. The electric vehicle aggregator energy market optimization decision system based on reinforcement learning according to claim 1 is characterized in that: When issuing charge and discharge instructions, the control execution module meets the following constraints: State of charge constraints: Charge and discharge power limit: almost electric rate to the end. Power balance constraints: Adjust capacity maintenance time constraint: Among them, the initial and final state of charge constraints are: Among them, SOC i , t is the power state, and are the charging and discharging efficiencies, p i ,t ch and p i ,t dis are charging and discharging power respectively, Δt is the time interval, u i , t is the state variable, p i ,t dis(ch) is the upper power limit, δ s To adjust the signal.

8. The electric vehicle aggregator energy market optimization decision system based on reinforcement learning according to claim 1 is characterized in that: The market clearing feedback module aims to maximize social welfare and optimize the winning electricity volume and node marginal electricity price, where the objective function is expressed as: The market clearing constraint is: The optimization results are fed back to the reinforcement learning optimization module through an iterative coupling mechanism, wherein the iterative coupling mechanism includes: the upper-level reinforcement learning outputs the bidding strategy; the lower-level market clearing calculates the winning electricity volume and node marginal electricity price; and the winning electricity volume and node marginal electricity price are used as environmental feedback to update the reinforcement learning strategy.

9. The electric vehicle aggregator electric energy market optimization decision system based on reinforcement learning according to claim 1 is characterized in that: The training process of the reinforcement learning optimization module includes: initializing the Actor network and the Critic network; observing the current market state and generating a bidding strategy; calculating the market clearing results and storing the experience samples; updating the network parameters to minimize the temporal difference error, where the temporal difference error is: TDerror=E[(Q target -Q(s,a|θ Q )) 2 ] and soft-update the target network, where the target network is updated as: θ Q’ ←τθ Q +(1-τ)θ Q’ θ μ’ ←τθ μ +(1-τ)θ μ’ . Determine whether the cumulative reward fluctuation is less than the set threshold to determine training convergence, where the cumulative reward is expressed as: Cumulative reward = ∑ t =1 T γ t-1 r t Where γ is the discount factor, r t is the reward at time t, and τ is the soft update coefficient.

10. The electric vehicle aggregator electric energy market optimization decision system based on reinforcement learning according to claim 1 is characterized in that: The system also includes a performance analysis module for evaluating the convergence, computational efficiency, and robustness of the system, wherein: Convergence is analyzed by cumulative reward curve, and cumulative reward is expressed as: cumulative reward = ∑ t =1 T γ t-1 r t The computational efficiency is evaluated by comparing the solution time of the reinforcement learning algorithm and the traditional stochastic programming. The robustness is analyzed by testing the revenue fluctuation amplitude in the perturbed electricity price scenario.

11. An optimization decision-making method for an electric vehicle aggregator electric energy market optimization decision-making system based on reinforcement learning according to any one of claims 1 to 10, characterized in that: The following steps are involved: Collect information on the remaining battery capacity, charge and discharge power limits, and real-time market electricity prices and capacity adjustment prices of electric vehicles; A two-layer optimization model is constructed, in which the upper layer model dynamically optimizes the bidding strategies in the day-ahead and real-time markets through reinforcement learning; the lower layer model allocates the winning electricity volume through a market clearing algorithm; and Lagrange multipliers are used to couple the time-scale interaction of the two models. Through a deep deterministic policy gradient algorithm, the energy bid amount, capacity bid amount, and charging and discharging power allocation strategy are dynamically adjusted based on real-time power consumption and node marginal electricity price feedback; The market clearing results are solved based on mixed integer linear programming, and the winning electricity volume and node marginal electricity price are fed back into the reinforcement learning strategy update; According to the optimized strategy, charging and discharging instructions are issued to electric vehicles to meet the state of charge constraints and power balance constraints.

12. The electric vehicle aggregator power market optimization decision-making method based on reinforcement learning according to claim 11 is characterized in that: The multi-timescale modeling includes: Day-ahead timescale modeling: forecasting electricity prices and regulation signal scenarios on an hourly basis, and determining the energy and regulation capacity bids of EV aggregators; Real-time timescale modeling: responding to regulation signals in real time on a minute-by-minute basis to adjust the charging and discharging power distribution of electric vehicles; Time scale coupling: The dual variables are transferred via Lagrange multipliers to couple the bidding decision of the day-ahead market with the power allocation in the real-time market to optimize the overall decision-making process.

13. The electric vehicle aggregator energy market optimization decision-making method based on reinforcement learning according to claim 11 is characterized in that: In the reinforcement learning optimization, the state space includes: current time, market information and electric vehicle status; the action space includes: bidding strategy and power allocation strategy; and the reward function includes: energy income, capacity adjustment income, mileage adjustment income, deployment adjustment income and battery degradation cost.

14. The electric vehicle aggregator power market optimization decision-making method based on reinforcement learning according to claim 11 is characterized in that: Uncertainty modeling is also included: generating electricity price scenarios and regulation signal scenarios through stochastic programming, and calculating the joint scenario probability.

15. The electric vehicle aggregator power market optimization decision-making method based on reinforcement learning according to claim 11 is characterized in that: The control execution satisfies the following constraints: state of charge constraint; charge and discharge power constraint; power balance constraint; regulation capacity maintenance time constraint; initial and final state of charge constraints.

16. The electric vehicle aggregator energy market optimization decision-making method based on reinforcement learning according to claim 11 is characterized in that: The market clearing feedback aims to maximize social welfare, optimizes the winning electricity volume and node marginal electricity price, and feeds the optimization results back to the reinforcement learning optimization module through an iterative coupling mechanism.

17. The electric vehicle aggregator power market optimization decision-making method based on reinforcement learning according to claim 11 is characterized in that: The training process of the reinforcement learning optimization module includes: initializing the Actor network and the Critic network; observing the current market state and generating a bidding strategy; calculating the market clearing results and storing experience samples; updating the network parameters to minimize the temporal difference error and soft-updating the target network; and judging whether the cumulative reward fluctuation is less than a set threshold to determine training convergence.

18. The electric vehicle aggregator power market optimization decision-making method based on reinforcement learning according to claim 11 is characterized in that: Performance analysis is also included: evaluating the convergence, computational efficiency, and robustness of the system.

19. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the electric vehicle aggregator electric energy market optimization decision-making method based on reinforcement learning as described in any one of claims 11 to 18 is implemented.

Citation Information

Cited By

  • Day-ahead bidding method and device for electric vehicle aggregator to participate in electric power spot market

    CN121458408A