A Charging Station Pricing and Charging Method Based on a Robust Reward Function of Multi-Agent Reinforcement Learning

Through the reinforcement learning of multiple agents, the robust reward function of electric vehicles is solved, and the pricing and charging decisions of charging stations are optimized, which solves the impact of random charging of electric vehicles on the distribution network in the existing technology, and realizes the efficient returns of charging stations and the stability of the power system.

CN118917901BActive Publication Date: 2025-07-22HOHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410823734.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2025-07-22
Estimated Expiration
2044-06-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the uncertainty effect of random charging of electric vehicles on the distribution network, resulting in increased computational complexity and economic losses, and lack of coordinated optimization decisions between distribution networks and electric vehicles under information barriers.

Method used

Using a robust reward function based on multi-agent reinforcement learning, the uncertainty of electric vehicle behavior is processed through historical data, and combined with the constraints of distribution networks and charging stations, a robust reward function is constructed to optimize charging station pricing and charging decisions.

Benefits of technology

It improves the revenue and utilization rate of charging stations, balances the load of the distribution network, improves the stability of the power system and the accuracy of decision-making under information privacy, and adapts to market changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118917901B_ABST
    Figure CN118917901B_ABST
Patent Text Reader

Abstract

The present invention discloses a charging station pricing and charging method based on a robust reward function of multi-agent reinforcement learning. By using the robust reward function to handle the uncertainty of electric vehicle behavior, on the basis of historical data, the optimal pricing and charging decision of the charging station are realized. Multi-agent deep reinforcement learning is used to mine the non-cooperative game actions of the charging station, realizing the efficient coordination of the distribution network - electric vehicle. A Markov process with uncertain states is adopted to construct a robust reward function in the worst-case scenario; the policy difference is converted into a reward difference through the total variation distance. Considering the operation constraints and optimization objectives of the power grid and the charging station, an optimal pricing and charging model of the charging station based on the robust reward function of multi-agent deep reinforcement learning is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of coordinated optimization of the electric - transportation network, and particularly relates to a charging station pricing and charging method based on a robust reward function of multi - agent reinforcement learning. Background Art

[0002] Currently, significant investment is being made in transportation electrification to promote energy transition and decarbonization. This includes large - scale deployment of electric transportation networks and distribution networks, and a close integration between vehicles and fast - charging stations. In recent years, the scale of electric vehicles has grown rapidly. By 2030, the number of electric vehicles is expected to reach 1.81 billion, consuming 43000 GWh of electricity annually. The distribution network and local congestion may be caused, which may lead to large - scale power outages. The large amount of electrical energy consumed by the random charging of electric vehicles brings a huge burden to electric vehicle users, causing serious economic losses and negative social impacts.

[0003] However, there are two important gaps in the existing literature regarding charging station decision - making. First, most of the existing modeling work encapsulates charging demand in the transportation network. In fact, this highly depends on the complete and private information of the origin - destination trips of electric vehicles. The real - world environment is highly random and dynamic, and accurate demand modeling may not be very realistic. Second, in some existing literature, electric vehicles are used as agents for decision - making. With the increase in the number of electric vehicles, the computational complexity of pricing optimization will increase significantly. Therefore, there is an urgent need to develop a charging station pricing and charging method based on a robust reward function of multi - agent reinforcement learning to consider the uncertainty of charging demand and the coupling relationship between the distribution network and electric vehicles. Summary of the Invention

[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide a charging station pricing and charging method based on a robust reward function of multi - agent reinforcement learning in view of the deficiencies of the existing technology. The present invention takes into account the efficient coordination among the distribution network, electric vehicles, and charging stations. The robust reinforcement learning algorithm calculates the charging fees and pricing decisions of charging stations based on historical data to handle the uncertainty of electric vehicle behavior. Grid - related constraints are introduced to perform coordinated scheduling of the power - charging stations. The present invention can not only promote the coordination between the distribution network and electric vehicles under the information barrier, but also help charging stations increase their own profits based on the dynamic market.

[0005] Technical Solution: To solve the above - mentioned technical problem, the present invention proposes a charging station pricing and charging method based on a robust reward function of multi - agent reinforcement learning, and the method includes the following steps:

[0006] Step 1: Obtain the network coefficients and operation coefficients of the power grid model. The network coefficients include the power grid topology, line resistance, and impedance, and the operation coefficients include the power generation coefficient of the generator set, the charge and discharge coefficient of the energy storage system, the photovoltaic inverter coefficient, and the charging station parameters;

[0007] Step 2: Obtain the power grid load demand, photovoltaic output, historical traffic flow of the charging station, and historical charging demand scenario data;

[0008] Step 3: For the obtained distribution network parameters, taking the power grid operation constraints and unit operation constraints as the constraint conditions, and taking the minimum distribution network operation cost as the objective function, establish a distribution network operation model based on optimal power flow. According to this model, obtain the nodal price of the node to which the charging station belongs;

[0009] Step 4: The charging station serves as an agent, taking the nodal price obtained in Step 3 and the historical data of the charging station as the input of the agent state. Taking the vehicle charging demand update constraint, parking time update constraint, and charging power constraint as the constraint conditions, and taking the maximum comprehensive income of the charging station as the objective function, establish a charging station operation model. According to this model, calculate the declared power of the charging station and the distribution network;

[0010] Step 5: Based on the nodal price obtained in Step 3 and the declared power obtained in Step 4, use the multi-agent robust proximal policy optimization algorithm to calculate the robust reward function of the charging station. The process of constructing the robust reward function is the process of quantifying the difference between the reward in the worst case and the deterministic reward, and relax the difference using the dual form of the Hőlder inequality and the Cauchy–Schwarz inequality.

[0011] Further, in Step 3, the distribution network operation model is:

[0012]

[0013] In the formula, F DN represents the distribution network operation cost, g represents the distributed power source, a g represents the second-order cost coefficient of the distributed power source g, b g represents the first-order cost coefficient of the distributed power source g, c g represents the constant cost coefficient of the distributed power source g, P g,t represents the output power of the distributed power source g at time t, P sub,t represents the electric power purchased from the root node at time t, c sub,t represents the electricity price for purchasing electricity from the root node at time t, c DNe represents the charge and discharge cost coefficient of the energy storage numbered DNe, represents the discharge power of the energy storage numbered DNe at time t, represents the charging power of the energy storage numbered DNe at time t, ccut Indicates the penalty coefficient that meets the curtailment, and j represents the load node. Indicates the load curtailment amount of node j at time t.

[0014] Furthermore, in step 3, the distribution network operation constraints are as follows:

[0015] P g,t -P g,t-1 ≤RU g (A-2)

[0016] P g,t-1 -P g,t ≤RD g (A-3)

[0017]

[0018] SOC DNe,min ≤SOC DNe,t ≤SOC DNe,max (A-5)

[0019]

[0020] P g,min ≤P g,t ≤P g,max (A-12)

[0021] Q g,min ≤Q g,t ≤Q g,max (A-13)

[0022]

[0023] P sub,min ≤P sub,t ≤P sub,max (A-15)

[0024]

[0025] In the formula, RU g and RD g respectively represent the upper and lower bounds of the power ramp of distributed power source g, SOC DNe,t , SOC DNe,t+1 respectively represent the state of charge of the energy storage numbered DNe at time t and time t+1. Represents the charging state of the energy storage numbered DNe at time t. Represents the discharging state of the energy storage numbered DNe at time t. η ch and η dis respectively represent the charging and discharging efficiencies, Δt DNIndicates the time interval of the distribution network, SOC DNe,min , SOC DNe,max respectively represent the minimum and maximum values of the state of charge of the energy storage numbered DNe, P j,t represents the active power load of node j at time t, represents the reduction amount of the load of node j at time t, f and k both represent nodes in the distribution network, N DN represents the set of nodes in the distribution network, F DN represents the set of lines flowing into node j in the distribution network, T DN represents the set of lines flowing out of node j in the distribution network, represents the active power load amount of node j after reduction at time t, Q j,t represents the reactive power load amount of node j at time t, P jf,t and P kj,t respectively represent the active power of line jf and line kj at time t, Q jf,t and Q kj,t respectively represent the reactive power of line jf and line kj at time t, R kj , X kj respectively represent the resistance and reactance of line kj; I kj,t , I jf,t respectively represent the currents of line kj and jf at time t; B DN represents the set of all lines in the distribution network, V f,t , V j,t , V k,t respectively represent the voltages of nodes f, j, and k at time t, V k,min , V k,max respectively represent the minimum and maximum limits of the voltage of node k, I jf,max represents the maximum value of the current of line jf, P g,min , P g,max respectively represent the minimum and maximum values of the active power output of distributed power source g, Q g represents the reactive power output of distributed power source g,.Q g,min . and Q g,max respectively represent the maximum and minimum values of the reactive power output of distributed power source g, P sub,min and P sub,max respectively represent the minimum and maximum values of the power purchase of the root node, and respectively represent the maximum power of discharging and charging of the energy storage numbered DNe.

[0026] Furthermore, in step 4, the operation model of the charging station with the goal of maximizing the profit is:

[0027]

[0028] Charging demand update constraint:

[0029]

[0030] Remaining parking time update constraint:

[0031]

[0032] Charging power constraint:

[0033]

[0034] Set definition:

[0035]

[0036] In the formula, C z,t represents the revenue of charging station z at time t, r z,t represents the service fee pricing of charging station z at time t, represents the arrival time of electric vehicle i, represents the set of vehicles arriving at charging station z at historical time t, D i (r z,t ) represents the charging demand of electric vehicle i under the service fee rz,t, pi represents the initial parking duration of electric vehicle i, λ j,t represents the marginal electricity price of node j at time t, ω represents the penalty coefficient of power deviation, e z,t represents the total charging rate of charging station z at time t, represents the actual charging rate of charging station z at time t, respectively represent the remaining charging demands of electric vehicle i at time t and t + 1, respectively represent the remaining parking times of electric vehicle i at time t and t + 1, x i,t represents the charging rate of electric vehicle i at time t, represents the maximum charging rate of electric vehicle i, represents the maximum total charging rate of charging station z, and respectively represent the sets of electric vehicles in charging station z at times t and t + 1.

[0037] Furthermore, in step 5, the calculation of the robust reward function is as follows:

[0038]

[0039] In the formula, respectively represent the strategies of the charging station in the uncertain state and the certain state, A represents the action set of the charging station, S represents the state set of the charging station, represents the strategy and π CS The difference in the reward function of the charging station z below, s t / s t+1 respectively represent the states of the charging station at times t and t + 1 respectively represent the policy π CS and The probability that the charging station is in state s at time t under t represents the policy π CS and The state value function of the charging station in state s at time t under t a t represents the action of the charging station at time t, r t represents the revenue of the charging station at time t respectively represent the overall revenues of the charging station under the policies represents the policy and π CS The total variation distance of the charging station state s under t

[0040] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0041] Compared with the traditional charging station decision-making scheme, the technical solution of the present invention deals with the uncertainty of electric vehicle behavior by introducing a robust reward function. At the same time, to address the problem of information barriers, the agent makes decisions about its own behavior based on historical data. The results of the numerical example tests show that the method proposed in the present invention helps the charging station to improve its own revenue in a more complex real environment compared with the existing methods, improves the utilization rate of charging piles, balances the peak and valley values of the distribution network load, and improves the stability of the power system. Considering the realistic situation of the existing information barrier between the current transportation network and the charging station, the present invention makes decisions through robust reinforcement learning based on historical data to ensure the accurate strategy of the charging station under information privacy, and realizes the optimal pricing and charging decision of the charging station. Brief Description of the Drawings

[0042] Figure 1 is the flowchart of the method of the present invention;

[0043] Figure 2 is the utilization situation of the charging station under different service fees;

[0044] Figure 3 is the comparison chart of the training results of the proposed algorithm and the existing algorithms. Detailed Embodiments

[0045] ​​​The present invention will be further illustrated below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art to the present invention fall within the scope defined by the appended claims of this application.

[0046] As Figure 1 shown, the present invention proposes a charging station pricing and charging method based on a robust reward function of multi-agent reinforcement learning. The method includes the following steps:

[0047] Step 1: Obtain the network coefficients and operating coefficients of the power grid model. The network coefficients include the power grid topology, line resistance, and impedance, and the operating coefficients include the power generation coefficient of the generator set, the charge and discharge coefficient of the energy storage system, the coefficient of the photovoltaic inverter, and the parameters of the charging station.

[0048] Step 2: Obtain the power grid load demand, photovoltaic output, historical traffic flow of the charging station, and historical charging demand scenario data.

[0049] Step 3: For the obtained distribution network parameters, taking the power grid operation constraints and unit operation constraints as constraint conditions and the minimum distribution network operation cost as the objective function, establish a distribution network operation model based on optimal power flow, and obtain the nodal price of the node where the charging station is located according to this model.

[0050] Step 4: Regarding the charging station as an agent, taking the nodal price obtained in Step 3 and the historical data of the charging station as the input of the agent state, taking the vehicle charging demand update constraint, parking time update constraint, and charging power constraint as constraint conditions, and the maximum comprehensive income of the charging station as the objective function, establish a charging station operation model, and calculate the declared power of the charging station and the distribution network according to this model.

[0051] Step 5: Based on the nodal price obtained in Step 3 and the declared power obtained in Step 4, use the multi-agent robust proximal policy optimization algorithm to calculate the robust reward function of the charging station. The process of constructing the robust reward function is the process of quantifying the difference between the reward in the worst case and the deterministic reward, and the dual form of the Hőlder inequality and the Cauchy–Schwarz inequality is used to relax this difference.

[0052] Further, in Step 3, the distribution network operation model is:

[0053]

[0054] In the formula, F DN represents the distribution network operation cost, g represents the distributed power source, a g represents the second-order cost coefficient of the distributed power source g, b g represents the first-order cost coefficient of the distributed power source g, c gThe constant cost coefficient representing the distributed power source g, P g,t The output power of the distributed power source g at time t, P sub,t The electric power purchased from the root node at time t, c sub,t The electricity price for purchasing electricity from the root node at time t, c DNe The charge-discharge cost coefficient representing the energy storage numbered DNe The discharge power of the energy storage numbered DNe at time t The charging power of the energy storage numbered DNe at time t, c cut The penalty coefficient representing compliance reduction, where j represents the load node The load reduction amount of node j at time t

[0055] Furthermore, in step 3, the operation constraints of the distribution network are as follows:

[0056] P g,t -P g,t-1 ≤RU g (A-2)

[0057] P g,t-1 -P g,t ≤RD g (A-3)

[0058]

[0059] SOC DNe,min ≤SOC DNe,t ≤SOC DNe,max (A-5)

[0060]

[0061] P g,min ≤P g,t ≤P g,max (A-12)

[0062] Q g,min ≤Q g,t ≤Q g,max (A-13)

[0063]

[0064] P sub,min ≤P sub,t ≤P sub,max (A-15)

[0065]

[0066] In the formula, RU g and RD grespectively represent the upper and lower bounds of the power ramp of distributed power source g, SOC DNe,t , SOC DNe,t+1 respectively represent the state of charge of the energy storage numbered DNe at time t and time t + 1, represents the charging state of the energy storage numbered DNe at time t, represents the discharging state of the energy storage numbered DNe at time t, η ch and η dis respectively represent the charging and discharging efficiencies, Δt DN represents the time interval of the distribution network, SOC DNe,min , SOC DNe,max respectively represent the minimum and maximum values of the state of charge of the energy storage numbered DNe, P j,t represents the active power load of node j at time t, represents the amount of load reduction of node j at time t, f and k both represent nodes in the distribution network, N DN represents the set of distribution network nodes, F DN represents the set of lines flowing into node j in the distribution network, T DN represents the set of lines flowing out of node j in the distribution network, represents the amount of active power load after reduction of node j at time t, Q j,t represents the reactive power load of node j at time t, P jf,t and P kj,t respectively represent the active power of line jf and line kj at time t, Q jf,t and Q kj,t respectively represent the reactive power of line jf and line kj at time t, R kj , X kj respectively represent the resistance and reactance of line kj; I kj,t , I jf,t respectively represent the current of line kj and line jf at time t; B DN represents the set of all lines in the distribution network, V f,t , V j,t , V k,t respectively represent the voltages of nodes f, j, and k at time t, V k,min , V k,max respectively represent the minimum and maximum limits of the voltage of node k, I jf,max represents the maximum value of the current of line jf, P g,min , P g,max respectively represent the minimum and maximum values of the active power output of distributed power source g, Q g represents the reactive power output of distributed power source g,.Q g,min . and Q g,max respectively represent the maximum and minimum values of the reactive power output of distributed power source g, Psub,min and P sub,max represent the minimum and maximum values of the electricity purchase volume of the root node respectively, and represent the maximum power of discharging and charging of the energy storage with the number DNe respectively.

[0067] Furthermore, in step 4, the operation model of the charging station aiming at maximizing the profit is:

[0068]

[0069] Charging demand update constraint:

[0070]

[0071] Remaining parking time update constraint:

[0072]

[0073] Charging power constraint:

[0074]

[0075] Set definition:

[0076]

[0077] In the formula, C z,t represents the profit of the charging station z at time t, r z,t represents the service fee pricing of the charging station z at time t, represents the arrival time of the electric vehicle i, represents the set of vehicles arriving at the charging station z at the historical time t, D i (r z,t ) represents the charging demand of the electric vehicle i under the service fee rz,t, pi represents the initial parking duration of the electric vehicle i, λ j,t represents the marginal electricity price of the node j at time t, ω represents the penalty coefficient of the electricity quantity deviation, e z,t represents the total charging rate of the charging station z at time t, represents the actual charging rate of the charging station z at time t, represent the remaining charging demands of the electric vehicle i at time t and t + 1 respectively, represent the remaining parking times of the electric vehicle i at time t and t + 1 respectively, x i,t represents the charging rate of the electric vehicle i at time t, represents the maximum charging rate of the electric vehicle i, represents the maximum total charging rate of the charging station z, and respectively represent the set of electric vehicles in the charging station z at times t and t+1.

[0078] Furthermore, in step 5, the calculation of the robust reward function is as follows:

[0079]

[0080] In the formula, respectively represent the strategies of the charging station in the uncertain state and the certain state, A represents the set of actions of the charging station, S represents the set of states of the charging station, represents the strategy and π CS the difference in the reward function of charging station z under t / s t+1 respectively represent the states of the charging station at times t and t+1, respectively represent the strategies π CS and the probability that the charging station is in state s t at time t under represents the strategy π CS and the state value function of the charging station in state s t at time t under t a represents the action of the charging station at time t, r t represents the revenue of the charging station at time t, respectively represent the comprehensive revenues of the charging station under the strategies represents the strategy and π CS the total variation distance of the charging station state s t under

[0081] Case study

[0082] The superiority of the charging station pricing and charging method based on the robust reward function of multi-agent reinforcement learning described in the present invention is illustrated by the following case study. The present invention uses Figure 2 the improved IEEE 33-node power system shown. To compare the superiority of the method proposed in the present invention, the charging station pricing and charging methods based on multi-agent proximal policy optimization and the charging station pricing and charging method based on robust multi-agent proximal policy optimization proposed in the present invention are respectively used for power-electric vehicle coordinated scheduling. The present invention is implemented through the Pytorch platform and the Gurobi solver is used to solve the non-linear programming problem.

[0083] ​Based on this example, a comparison of the performance and scheduling results of the charging station pricing and charging method based on multi-agent proximal policy optimization and the charging station pricing and charging method based on robust multi-agent proximal policy optimization proposed by the present invention (the results are shown in Table 1), as well as the utilization of the charging station under different service fees (the results are shown in Figure 2 ). It shows that compared with the model that ignores the uncertainty of electric vehicles, the method proposed by the present invention can bring higher revenue returns to the charging station in complex real-world scenarios (Table 1 and Figure 3 ), and the charging station shows sensitivity to environmental changes when formulating strategies, considering the market supply and demand relationship and the price sensitivity of electric vehicles.

[0084] Table 1 Comparison of prediction performance and scheduling results of different methods

[0085] Performance index Multi-agent Proximal Policy Optimization Robust Multi-agent Proximal Policy Optimization Number of iterations 60 121 Scheduling result / ¥ 1737.38 1139.95

[0086] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. A charging station pricing and charging method based on a robust reward function of multi-agent reinforcement learning, characterized in that The method includes the following steps: Step 1: Obtain the network coefficients and operation coefficients of the power grid model. The network coefficients include the power grid topology, line resistance, and impedance, and the operation coefficients include the power generation coefficient of the generator set, the charge and discharge coefficient of the energy storage system, the photovoltaic inverter coefficient, and the charging station parameters; Step 2: Obtain the power grid load demand, photovoltaic output, historical traffic flow of the charging station, and historical charging demand scenario data; Step 3: For the obtained distribution network parameters, with the power grid operation constraints and unit operation constraints as the constraint conditions, and the minimum distribution network operation cost as the objective function, establish a distribution network operation model based on optimal power flow, and obtain the nodal price of the node to which the charging station belongs according to this model; Step 4: The charging station is used as an agent, taking the nodal price obtained in Step 3 and the historical data of the charging station as the input of the agent state, with the charging demand update constraint, parking time update constraint, and charging power constraint as the constraint conditions, and the maximum comprehensive income of the charging station as the objective function, establish a charging station operation model, and calculate the declared power of the charging station and the distribution network according to this model; Step 5: Based on the nodal price obtained in Step 3 and the declared power obtained in Step 4, use the multi-agent robust proximal policy optimization algorithm to calculate the robust reward function of the charging station. The process of constructing the robust reward function is the process of quantifying the difference between the reward in the worst-case scenario and the deterministic reward, and relax the difference using the dual form of the Hőlder inequality and the Cauchy–Schwarz inequality.

2. The charging station pricing and charging method based on a robust reward function of multi-agent reinforcement learning according to claim 1, characterized in that In Step 3, the distribution network operation model is: Where F DN represents the operation cost of the distribution network, g represents the distributed power source, and a g represents the second-order cost coefficient of the distributed power source g, and b g represents the first-order cost coefficient of the distributed power source g, and c g represents the constant cost coefficient of the distributed power source g, P g,t represents the output power of the distributed power source g at time t, P sub,t represents the electric power purchased from the root node at time t, c sub,t represents the electricity price for purchasing electricity from the root node at time t, c DNe represents the charge and discharge cost coefficient of the energy storage numbered DNe, represents the discharge power of the energy storage numbered DNe at time t, represents the charge power of the energy storage numbered DNe at time t, c cut represents the penalty coefficient for load curtailment, j represents the load node, represents the load curtailment amount of node j at time t.

3. The charging station pricing and charging method based on a robust reward function of multi-agent reinforcement learning according to claim 2, characterized in that, In Step 3, the distribution network operation constraints are: P g,t -P g,t-1 ≤RU g (A - 2) P g,t-1 -P g,t ≤RD g (A - 3) SOC DNe,min ≤SOC DNe,t ≤SOC DNe,max (A - 5) P g,min ≤P g,t ≤P g,max (A - 12) Q g,min ≤Q g,t ≤Q g,max (A - 13) P sub,min ≤P sub,t ≤P sub,max (A - 15) Wherein, RU g and RD g respectively represent the upper and lower bounds of the power ramp of the distributed power source g, SOC DNe,t , SOC DNe,t+1 respectively represent the state of charge of the energy storage numbered DNe at time t and time t + 1. represents the charging state of the energy storage numbered DNe at time t. represents the discharging state of the energy storage numbered DNe at time t. η ch and η dis respectively represent the charging and discharging efficiencies, Δt DN represents the time interval of the distribution network, SOC DNe,min , SOC DNe,max respectively represent the minimum and maximum values of the state of charge of the energy storage numbered DNe, P j,t represents the active power load of node j at time t. represents the reduction amount of the load of node j at time t. f and k both represent nodes in the distribution network, N DN represents the set of nodes in the distribution network, F’ DN represents the set of lines flowing into node j in the distribution network, T DN represents the set of lines flowing out of node j in the distribution network. represents the reduced active power load of node j at time t, Q j,t represents the reactive power load of node j at time t, P jf,t and P kj,t respectively represent the active power of line jf and line kj at time t, Q jf,t and Q kj,t respectively represent the reactive power of line jf and line kj at time t, R kj , X kj respectively represent the resistance and reactance of line kj; I kj,t , I jf,t respectively represent the current of line kj and line jf at time t; B DN represents the set of all lines in the distribution network, V f,t , V j,t , V k,t respectively represent the voltages of nodes f, j, and k at time t, V k,min , V k,max respectively represent the minimum and maximum limits of the voltage of node k, I jf,max represents the maximum value of the current of line jf, P g,min , P g,max respectively represent the minimum and maximum values of the active power output of the distributed power source g, Q g represents the reactive power output of the distributed power source g, Q g,min and Q g,max respectively represent the minimum and maximum reactive power outputs of distributed power source g, P sub,min and P sub,max respectively represent the minimum and maximum purchased power quantities of the root node, and respectively represent the maximum charging and discharging powers of the energy storage numbered DNe, λ j,t represents the marginal electricity price of node j at time t.

4. The charging station pricing and charging method based on a robust reward function of multi-agent reinforcement learning according to claim 3, characterized in that, In Step 4, the charging station operation model with the maximization of income as the objective is: Charging demand update constraint: Remaining parking time update constraint: Charging power constraint: Set definition: where, C z,t represents the revenue of charging station z at time t, r z,t represents the service fee pricing of charging station z at time t, represents the arrival time of electric vehicle i, represents the set of vehicles arriving at charging station z at historical time t, D i (r z,t ) represents the charging demand of electric vehicle i at service fee r z,t , pi represents the initial parking duration of electric vehicle i, λ j,t represents the marginal electricity price of node j at time t, ω represents the penalty coefficient of power deviation, e z,t represents the total charging rate of charging station z at time t, represents the actual charging rate of charging station z at time t, respectively represent the remaining charging demands of electric vehicle i at times t and t + 1, respectively represent the remaining parking times of electric vehicle i at times t and t + 1, x i,t represents the charging rate of electric vehicle i at time t, represents the maximum charging rate of electric vehicle i, represents the maximum total charging rate of charging station z, and respectively represent the sets of electric vehicles in charging station z at times t and t + 1.

5. A charging station pricing and charging method based on a robust reward function of multi-agent reinforcement learning according to claim 4, characterized in that, In Step 5, the calculation of the robust reward function is: where, π CS and represent the strategies of the charging station in the uncertain state and the certain state respectively. A represents the action set of the charging station, S represents the state set of the charging station, represents the strategy and π CS is the difference in the reward function of the charging station z under the strategies t and s t+1 represent the states of the charging station at time t and t + 1 respectively. and represent the probabilities that the charging station is in the state s CS and at time t under the strategies π t respectively. and represent the state value functions of the charging station in the state s CS and at time t under the strategies π t respectively. a t represents the action of the charging station at time t, r t represents the revenue of the charging station at time t, J(π CS ) and represent the comprehensive revenues of the charging station under the strategies π CS and respectively. represents the total variation distance of the state s and π CS of the charging station. t ​

Citation Information

Patent Citations

  • Time-phased joint pricing method and system for electric vehicle photovoltaic charging station network

    CN113052402A

  • Optical storage charging station operation optimization method and system based on near-end strategy optimization algorithm

    CN115986834A