A two-tier decision-making guidance method for electric vehicles in a coupled transportation electrification system
By improving the two-layer FMDP model trained by the Rainbow algorithm, the real-time interaction problem between electric vehicle charging and driving decision-making is solved, user costs are reduced, grid operation is optimized, and coordinated optimization of electric vehicles and the grid is achieved.
Patent Information
- Application Number
- CN202410264436.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-08
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-03-08
AI Technical Summary
Existing technologies make it difficult to achieve real-time interaction between electric vehicle charging and driving decisions, and are unable to effectively reduce user charging costs and ensure stable operation of the power grid. In addition, DRL-based methods are limited by the limited action output space and cannot simultaneously handle charging station recommendations and route navigation.
The improved Rainbow algorithm based on the DQN architecture is used to train the two-layer FMDP model. By decoupling the multi-objective optimization model, the action and reward functions of the upper and lower layer agents are designed. The Double DQN mechanism, Dueling DQN mechanism, priority replay cache mechanism and learning rate decay strategy are integrated to improve the learning ability and decision-making performance of the agent.
It realizes the real-time guidance of electric vehicle charging and driving, reduces the comprehensive cost of users, optimizes the voltage distribution of power grid nodes, and improves the generalization ability and decision-making efficiency of intelligent agents.
Smart Images

Figure CN118395829B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an electric vehicle double-layer decision-making guidance method for a traffic electrification coupling system, and belongs to the technical field of electric vehicle and power grid interaction. Background Art
[0002] Electric vehicles have garnered significant attention in recent years as environmentally friendly transportation options. Governments worldwide view the electrification of transportation as a path to addressing energy and environmental challenges. By combining historical travel data with behavioral economics, this approach considers the subjective and stochastic nature of electric vehicle owners and focuses on providing a boundedly rational, time-ahead charging decision-making method. While this approach can more realistically quantify the owner's decision-making process, helping to understand the nature of vehicle-grid interactions and enabling orderly vehicle guidance, it can also mitigate the impact of clustered charging. However, this behavioral economics-based guidance approach remains an offline operation, requiring centralized control of a large number of electric vehicles and consuming significant computing power. With the increasing number of electric vehicle users in cities, it is necessary to develop a real-time interactive charging and driving decision-making solution that can provide owners with optimal charging stations and driving routes. Furthermore, through multi-network information fusion and multi-objective optimization across the "vehicle-station-road-grid" network, it can reduce the interaction costs associated with charging and driving, as well as grid security.
[0003] In the multi-agent interaction and multi-objective optimization problem of electric vehicle charging guidance, electric vehicle owners, acting as intelligent agents, perceive information about the electrified transportation environment, including road network traffic conditions, charging prices, and charging costs. By perceiving the charging and driving conditions in the coupled network, the intelligent agent receives corresponding reward feedback to make decisions and evaluations, sequentially selecting appropriate charging stations for energy replenishment and the optimal driving route for travel until the action is completed and the destination is reached. Therefore, the decision-making process for electric vehicle charging guidance fully conforms to the relevant definition of a finite Markov chain. To effectively address the dimensionality curse caused by the intelligent agent's perception of the complex state of the electrified transportation coupling network, achieve real-time guidance for electric vehicle charging and driving, and address the problem that existing DRL-based methods are limited by the limited action output space and cannot simultaneously handle charging station recommendations and route navigation, it is necessary to develop a real-time and effective electric vehicle charging and driving decision guidance method.
[0004] The information disclosed in this background section is only intended to enhance understanding of the overall background of the invention and should not be considered as an admission or any form of suggestion that the information constitutes the prior art already known to a person of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide an electric vehicle two-tier decision guidance method for a transportation electrification coupled system. The improved Rainbow algorithm based on the DQN architecture improves the learning ability, decision performance, and generalization ability of the two-tier FMDP model. The trained two-tier FMDP model is used to implement two-tier decision guidance for electric vehicles, thereby reducing user charging costs while ensuring the stable operation of the power grid.
[0006] To achieve the above object, the present invention is implemented by adopting the following technical solutions:
[0007] The present invention discloses a two-layer decision-making guidance method for electric vehicles in a traffic electrification coupling system, comprising the following steps:
[0008] Obtain environmental data for electric vehicles;
[0009] Inputting the environmental data into a trained two-layer FMDP model for solving electric vehicle charging station recommendation and route navigation to obtain a decision-making guidance result for the electric vehicle;
[0010] The two-layer FMDP model is obtained by decoupling a pre-built multi-objective optimization model of the electric vehicle and transportation electrification coupling system. The objective function of the optimization model is to reduce the comprehensive cost of electric vehicle users and reduce the deviation of the grid voltage. The constraints of the optimization model include electric vehicle constraints, distribution network flow constraints, and operation safety constraints.
[0011] The two-layer FMDP model is trained and solved by an improved Rainbow algorithm based on the DQN architecture to obtain a trained two-layer FMDP model; the improved Rainbow algorithm based on the DQN architecture includes a Double DQN mechanism, a Dueling DQN mechanism, a priority replay cache mechanism, a learning rate decay strategy, and a dropout layer technology.
[0012] Furthermore, the objective function of the multi-objective optimization model of the electric vehicle and transportation electrification coupling system is expressed as follows:
[0013]
[0014]
[0015]
[0016]
[0017]
[0018]
[0019]
[0020] Where f represents the objective function of the optimization model; f1 represents the charging and travel costs of electric vehicle owners; f2 represents the penalty cost of power grid operation safety; δ mn A 0-1 variable indicating path selection; represents the charging station selection variable, Indicates that electric vehicle user i is recommended to the kth charging station;
[0021] represents the energy consumption cost of the electric vehicle user i; represents the charging cost of electric vehicle user i; π represents the unit time cost; T i tr represents the travel time of electric vehicle user i; T i wt represents the charging waiting time of electric vehicle user i; T i ch represents the charging time of electric vehicle user i; i∈Ω EV ,Ω EV is the set of electric vehicle numbers;
[0022] U t,p represents the real-time voltage of node p; Represents the rated voltage of node p; p∈N PG , N PG Indicates the power grid G PG The number of nodes; T represents the control time;
[0023] represents the average charging price at the charging station; ε r Represents the power consumption model per unit mileage of different levels of roads, r = 1, 2, 3; v mn Represents the traffic network G TN The road; i represents the path selection set of electric vehicle user i; v mn ∈Ω TN ,Ω TN For traffic network G TN A collection of road segments; mn Represents the traffic network G TN Road v mn length; A 0-1 variable indicating path selection;
[0024] represents the start time of charging of electric vehicle user i; Indicates the end charging time of electric vehicle user i; represents the real-time electricity price of the kth charging station, k∈Ω CS ,Ω CS is the collection of charging stations; P ch represents the charging power; Δt represents the simulation step length;
[0025] Represents the traffic network G TN Road v mn average traffic speed;
[0026] represents the battery capacity of electric vehicle user i; represents the SOC value of electric vehicle user i at the time of arrival; represents the expected final SOC value of electric vehicle user i; f ch Indicates the SOC value of the charging capacity; P ch Indicates the charging power of the charging pile; η ch Indicates the charging efficiency of the charging pile.
[0027] Furthermore, the electric vehicle constraint is expressed as follows:
[0028]
[0029]
[0030]
[0031] Where, Indicates the SOC value of electric vehicle user i when charging; ε r Represents the power consumption model per unit mileage of different levels of roads, r = 1, 2, 3; v mn Represents the traffic network G TN The road; i represents the path selection set of electric vehicle user i; l mn Represents the traffic network G TN Road v mn length; A 0-1 variable indicating path selection; represents the battery capacity of electric vehicle user i; e fl Indicates the minimum SOC value of the vehicle battery. If it is lower than this value, the electric vehicle is considered to have broken down.
[0032] represents the charging station selection variable, Indicates that electric vehicle user i is recommended to the kth charging station; Ω CS A collection of charging stations;
[0033] δ mn A 0-1 variable representing the path selection, Indicates the traffic node v m Adjacent road segment v mn gather;
[0034] The distribution network power flow constraint is expressed as follows:
[0035]
[0036]
[0037] Where, represents the charging active load of node p; represents the normal active load of node p; U t,p Represents the real-time voltage of the grid node p; G pq represents branch conductance; B pq represents branch susceptance; θ t,pq represents the phase angle difference; Represents the charging reactive load of the node; Indicates conventional reactive load;
[0038] The expression of the operational safety constraint is as follows:
[0039]
[0040]
[0041] Where U t,p Represents the real-time voltage of the grid node p; Indicates the lower limit of node voltage; Indicates the upper limit of node voltage; I t,pq Indicates the real-time current of line pq; Indicates the lower limit of current; Indicates the upper current limit.
[0042] Furthermore, the two-layer FMDP model includes an upper-layer agent and a lower-layer agent, wherein the upper-layer agent is used to couple the mapping between the network environment and the optimal charging station, and the lower-layer agent is used to couple the mapping between the network environment and the optimal driving path;
[0043] The training steps of the two-layer FMDP model are as follows:
[0044] Initialize the network parameters of the two-layer FMDP model;
[0045] The initialized two-layer FMDP model is iteratively trained using the improved Rainbow algorithm based on the DQN architecture until the preset iteration termination condition is reached, resulting in a trained two-layer FMDP model. Each training round includes the following steps:
[0046] Initializing a training environment, wherein the training environment includes electric vehicle data, charging station data, distribution network data, and traffic network data;
[0047] According to the training environment, for each electric vehicle user, the Double DQN mechanism, DuelingDQN mechanism, and priority replay cache mechanism are used to update the network parameters of the lower-layer agent and the upper-layer agent respectively;
[0048] After completing the network parameter update for the last electric vehicle user, the learning rate decay strategy is used to update the learning rates of the upper and lower intelligent agents respectively;
[0049] In each training round, the dropout layer technology is used to adaptively select and discard network neurons for the upper and lower layer agents based on the preset variable probabilities.
[0050] Furthermore, the lower-layer agent includes a lower-layer evaluation network, a lower-layer target network, and a lower-layer experience playback unit, and updating the network parameters of the lower-layer agent includes the following steps:
[0051] By observing the upper-level state, the upper-level action decision is obtained;
[0052] According to the upper-layer action decision, obtaining the lower-layer action decision by acquiring the lower-layer state;
[0053] Based on the lower-level action decision, the lower-level reward is calculated by observing the new lower-level state;
[0054] Obtaining corresponding lower-layer experience samples according to the lower-layer state, the lower-layer action decision, the new lower-layer state, and the lower-layer reward; and storing the lower-layer experience samples in a lower-layer experience replay unit based on a priority replay cache mechanism;
[0055] Calculating the loss values of multiple small samples in the lower-layer experience replay unit based on the Double DQN mechanism and the Dueling DQN mechanism; optimizing the network parameters of the lower-layer evaluation network using the gradient descent method with the goal of minimizing the loss value; assigning the network parameters of the lower-layer evaluation network to the lower-layer target network each time a preset optimization step threshold is passed;
[0056] After the network parameters are optimized, according to the new lower-level state, in response to the electric vehicle user not arriving at the target charging station, the step of obtaining the lower-level state is returned to for loop iteration, otherwise the step of updating the network parameters of the upper-level intelligent agent is entered.
[0057] Furthermore, the upper-layer agent includes an upper-layer evaluation network, an upper-layer target network, and an upper-layer experience playback unit. The network parameter update of the upper-layer agent includes the following steps:
[0058] By observing the new upper state, calculate the upper reward;
[0059] Obtaining corresponding upper-layer experience samples according to the upper-layer state, upper-layer action decision, new upper-layer state, and upper-layer reward; and storing the upper-layer experience samples in an upper-layer experience replay unit based on a priority replay cache mechanism;
[0060] The loss values of multiple small samples in the upper-layer experience replay unit are calculated based on the Double DQN mechanism and the Dueling DQN mechanism; the network parameters of the upper-layer evaluation network are optimized using the gradient descent method with the goal of minimizing the loss value; wherein, each time a preset optimization step threshold is passed, the network parameters of the upper-layer evaluation network are assigned to the upper-layer target network.
[0061] Furthermore, the expression of the upper state is as follows:
[0062]
[0063] Where, represents the upper-level state of electric vehicle user i; EV represents electric vehicle data; CS represents charging station data; PG represents distribution network data; t represents real-time time; represents the real-time SOC value of electric vehicle user i; represents the real-time location of electric vehicle user i; represents the real-time electricity price of the k-th charging station; represents the state variable of the kth charging station at time t, Indicates the number of free piles in the station, otherwise it indicates the number of people waiting in line; Indicates the location of the charging station; represents the real-time load of node p; represents the real-time voltage of node p;
[0064] The expression of the lower state is as follows:
[0065]
[0066] Where, represents the lower-level state of electric vehicle user i; TN represents the traffic network data; represents the location of the target charging station k for electric vehicle user i; Represents the traffic network G TN Road v mn Average traffic speed; mn Represents the traffic network G TN Road v mn length;
[0067] The expression of the upper-level action decision is as follows:
[0068]
[0069] Where, Represents the upper-level action decision; represents the location of the target charging station k for electric vehicle user i; Ω CS A collection of charging stations;
[0070] The expression of the lower-level action decision is as follows:
[0071]
[0072] Where, Indicates the lower-level action decision; v mn Represents the traffic network G TN the road; Indicates the traffic node v m Adjacent road segment v mn gather;
[0073] The expression of the upper layer reward is as follows:
[0074]
[0075] Where r i upp Indicates upper-level rewards; represents the charging cost of electric vehicle user i; π represents the unit time cost; T i wt represents the charging waiting time of electric vehicle user i; T i ch represents the charging time of electric vehicle user i; Represents the voltage penalty factor; N PG Indicates the power grid G PG Number of nodes; U t,p represents the real-time voltage of node p; represents the rated voltage of node p;
[0076] The expression of the lower-level reward is as follows:
[0077]
[0078] Where, Indicates the lower-level reward; l mn Represents the traffic network G TN Road v mn length; A 0-1 variable indicating path selection; represents the average charging price at the charging station; εr Represents the power consumption model per unit mileage of different levels of roads, r = 1, 2, 3; π represents the unit time cost; Represents the traffic network G TN Road v mn average traffic speed; represents the location of electric vehicle user i at the next moment; represents the location of the target charging station; ω arr Indicates navigation success reward; ω tow represents the navigation failure penalty, i.e., the towing cost in the area; represents the real-time SOC value of electric vehicle user i; represents the battery capacity of electric vehicle user i; e fl Indicates the lowest SOC value of the vehicle battery. If the SOC value is lower than this value, the electric vehicle is considered to have broken down.
[0079] Furthermore, the calculation of the loss value includes the following steps:
[0080] Based on the Double DQN mechanism, the action decision that can obtain the maximum Q value is obtained through the upper evaluation network or the lower evaluation network, and the Q value corresponding to the action decision is calculated through the upper target network or the lower target network; based on the Q value, the loss value is calculated. The expression of the loss value is as follows:
[0081]
[0082] Where L(ω) represents the loss value; r t represents the reward value; γ represents the discount factor;
[0083] Represents the estimated Q value of the target network;
[0084] Represents the estimated Q value of the evaluation network;
[0085] s t+1 Indicates the state at time t+1; a t+1 represents the action decision at time t+1; ω - Represents the neural network parameters of the target network; represents the neural network parameters of the evaluation network;
[0086] The upper evaluation network or the lower evaluation network or the upper target network or the lower target network is based on the Dueling DQN mechanism, which divides the network structure into state value and action advantage to calculate the Q value. The expression of the Dueling DQN mechanism is as follows:
[0087]
[0088] Where, Q(s t ,a t ) means in state s t With action a t The Q value under Indicates state s t The state value of Represents action decision a t The action value of Represents action decision a t+1 The action value of |A| represents the number of actions in the action space A.
[0089] Furthermore, the priority playback cache mechanism includes a priority storage mechanism and a priority sample extraction mechanism, which specifically includes the following steps:
[0090] Based on the priority storage mechanism, in response to the reward value corresponding to the upper-layer experience sample or the lower-layer experience sample being within the preset reward range, the upper-layer experience sample or the lower-layer experience sample is stored in the upper-layer experience playback unit or the lower-layer experience playback unit; otherwise, the upper-layer experience sample or the lower-layer experience sample is discarded;
[0091] Based on the priority sample extraction mechanism, the temporal difference deviation is used as the evaluation index to determine the sampling probability of the upper experience sample or the lower experience sample in the upper experience playback unit or the lower experience playback unit being sampled as a small sample. The expression of the sampling probability is as follows:
[0092]
[0093] Where, P j represents the sampling probability of the jth upper layer experience sample or lower layer experience sample; γ j Indicates the priority of the jth upper layer experience sample or lower layer experience sample; μ represents the control priority influencing factor, and its value range is [0,1]. When μ = 0, it indicates the original uniform sampling mechanism, and when μ = 1, it indicates the time difference deviation sampling mechanism; N bat Indicates the number of upper-layer experience samples or lower-layer experience samples in the upper-layer experience playback unit or the lower-layer experience playback unit.
[0094] Furthermore, the expression of the learning rate decay strategy is as follows:
[0095]
[0096] Where, α n represents the learning rate of the nth training round; α 0 represents the initial learning rate; τ represents the decay coefficient; n represents the current training round; n d Indicates a decay round.
[0097] Compared with the prior art, the present invention has the following beneficial effects:
[0098] In this paper, a two-layer FMDP model is first derived by decoupling the multi-objective optimization model of the coupled electric vehicle and transportation electrification system. Second, the two-layer FMDP model is trained and solved using the Rainbow algorithm, which integrates the Double DQN mechanism, the Dueling DQN mechanism, a prioritized replay cache mechanism, a learning rate decay strategy, and a dropout layer technique. Finally, the trained two-layer FMDP model enables two-tier decision guidance for electric vehicles. Compared with offline guidance strategies and online guidance strategies based on basic DQN, this method significantly reduces overall costs while optimizing node voltage distribution.
[0099] The present invention first decouples charging station recommendation and path navigation tasks into upper and lower-level decision-making problems, and specifically designs upper and lower-level action and reward functions to improve the efficiency of mutual collaboration between intelligent agents. Secondly, the improved Rainbow algorithm based on the DQN architecture effectively improves the generalization ability of the intelligent agent. Compared with the traditional DQN method, it can track and learn scenes outside the training data in real time. At the same time, the use of the priority replay cache mechanism and the learning rate decay strategy improves the learning ability and decision-making performance of the dual intelligent agents. Through the collaboration of the upper-lower-level dual intelligent agents, not only the average charging cost of EV users is reduced, but also the node voltage distribution is optimized. Compared with the offline guidance strategy and the online guidance strategy based on the basic DQN, the overall cost is effectively reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] Figure 1 It is a schematic diagram of the structure of the two-tier decision-making guidance method for electric vehicles in the transportation electrification coupling system;
[0101] Figure 2 It is a training flow chart of the two-tier decision guidance method for electric vehicles in the transportation electrification coupled system;
[0102] Figure 3 is a schematic diagram of the structure of an upper-layer agent or a lower-layer agent provided by an embodiment;
[0103] Figure 4 is a reward distribution diagram for training the upper-layer agent provided in the embodiment;
[0104] Figure 5 is a sliding average reward distribution diagram of the training of the upper-layer agent provided in the embodiment;
[0105] Figure 6 is a reward distribution diagram for training the lower-level agent provided in the embodiment;
[0106] Figure 7is a sliding average reward distribution diagram for the training of the lower-level agent provided in the embodiment;
[0107] Figure 8 1 is a reward distribution diagram for training an upper-layer agent comparing the four methods provided in the embodiment;
[0108] Figure 9 1 is a reward distribution diagram for training the lower-level agent comparing the four methods provided in the embodiment;
[0109] Figure 10 This is a comparison chart of the cumulative average total costs of the offline strategy and the online strategy of the five methods provided in the embodiment for 100 days;
[0110] Figure 11 This is a schematic diagram of the average charging service balance of BDRL provided in the embodiment;
[0111] Figure 12 This is a schematic diagram of the average charging service balance of BDRL provided in the embodiment;
[0112] Figure 13 This is a schematic diagram of the average charging service balance of Dueling-DQN provided in the embodiment;
[0113] Figure 14 This is a schematic diagram of the average charging service balance of DQN provided by the embodiment;
[0114] Figure 15 Schematic diagram of the average charging service balance of DDG provided in the embodiment. DETAILED DESCRIPTION
[0115] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0116] This embodiment provides a two-tier decision-making guidance method for electric vehicles in a traffic electrification coupling system, including the following steps:
[0117] Obtain environmental data for electric vehicles;
[0118] The environmental data is input into the trained two-layer FMDP model for solving electric vehicle charging station recommendation and route navigation, and the decision guidance results of electric vehicles are obtained;
[0119] Among them, the two-layer FMDP model is obtained by decoupling the pre-built multi-objective optimization model of the electric vehicle and transportation electrification coupling system. The objective function of the optimization model is to reduce the comprehensive cost of electric vehicle users and reduce the deviation of grid voltage. The constraints of the optimization model include electric vehicle constraints, distribution network flow constraints and operation safety constraints.
[0120] The two-layer FMDP model is trained and solved through the improved Rainbow algorithm based on the DQN architecture to obtain a trained two-layer FMDP model; the improved Rainbow algorithm based on the DQN architecture includes the Double DQN mechanism, Dueling DQN mechanism, priority replay cache mechanism, learning rate decay strategy and dropout layer technology.
[0121] The technical concept of this invention is as follows: first, a two-layer FMDP model is derived through optimization model decoupling. Second, an improved Rainbow algorithm based on the DQN architecture is used to train and solve the two-layer FMDP model. This improves the learning ability, decision-making performance, and generalization capability of the two-layer FMDP model. Compared with offline guidance strategies and online guidance strategies based on basic DQN, this method effectively reduces overall costs while optimizing node voltage distribution.
[0122] The specific steps of the electric vehicle two-tier decision-making guidance method for the transportation electrification coupling system provided in this embodiment are as follows:
[0123] S1: Construct a multi-objective optimization model for the coupled system of electric vehicles and transportation electrification.
[0124] With the goal of reducing the comprehensive cost of electric vehicle users and minimizing the deviation of grid voltage, a coordinated optimization target of the electric vehicle and transportation electrification coupling system is established.
[0125] S1.1: Problem modeling.
[0126] The objective function of the multi-objective optimization model of the electric vehicle and transportation electrification coupling system is expressed as follows:
[0127]
[0128]
[0129]
[0130]
[0131]
[0132]
[0133]
[0134] Where f represents the objective function of the optimization model; f1 represents the charging and travel costs of electric vehicle owners; f2 represents the penalty cost of power grid operation safety; δ mn A 0-1 variable indicating path selection; represents the charging station selection variable, Indicates that electric vehicle user i is recommended to the kth charging station;
[0135] represents the energy consumption cost of the electric vehicle user i; represents the charging cost of electric vehicle user i; π represents the unit time cost; T i tr represents the travel time of electric vehicle user i; T i wt represents the charging waiting time of electric vehicle user i; T i ch represents the charging time of electric vehicle user i; i∈Ω EV ,Ω EV is the set of electric vehicle numbers;
[0136] U t,p represents the real-time voltage of node p; Represents the rated voltage of node p; p∈N PG , N PG Indicates the power grid G PG The number of nodes; T represents the control time;
[0137] represents the average charging price at the charging station; ε r Represents the power consumption model per unit mileage of different levels of roads, r = 1, 2, 3; v mn Represents the traffic network G TN The road; i represents the path selection set of electric vehicle user i; v mn ∈Ω TN ,Ω TN For traffic network G TN A collection of road segments; mn Represents the traffic network G TN Road v mn length; A 0-1 variable indicating path selection;
[0138] represents the start time of charging of electric vehicle user i; Indicates the end charging time of electric vehicle user i; represents the real-time electricity price of the kth charging station, k∈Ω CS ,Ω CS is the collection of charging stations; P ch represents the charging power; Δt represents the simulation step length;
[0139] Represents the traffic network GTN Road v mn average traffic speed;
[0140] represents the battery capacity of electric vehicle user i; represents the SOC value of electric vehicle user i at the time of arrival; represents the expected final SOC value of electric vehicle user i; f ch Indicates the SOC value of the charging capacity; P ch Indicates the charging power of the charging pile; η ch Indicates the charging efficiency of the charging pile.
[0141] S1.2: Coordination constraints for the coupled system of electric vehicles and transportation electrification.
[0142] S1.2.1: The electric vehicle constraint is expressed as follows:
[0143]
[0144]
[0145]
[0146] Where, Indicates the SOC value of electric vehicle user i when charging; ε r Represents the power consumption model per unit mileage of different levels of roads, r = 1, 2, 3; v mn Represents the traffic network G TN The road; i represents the path selection set of electric vehicle user i; l mn Represents the traffic network G TN Road v mn length; A 0-1 variable indicating path selection; represents the battery capacity of electric vehicle user i; e fl Indicates the minimum SOC value of the vehicle battery. If it is lower than this value, the electric vehicle is considered to have broken down.
[0147] represents the charging station selection variable, Indicates that electric vehicle user i is recommended to the kth charging station; Ω CS A collection of charging stations;
[0148] δ mn A 0-1 variable representing the path selection, Indicates the traffic node v m Adjacent road segment v mn S1.2.2: The distribution network flow constraint is expressed as follows:
[0149]
[0150]
[0151] Where, represents the charging active load of node p; represents the normal active load of node p; U t,p Represents the real-time voltage of the grid node p; G pq represents branch conductance; B pq represents branch susceptance; θ t,pq represents the phase angle difference; Represents the charging reactive load of the node; Indicates normal reactive load.
[0152] S1.2.3: The operational safety constraint is expressed as follows:
[0153]
[0154]
[0155] Where U t,p Represents the real-time voltage of the grid node p; Indicates the lower limit of node voltage; Indicates the upper limit of node voltage; I t,pq Indicates the real-time current of line pq; Indicates the lower limit of current; Indicates the upper current limit.
[0156] S2: Decouple the optimization model in S1 to obtain a two-layer FMDP model for solving electric vehicle charging station recommendation and route navigation.
[0157] Specifically, such as Figure 1-3 As shown in the figure, the input data of the optimization model is decoupled into the state of the two-layer FMDP model; the output data of the optimization model is decoupled into the action decision of the two-layer FMDP model; and the optimization objectives and constraints of the optimization model are decoupled into the rewards of the two-layer FMDP model.
[0158] The two-layer FMDP model includes an upper-layer agent and a lower-layer agent. The upper-layer agent is used to couple the mapping between the network environment and the optimal charging station, while the lower-layer agent is used to couple the mapping between the network environment and the optimal driving path.
[0159] S2.1: Status.
[0160] The state represents the agent's real-time perception of the environment, while the state space represents the set of all possible states.
[0161] The expression of the upper state is as follows:
[0162]
[0163] Where, represents the upper-level state of electric vehicle user i; EV represents electric vehicle data; CS represents charging station data; PG represents distribution network data; t represents real-time time; represents the real-time SOC value of electric vehicle user i; represents the real-time location of electric vehicle user i; represents the real-time electricity price of the k-th charging station; represents the state variable of the kth charging station at time t, Indicates the number of free piles in the station, otherwise it indicates the number of people waiting in line; Indicates the location of the charging station; represents the real-time load of node p; represents the real-time voltage of node p;
[0164] The expression of the lower state is as follows:
[0165]
[0166] Where, represents the lower-level state of electric vehicle user i; TN represents the traffic network data; represents the location of the target charging station k for electric vehicle user i; Represents the traffic network G TN Road v mn Average traffic speed; mn Represents the traffic network G TN Road v mn length;
[0167] S2.2: Action decision.
[0168] Actions are decisions made by an agent given an environment.
[0169] The expression of the upper-level action decision is as follows:
[0170]
[0171] Where, represents the upper-level action decision of electric vehicle user i; represents the location of the target charging station k for electric vehicle user i; Ω CS A collection of charging stations;
[0172] The expression of the lower-level action decision is as follows:
[0173]
[0174] Where, represents the lower-level action decision of electric vehicle user i; v mn Represents the traffic network G TN the road; Indicates the traffic node v m Adjacent road segment v mn gather;
[0175] Formula (18) represents the traffic network G TN The road section is v mn is the action space of the lower-level agent. And the series of decision actions obtained Connect sequentially to build the optimal navigation path As shown in formula (19). Electric vehicles follow this path From current location Drive until you reach the target charging station
[0176]
[0177] S2.3: Rewards.
[0178] Rewards represent the immediate feedback an agent receives after selecting an action under a specific state. They are the most important part of training an agent to learn a certain ability or achieve a certain goal.
[0179] Upper level - Charging station recommendation: The action selected at the upper level charging station will directly affect the charging cost after the user arrives at the station Waiting time T i wt And charging time T i ch At the same time, the charging options of large-scale EVs will affect the operating status of the distribution network. Therefore, these three items are used as the reward function of the upper-level intelligent agent. The expression of the upper-level reward is as follows:
[0180]
[0181] Where r i upp Indicates upper-level rewards; represents the charging cost of electric vehicle user i; π represents the unit time cost; T i wt represents the charging waiting time of electric vehicle user i; T i ch represents the charging time of electric vehicle user i; Represents the voltage penalty factor; N PG Indicates the power grid G PGNumber of nodes; U t,p represents the real-time voltage of node p; represents the rated voltage of node p;
[0182] Lower layer - driving route navigation:
[0183] Combining the overall optimization objective and the upper-level reward design, the rewards for the lower-level agent primarily consist of the battery charge cost and road time cost during the user's navigation process. Once the vehicle reaches the target charging station, the agent receives a positive reward. Conversely, if the EV fails to reach its destination before its battery runs out, the user must call a tow truck. Consequently, the agent receives a negative penalty. The lower-level reward is expressed as follows:
[0184]
[0185] Where, Indicates the lower-level reward; l mn Represents the traffic network G TN Road v mn length; A 0-1 variable indicating path selection; represents the average charging price at the charging station; ε r Represents the power consumption model per unit mileage of different levels of roads, r = 1, 2, 3; π represents the unit time cost; Represents the traffic network G TN Road v mn average traffic speed; represents the location of electric vehicle user i at the next moment; represents the location of the target charging station; ω arr Indicates navigation success reward; ω tow represents the navigation failure penalty, i.e., the towing cost in the area; represents the real-time SOC value of electric vehicle user i; represents the battery capacity of electric vehicle user i; e fl Indicates the lowest SOC value of the vehicle battery. If the SOC value is lower than this value, the electric vehicle is considered to have broken down.
[0186] S2.4: Action-Value Function
[0187] The state-action value function is used to evaluate the cumulative expected reward that the agent can obtain after performing actions based on the current strategy.
[0188] Although the upper and lower agents rely on the strategy ψ upp With ψ low , their state-action value function Q ψ (s,a)(Q-value) are the same.
[0189]
[0190] Where: h∈H represents the time step; γ represents the discount factor.
[0191] In the EV charging navigation problem, the agent’s goal is to find the optimal strategy ψ*, which is equivalent to finding the strategy that can obtain the maximum Q ψ Strategy for (s,a):
[0192]
[0193] S3: The two-layer FMDP model is trained and solved using the improved Rainbow algorithm based on the DQN architecture to obtain a trained two-layer FMDP model.
[0194] like Figure 3 As shown, the agent mainly includes an evaluation network, a target network and an experience replay unit. The structures of the upper-layer agent and the lower-layer agent are consistent. Specifically, the lower-layer agent in this embodiment includes a lower-layer evaluation network, a lower-layer target network and a lower-layer experience replay unit, and the upper-layer agent includes an upper-layer evaluation network, an upper-layer target network and an upper-layer experience replay unit.
[0195] like Figure 2 As shown in Figure 2, the training steps of the two-layer FMDP model are as follows:
[0196] S3.1 Initialize the network parameters of the two-layer FMDP model; including the upper layer evaluation network parameters Upper target network parameter ω upp,- , lower layer evaluation network parameters Lower layer target network parameter ω low,- , discount factor γ, initial learning rate α 0 , attenuation coefficient τ, attenuation round n d .
[0197] S3.2 Iteratively train the initialized two-layer FMDP model using the improved Rainbow algorithm based on the DQN architecture until a preset iteration termination condition is reached, thereby obtaining a trained two-layer FMDP model. In this embodiment, the iteration termination condition is set to reach a preset maximum number of training rounds.
[0198] Each training round consists of the following steps:
[0199] S3.2.1 Initialize the training environment, which includes electric vehicle data, charging station data, distribution network data, and traffic network data.
[0200] S3.2.2 For the first electric vehicle user, by observing its upper state Get upper-level action decision That is, select a target charging station.
[0201] S3.2.3 Decision-making based on upper-level actions By getting the underlying status Get the lower-level action decision That is, the electric vehicle user formulates a route selection plan based on the target charging station in S3.2.2 and the real-time status of the electric vehicle user.
[0202] S3.2.4 Decision-making based on lower-level actions By observing the new underlying state Calculating lower-level rewards
[0203] S3.2.5 According to the lower layer status Lower-level action decision New lower level state and lower-level rewards Get the corresponding lower-level experience sample The lower-level experience samples are stored in the lower-level experience replay unit based on the priority replay cache mechanism.
[0204] The prioritized replay buffer mechanism works as follows: During training, the prioritized replay buffer specifies the probability of sampling each sample in the replay buffer based on the time difference error (TD-error). This allows samples with larger TD-errors to be sampled with a higher probability for network parameter optimization, speeding up training.
[0205]
[0206]
[0207] Where, δ j represents the timing error of sample j in the playback buffer; r j represents the reward of the sample; γ represents the discount factor; Indicates the estimated Q value of the target network; s j+1 Indicates the state of the sample at the next moment; Represents the action a obtained based on the target network t+1 ;a t+1 Indicates the action at the next moment; represents the neural network parameters of the evaluation network; ω - Represents the neural network parameters of the target network; Represents the estimated Q value of the evaluation network;
[0208] p jIt represents the probability that sample j is sampled into the small sample; rank(j) represents the loss size ranking of sample j.
[0209] The priority playback cache mechanism in this embodiment includes a priority storage mechanism and a priority sample extraction mechanism, which specifically includes the following steps:
[0210] Based on the priority storage mechanism, in response to the reward value corresponding to the upper-layer experience sample or the lower-layer experience sample being within the preset reward range, the upper-layer experience sample or the lower-layer experience sample is stored in the upper-layer experience playback unit or the lower-layer experience playback unit, otherwise it is discarded.
[0211] Specifically, in the later stages of the algorithm's utilization, random "noise" samples (i.e., samples with excessively low or high reward values) are removed from the upper or lower experience replay units. Once these samples are extracted for gradient updates, they can cause the agent to fall into a local optimum, reducing the stability of the algorithm in the later stages.
[0212] Through previous pre-training (i.e., using a basic sample storage mechanism to train the algorithm), determine the round e at which the agent roughly enters the exploitation phase, and calculate the minimum experience sample reward for each round The mean And the maximum experience sample reward per round The mean
[0213]
[0214]
[0215] Where: Represents the mean of the minimum experience sample reward per round; Indicates the minimum experience sample reward per round; N epi represents the total number of training rounds; e represents the round in which the agent enters the exploitation phase; Represents the mean of the maximum experience sample reward per round; Indicates the maximum experience sample reward per round.
[0216] If the reward value r corresponding to the extracted experience sample j j Within this range If the sample is not good, it will be considered as an excellent experience sample and stored in the experience replay pool, that is, the upper experience replay unit or the lower experience replay unit. Otherwise, it will be considered as a noise experience sample and discarded.
[0217] Based on the priority sample extraction mechanism, the temporal difference deviation is used as the evaluation index to determine the sampling probability of the upper experience sample or the lower experience sample in the upper experience playback unit or the lower experience playback unit being sampled as a small sample.
[0218] Specifically, for the DQN algorithm, experience replay and other probabilistic random sampling methods are used to evaluate network training. This fails to distinguish the quality of different samples, resulting in low training efficiency. Therefore, to assess the value and priority of different experience samples, temporal difference deviation δ is used as an evaluation metric to determine the probability of a sample being sampled.
[0219] Define the priority γ of experience sample j j It is expressed as follows:
[0220]
[0221] Where, γ j Indicates the priority of the jth experience sample; rank(j) indicates the loss size ranking of sample j.
[0222] The expression of sampling probability is as follows:
[0223]
[0224] Where, P j represents the sampling probability of the jth upper layer experience sample or lower layer experience sample; γ j Indicates the priority of the jth upper layer experience sample or lower layer experience sample; μ represents the control priority influencing factor, and its value range is [0,1]. When μ = 0, it indicates the original uniform sampling mechanism, and when μ = 1, it indicates the time difference deviation sampling mechanism; N bat Indicates the number of upper-layer experience samples or lower-layer experience samples in the upper-layer experience playback unit or the lower-layer experience playback unit.
[0225] In order to prevent the network from overfitting during training, the importance sampling weight w is used j To correct the parameters of the neural network:
[0226]
[0227] Where: w j P represents the importance sampling weight of the jth upper layer experience sample or lower layer experience sample; j P represents the sampling probability of the jth upper layer experience sample or lower layer experience sample; min represents the minimum sampling probability; β cor In this paper, β cor Set to 0.6.
[0228] S3.2.6 calculates the loss values of multiple small samples in the lower-level experience replay unit based on the Double DQN mechanism and the Dueling DQN mechanism; optimizes the network parameters of the lower-level evaluation network using the gradient descent method with the goal of minimizing the loss value; wherein, each time a preset optimization step threshold is passed, the network parameters of the lower-level evaluation network are assigned to the lower-level target network.
[0229] The Double DQN mechanism works as follows: In basic DQN, the maximum Q-value is used for iterative updates, which often leads to overestimation of Q-values. To address this, the Double Q Network (DQN) modifies the Q-value iteration rule based on DQN. The evaluation network determines the action that achieves the maximum Q-value, and the target network then calculates the Q-value corresponding to that action, effectively mitigating the overestimation of Q-values.
[0230]
[0231] Where: Represents the estimated Q value of the evaluation network;
[0232] α represents the learning rate, α∈[0,1], indicating the extent to which the Q-value is updated; α=0 means only using prior knowledge, α=1 means only considering the current estimate and ignoring previous information;
[0233] r t represents the reward; γ represents the discount factor;
[0234] Represents the estimated Q value of the target network;
[0235] Indicates the action that maximizes the Q value based on the evaluation network;
[0236] s t+1 Indicates the state at time t+1;
[0237] a t+1 represents the action decision at time t+1;
[0238] represents the neural network parameters of the evaluation network;
[0239] ω - Represents the neural network parameters of the target network;
[0240] Represents the estimated Q value of the evaluation network.
[0241] The calculation of loss value includes the following steps:
[0242] Based on the Double DQN mechanism, the action decision that can obtain the maximum Q value is obtained through the upper evaluation network or the lower evaluation network, and the Q value corresponding to the action decision is calculated through the upper target network or the lower target network; based on the Q value, the loss value is calculated. The expression of the loss value is as follows:
[0243]
[0244] Where L(ω) represents the loss value; r t represents the reward; γ represents the discount factor;
[0245] Represents the estimated Q value of the target network;
[0246] Represents the estimated Q value of the evaluation network;
[0247] s t+1 Indicates the state at time t+1; a t+1 represents the action decision at time t+1; ω - Represents the neural network parameters of the target network; Represents the neural network parameters of the evaluation network.
[0248] When Q(s t ,a t ) converges to Q * (s t ,a t ), the optimal strategy can be expressed as follows using the greedy strategy:
[0249]
[0250] Among them, the upper evaluation network or the lower evaluation network or the upper target network or the lower target network is based on the Dueling DQN mechanism to divide the network structure into state value and action advantage to calculate the Q value.
[0251] The principle of Dueling-DQN mechanism is as follows: by dividing the network structure into state value And action advantages The Q-value is output in two aspects. indicates the quality of the state obtained, A higher value indicates that the state is more conducive to the agent's learning. Indicates the quality of the obtained action, The higher the value, the higher the reward for the action compared to other actions. The improvement of the network structure removes redundant degrees of freedom and improves the efficiency of the algorithm. The mathematical expression of Dueling-DQN is:
[0252]
[0253] Where, Q(s t ,a t ) means in state s t With action a t The Q value under Indicates state s t The state value of Represents action decision a t The action value of Represents action decision a t+1 The action value of |A| represents the number of actions in the action space A.
[0254] After the network parameters are optimized in S3.2.7, based on the new lower-level state, in response to the electric vehicle user not arriving at the target charging station, the process returns to the step of obtaining the lower-level state in S3.2.3 for loop iteration, otherwise, the process proceeds to the step of updating the network parameters of the upper-level intelligent agent in 3.2.8.
[0255] 3.2.8 Calculate the upper layer reward by observing the new upper layer state.
[0256] 3.2.9 Based on the upper-level state, upper-level action decision, new upper-level state, and upper-level reward, obtain the corresponding upper-level experience sample; based on the priority replay cache mechanism, store the upper-level experience sample in the upper-level experience replay unit; this step is the same as step 3.2.5 and will not be repeated here.
[0257] 3.2.10 Calculate the loss of multiple small samples in the upper-layer experience replay unit based on the Double DQN and Dueling DQN mechanisms. Optimize the network parameters of the upper-layer evaluation network using gradient descent to minimize the loss. Each time the optimization step exceeds a preset threshold, assign the network parameters of the upper-layer evaluation network to the upper-layer target network. This step is similar to step 3.2.6 and will not be repeated here.
[0258] 3.2.11 At this point, the network parameter update of the lower-layer intelligent agent and the upper-layer intelligent agent for the first electric vehicle user is completed, and the process returns to step 3.2.2 to update the network parameters of the lower-layer intelligent agent and the upper-layer intelligent agent for the second electric vehicle user. The process iterates until the network parameter update of the lower-layer intelligent agent and the upper-layer intelligent agent for the last electric vehicle user is completed. The learning rate decay strategy is then used to update the learning rates of the upper-layer intelligent agent and the lower-layer intelligent agent respectively, and then a new round of training is carried out.
[0259] The learning rate in the basic DQN is fixed throughout the training process and cannot be dynamically adjusted as training rounds increase, which reduces the algorithm's ability to explore and exploit. To this end, we set the learning rate to decrease based on a reciprocal decay model to balance the agent's early exploration ability and later exploitation ability. The expression of the learning rate decay strategy is as follows:
[0260]
[0261] Where, α n represents the learning rate of the nth training round; α 0 represents the initial learning rate; τ represents the decay coefficient; n represents the current training round; n d Indicates a decay round.
[0262] In this embodiment, in each training round, the dropout layer technology is used to adaptively select and discard network neurons for the upper-layer agents and the lower-layer agents based on the preset variable probabilities.
[0263] While deep neural networks with numerous parameters are powerful machine learning systems, they suffer from overfitting and low computational performance during training and testing. To overcome these drawbacks, a dropout layer was introduced. The dropout layer is an improvement mechanism, not an adjustment to the loss function. Overall, the algorithm's architecture still involves adjusting the parameters of each neural network during training. During testing, only the input and output mapping is performed, without adjusting the neural network parameters. The output is state information (vehicle-station-road-network information) and the system's decision-making results, namely charging station locations at the upper layer and traffic route information at the lower layer. Therefore, the dropout layer technology, by adaptively selecting and discarding neurons, can alleviate overfitting in the evaluation network and effectively improve the generalization ability of the trained model.
[0264] The main purpose of introducing a dropout layer is to improve the generalization performance of the original neural network. During each training iteration, neurons are temporarily deleted with a probability p. During testing, these deleted neurons are restored with a probability 1-p. Furthermore, the training process is repeated over multiple iterations. During training, the neuron parameters are adjusted, forming a process of discarding and restoring. The neuron parameters remain unchanged; they are temporarily deactivated during training and then recalled during testing.
[0265] For a neural network with L layers, l∈{1,2,..,L} represents the hidden layer retrieval. w, b, and x represent the weight, bias, and input in layer l, respectively.
[0266] During the training phase, the feedforward neural network with dropout is operated, and the output layer It is expressed as follows:
[0267]
[0268] Where: represents the output layer in the training phase; w l represents the weight of the hidden layer L; x l represents the input of the hidden layer L; b l Represents the bias value of the hidden layer L; represents the dot product operator; μ l represents a vector of independent Bernoulli random variables, each with probability p equal to 0; B represents a Bernoulli distribution.
[0269] In the test phase, the feedforward neural network with dropout is operated, and the output layer It is expressed as follows:
[0270]
[0271] Standard dropout is equivalent to adding another layer after a layer of neurons, setting the values to zero with a certain probability during training, and then multiplying by 1-p during testing.
[0272] The calculation example configuration in this embodiment is as follows:
[0273] In order to match the scale of the urban road network, the IEEE-33 node system is used as the urban distribution network. The charging stations are connected to the distribution network nodes at the following locations: 2, 4, 6, 7, 9, 11, 13, 16, 20, 23, 24, 26, 28 and 32. At the upper level, the voltage penalty cost Set to 100 yuan. In the lower layer, EV navigation successfully rewards w arr With failure penalty w tow The values are set to 100 yuan and 200 yuan, respectively. Each round dispatches 1,000 EVs that require charging. Table 1 also lists the parameters for the improved Rainbow algorithm. The training environment is configured with a CPU i99960X, a GPU RTX2070, and 32GB of RAM.
[0274] Table 1 Improved Rainbow algorithm parameter configuration
[0275] parameter Upper level - recommended charging stations Lower layer - driving route navigation Number of hidden layers {100,80} {120,100} Discount rate γ 0.95 0.95 <![CDATA[Initialize the learning rate α 0 > 0.55 0.15 Attenuation coefficient τ 0.85 0.85 <![CDATA[Decay round n d > 70 150 Dropout probability p 0.20 0.20 <![CDATA[Small sample size N b > 68 128 <![CDATA[Priority experience replay buffer size Ξ > 6000 8000
[0276] Figure 4-7 The reward distribution of each training is shown. The total training time for 1000 rounds is 4.25 hours. Figure 4-7As can be seen, in the early stages of training, the agents were encouraged to explore the environment at a high learning rate, and the reward values fluctuated significantly. The upper-layer agent continuously accumulated historical experience and learned to recommend the optimal CS, achieving stability after 200 rounds, with an average convergence reward of -90.64 yuan. However, due to the complex road network navigation environment and unstable outputs from the upper layer, the lower-layer agent only achieved basic stability after approximately 450 rounds. Specifically, in the first 200 rounds, the lower-layer agent struggled to perform the navigation task, resulting in battery depletion in some guided vehicles. As the number of training rounds increased, the lower-layer reward eventually stabilized at around 13.24 yuan. Through coordinated cooperation, the two-layer agents achieved charging and driving decision guidance for electric vehicle users.
[0277] Figure 8 is the reward value graph obtained by the upper-level intelligent agent, Figure 9 This is a graph of the reward values obtained by the upper-layer agent. In this example, the offline strategy (DDG) and the traditional DRL strategies (DQN, Double-DQN, and Dueling-DQN) are selected to comprehensively compare the implementation effects of BDRL. In particular, only the constructed two-layer FMDP problem is solved using the DQN, Double-DQN, and Dueling-DQN methods, and 1000 training rounds are also required. The time taken to complete the entire training for DQN, Double-DQN, and Dueling-DQN is 3.29 hours, 3.79 hours, and 3.55 hours, respectively.
[0278] Depend on Figure 8 and Figure 9 Overall, while all four algorithms achieved stable convergence from exploring the external environment, significant differences existed in convergence speed, stability, and solution quality. Because the proposed method incorporates advanced extended and improved mechanisms from other methods, the learning rate can be dynamically adjusted based on the training rounds. It also eliminates the interference of noise samples on the agent's learning process, allowing for a faster convergence. Compared to other DQN methods, the BDRL method achieved the highest rewards for both the upper and lower layers, at -90.64 and 13.24 yuan, respectively. The basic DQN algorithm achieved relatively fast convergence speeds for both the upper and lower layers, reaching approximately 150 and 350 rounds, respectively. However, due to its simple network structure and training mechanism, this method achieved relatively poor solution quality, with final rewards stabilizing at -93.14 and 8.16 yuan, respectively. Furthermore, the Double-DQN and Dueling-DQN algorithms, while improving their neural network architecture and Q-value calculation paradigms, could be further improved in balancing the agent's initial exploration speed with later solution quality.
[0279] Figure 10is the cumulative average total cost of the offline strategy and the online strategy for 100 days (i.e. the cumulative sum of the average daily time and expenses of all car owners. Figure 10 It can be seen that the method proposed in the present invention can effectively reduce the total cost.
[0280] Figure 11-15 Table 2 shows the average charging service balance and specific evaluation index values for the above strategies. It can be seen that DDG, as a static guidance strategy, cannot automatically adjust its decision output based on changes in real-time information. Therefore, its cumulative cost is the highest, reaching 10,518.36 yuan. Compared with the DDG strategy, DQN, Double-DQN, Dueling-DQN, and BDRL, as dynamic guidance strategies, have the ability to perceive the environment in real time and adaptively adjust decisions. Their average cumulative rewards are reduced by 7.06%, 11.60%, 9.47%, and 13.75%, respectively.
[0281] Table 2 Comparison of evaluation indicators of offline and online strategies
[0282]
[0283] Regarding specific evaluation metrics, charging costs accounted for the highest proportion among the five strategies, exceeding 40% for all strategies. This indicates that charging costs are the most significant expense for vehicle owners. Because DDG uses a static guidance model with nearby CS recommendations and shortest path planning, its costs for all metrics, except energy consumption, are higher than those of DRL-based methods. This indicates that energy consumption is proportional to travel distance. DRL-based online decision-making methods achieve real-time decision mapping, with decision times within seconds. The proposed BDRL method, based on an improved Rainbow architecture, enhances the capabilities of both offline learning of discrete actions and online decision-making, outperforming other DQN-based methods in all metrics. Furthermore, the average charging service balance scores for DQN, Double-DQN, Dueling-DQN, and BDRL are 1.50, 1.51, 1.54, and 0.96, respectively. These results further demonstrate the superiority of the proposed method in electric vehicle charging guidance decision-making.
[0284] The terms involved in this embodiment are explained as follows:
[0285] DRL: Deep Reinforcement Learning, deep reinforcement learning;
[0286] FMDP: Finite Markov Decision Process, finite Markov decision process;
[0287] DQN: Deep Q Network, deep Q network;
[0288] DDG: Data Dependence Graph, data dependency graph;
[0289] BDRL: Bi-level Deep Reinforcement Learning, double-level deep reinforcement learning.
[0290] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0291] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0292] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0293] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0294] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A two-tier decision-making guidance method for electric vehicles in a transportation electrification coupled system, characterized by: The steps include: Obtain environmental data from electric vehicles; Inputting the environmental data into a trained two-layer FMDP model for solving electric vehicle charging station recommendation and route navigation to obtain a decision-making guidance result for the electric vehicle; The two-layer FMDP model is obtained by decoupling a pre-built multi-objective optimization model of the electric vehicle and transportation electrification coupling system. The objective function of the optimization model is to reduce the comprehensive cost of electric vehicle users and reduce the deviation of the grid voltage. The constraints of the optimization model include electric vehicle constraints, distribution network flow constraints, and operation safety constraints. The two-layer FMDP model is trained and solved by an improved Rainbow algorithm based on the DQN architecture to obtain a trained two-layer FMDP model; the improved Rainbow algorithm based on the DQN architecture includes a Double DQN mechanism, a DuelingDQN mechanism, a prioritized replay cache mechanism, a learning rate decay strategy, and a dropout layer technology; The two-layer FMDP model includes an upper-layer agent and a lower-layer agent, wherein the upper-layer agent is used to couple the mapping between the network environment and the optimal charging station, and the lower-layer agent is used to couple the mapping between the network environment and the optimal driving path; The training steps of the two-layer FMDP model are as follows: Initialize the network parameters of the two-layer FMDP model; The initialized two-layer FMDP model is iteratively trained using the improved Rainbow algorithm based on the DQN architecture until the preset iteration termination condition is reached, resulting in a trained two-layer FMDP model. Each training round includes the following steps: Initializing a training environment, wherein the training environment includes electric vehicle data, charging station data, distribution network data, and traffic network data; According to the training environment, for each electric vehicle user, the Double DQN mechanism, Dueling DQN mechanism, and priority replay cache mechanism are used to update the network parameters of the lower-layer agent and the upper-layer agent respectively; After completing the network parameter update for the last electric vehicle user, the learning rate decay strategy is used to update the learning rates of the upper and lower intelligent agents respectively; In each training round, the dropout layer technology is used to adaptively select and discard network neurons for the upper and lower layer agents based on the preset variable probabilities.
2. The electric vehicle two-tier decision-making guidance method for a transportation electrification coupling system according to claim 1 is characterized in that: The objective function of the multi-objective optimization model of the electric vehicle and transportation electrification coupling system is expressed as follows: ; ; ; ; ; ; ; Where, represents the objective function of the optimization model; represents the charging and toll costs for electric vehicle owners; represents the penalty cost for power grid operation safety; A 0-1 variable indicating path selection; represents the charging station selection variable, Indicates that electric vehicle user i is recommended to the kth charging station; represents the energy consumption cost of the electric vehicle user i; represents the charging cost of electric vehicle user i; Represents the unit time cost; represents the travel time of electric vehicle user i; Indicates the charging waiting time of electric vehicle user i; represents the charging time of electric vehicle user i; , is the set of electric vehicle numbers; represents the real-time voltage of node p; represents the rated voltage of node p; , Indicates the power grid Number of nodes; Indicates control time; represents the average charging price at the charging station; Represents the power consumption model per unit mileage of different levels of roads, ; Represents the traffic network the road; represents the path selection set of electric vehicle user i; , For transportation network A collection of road segments; Represents the traffic network Road length; A 0-1 variable indicating path selection; represents the start time of charging of electric vehicle user i; Indicates the end charging time of electric vehicle user i; represents the real-time electricity price of the k-th charging station, , A collection of charging stations; Indicates charging power; represents the simulation step size; Represents the traffic network Road average traffic speed; represents the battery capacity of electric vehicle user i; represents the SOC value of electric vehicle user i at the time of arrival; represents the expected final SOC value of electric vehicle user i; Indicates the SOC value of the charging capacity; Indicates the charging power of the charging pile; Indicates the charging efficiency of the charging pile.
3. The electric vehicle two-tier decision-making guidance method for a transportation electrification coupling system according to claim 1 is characterized in that: The expression of the electric vehicle constraint is as follows: ; ; ; Where, Indicates the SOC value of electric vehicle user i when charging is required; Represents the power consumption model per unit mileage of different levels of roads, ; Represents the traffic network the road; represents the path selection set of electric vehicle user i; Represents the traffic network Road length; A 0-1 variable indicating path selection; represents the battery capacity of electric vehicle user i; Indicates the minimum SOC value of the vehicle battery. If it is lower than this value, the electric vehicle is considered to have broken down. represents the charging station selection variable, Indicates that electric vehicle user i is recommended to the kth charging station; A collection of charging stations; A 0-1 variable representing the path selection, Representation and transportation nodes Adjacent road sections gather; The distribution network power flow constraint is expressed as follows: ; ; Where, represents the charging active load of node p; represents the conventional active load of node p; Represents the real-time voltage of the grid node p; represents branch conductance; Indicates branch susceptance; represents the phase angle difference; Represents the charging reactive load of the node; Indicates conventional reactive load; The expression of the operational safety constraint is as follows: ; ; Where, Represents the real-time voltage of the grid node p; Indicates the lower limit of node voltage; Indicates the upper limit of node voltage; Indicates the real-time current of line pq; Indicates the lower limit of current; Indicates the upper current limit.
4. The electric vehicle two-tier decision-making guidance method for a transportation electrification coupling system according to claim 1 is characterized in that: The lower-layer agent includes a lower-layer evaluation network, a lower-layer target network, and a lower-layer experience playback unit. The network parameter update of the lower-layer agent includes the following steps: By observing the upper-level state, the upper-level action decision is obtained; According to the upper-layer action decision, obtaining the lower-layer action decision by acquiring the lower-layer state; Based on the lower-level action decision, the lower-level reward is calculated by observing the new lower-level state; Obtaining corresponding lower-layer experience samples according to the lower-layer state, lower-layer action decision, new lower-layer state, and lower-layer reward; and storing the lower-layer experience samples in a lower-layer experience replay unit based on a priority replay cache mechanism; Calculating the loss values of multiple small samples in the lower-layer experience replay unit based on the Double DQN mechanism and the Dueling DQN mechanism; With the goal of minimizing the loss value, the gradient descent method is used to optimize the network parameters of the lower evaluation network; wherein, each time the preset optimization step threshold is passed, the network parameters of the lower evaluation network are assigned to the lower target network; After the network parameters are optimized, according to the new lower-level state, in response to the electric vehicle user not arriving at the target charging station, the step of obtaining the lower-level state is returned to for loop iteration, otherwise the step of updating the network parameters of the upper-level intelligent agent is entered.
5. The electric vehicle double-layer decision-making guidance method for the transportation electrification coupling system according to claim 4 is characterized in that: The upper-layer agent includes an upper-layer evaluation network, an upper-layer target network, and an upper-layer experience playback unit. The network parameter update of the upper-layer agent includes the following steps: By observing the new upper state, calculate the upper reward; Obtaining a corresponding upper-layer experience sample according to the upper-layer state, the upper-layer action decision, the new upper-layer state, and the upper-layer reward; and storing the upper-layer experience sample in an upper-layer experience replay unit based on a priority replay cache mechanism; Calculating the loss values of multiple small samples in the upper-layer experience replay unit based on the Double DQN mechanism and the Dueling DQN mechanism; With the goal of minimizing the loss value, the gradient descent method is used to optimize the network parameters of the upper evaluation network; wherein, each time the preset optimization step threshold is passed, the network parameters of the upper evaluation network are assigned to the upper target network.
6. The electric vehicle two-tier decision-making guidance method for a transportation electrification coupling system according to claim 5 is characterized in that: The expression of the upper state is as follows: ; Where, represents the upper state of electric vehicle user i; Represents electric vehicle data; Represents charging station data; Represents distribution network data; Indicates real time; represents the real-time SOC value of electric vehicle user i; represents the real-time location of electric vehicle user i; represents the real-time electricity price of the k-th charging station; represents the state variable of the kth charging station at time t, Indicates the number of free piles in the station, otherwise it indicates the number of people waiting in line; Indicates the location of the charging station; represents the real-time load of node p; represents the real-time voltage of node p; The expression of the lower state is as follows: ; Where, represents the lower level state of electric vehicle user i; Represents traffic network data; represents the location of the target charging station k for electric vehicle user i; Represents the traffic network Road average traffic speed; Represents the traffic network Road length; The expression of the upper-level action decision is as follows: ; Where, Represents the upper-level action decision; represents the location of the target charging station k for electric vehicle user i; A collection of charging stations; The expression of the lower-level action decision is as follows: ; Where, Represents the lower-level action decision; Represents the traffic network the road; Representation and transportation nodes Adjacent road sections gather; The expression of the upper layer reward is as follows: ; Where, Indicates upper-level rewards; represents the charging cost of electric vehicle user i; Indicates the unit time cost; Indicates the charging waiting time of electric vehicle user i; represents the charging time of electric vehicle user i; represents the voltage penalty factor; Indicates the power grid Number of nodes; represents the real-time voltage of node p; represents the rated voltage of node p; The expression of the lower-level reward is as follows: ; Where, Indicates lower-level rewards; Represents the traffic network Road length; A 0-1 variable indicating path selection; represents the average charging price at the charging station; Represents the power consumption model per unit mileage of different levels of roads, ; Indicates the unit time cost; Represents the traffic network Road average traffic speed; represents the location of electric vehicle user i at the next moment; Indicates the location of the target charging station; Indicates navigation success reward; represents the navigation failure penalty, i.e., the towing cost in the area; represents the real-time SOC value of electric vehicle user i; represents the battery capacity of electric vehicle user i; Indicates the lowest SOC value of the vehicle battery. If the SOC value is lower than this value, the electric vehicle is considered to have broken down.
7. The electric vehicle two-tier decision-making guidance method for a transportation electrification coupling system according to claim 5 is characterized in that: The calculation of the loss value includes the following steps: Based on the Double DQN mechanism, the action decision that can obtain the maximum Q value is obtained through the upper evaluation network or the lower evaluation network, and the Q value corresponding to the action decision is calculated through the upper target network or the lower target network; based on the Q value, the loss value is calculated. The expression of the loss value is as follows: ; Where, Indicates the loss value; Represents the reward value; represents the discount factor; Represents the estimated Q value of the target network; Represents the estimated Q value of the evaluation network; express The state of the moment; express Moment-by-moment action decisions; Represents the neural network parameters of the target network; represents the neural network parameters of the evaluation network; The upper evaluation network or the lower evaluation network or the upper target network or the lower target network is based on the Dueling DQN mechanism, which divides the network structure into state value and action advantage to calculate the Q value. The expression of the Dueling DQN mechanism is as follows: ; Where, Indicates that the status and action The Q value under Indicates status The state value of Representing action decisions The action value of Representing action decisions The action value of Represents the number of actions in the action space A.
8. The electric vehicle two-tier decision-making guidance method for a transportation electrification coupling system according to claim 5 is characterized in that: The priority playback cache mechanism includes a priority storage mechanism and a priority sample extraction mechanism, and specifically includes the following steps: Based on the priority storage mechanism, in response to the reward value corresponding to the upper-layer experience sample or the lower-layer experience sample being within the preset reward range, the upper-layer experience sample or the lower-layer experience sample is stored in the upper-layer experience playback unit or the lower-layer experience playback unit; otherwise, the upper-layer experience sample or the lower-layer experience sample is discarded; Based on the priority sample extraction mechanism, the temporal difference deviation is used as the evaluation index to determine the sampling probability of the upper experience sample or the lower experience sample in the upper experience playback unit or the lower experience playback unit being sampled as a small sample. The expression of the sampling probability is as follows: ; Where, Indicates the The sampling probability of an upper experience sample or a lower experience sample; Indicates the the priority of an upper or lower level experience sample; Indicates the control priority impact factor, its value range is [0,1]. When the original is based on the uniform sampling mechanism, Time representation is based on the time difference deviation sampling mechanism; Indicates the number of upper-layer experience samples or lower-layer experience samples in the upper-layer experience playback unit or the lower-layer experience playback unit.
9. The electric vehicle two-tier decision-making guidance method for a transportation electrification coupling system according to claim 4 is characterized in that: The expression of the learning rate decay strategy is as follows: ; Where, Indicates the learning rate of the nth training round; represents the initial learning rate; represents the attenuation coefficient; Indicates the current training round; Indicates a decay round.
Citation Information
Patent Citations
Sustainable deep learning charging station / pile recommendation system and recommendation method
CN110888908A
Electric vehicle charging navigation system design based on voice recognition
CN110986986A