Electric vehicle charging navigation method based on electricity price prediction and electronic equipment

By establishing a power-traffic coupling network model and deep reinforcement learning of multiple agents, and optimizing the charging navigation of electric vehicles, the overload problem of centralized charging of electric vehicles on the power grid is solved, and cost optimization and efficient travel are achieved.

CN120373700AActive Publication Date: 2025-07-25SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510301098.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-25
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

In the prior art, centralized charging of electric vehicles brings overload risks to the power distribution network, increases waiting time and additional costs, and lacks an effective multi-agent deep reinforcement learning method to optimize charging navigation.

Method used

Establish a power-traffic coupling network model, combine multi-agent deep reinforcement learning, optimize charging navigation through electricity price prediction information, design state variables, action variables and reward functions, and build a multi-agent SAC training network to realize intelligent decision-making to reduce costs.

Benefits of technology

Provide optimal paths and charging station selection, improve travel efficiency, reduce driving and charging costs, and enhance system autonomy and decision-making optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373700A_ABST
    Figure CN120373700A_ABST
Patent Text Reader

Abstract

The invention discloses an electric vehicle charging navigation method based on electricity price prediction and electronic equipment, and the method comprises the steps: building a power-traffic coupling network model according to a power distribution network, a traffic network, a charging station and an electric vehicle model; according to the electric power-traffic coupling network model, establishing a charging navigation economic optimization model of the electric vehicle, and determining variables and constraints of the model; designing a state variable set, an action variable set and a reward function in multi-agent deep reinforcement learning according to the variables and the constraints; constructing a multi-agent SAC training network structure considering electricity price prediction information, and setting parameters of an action network and an evaluation network of the SAC training network structure; through an interactive power-traffic coupling network model, a training agent learns to make an optimal decision under the actual traffic condition and the dynamic electricity price so as to maximize a reward function, and cost optimization of an electric vehicle user in the driving and charging process is realized. According to the invention, the economic driving and charging cost of the electric vehicle can be effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of coordinated optimization of power systems and transportation networks, and particularly to an electric vehicle charging navigation method and an electronic device based on electricity price prediction. Background Art

[0002] With the continuous popularization of electric vehicles, the number of charging facilities is increasing day by day to continuously meet the expanding market demand. This trend has greatly increased the proportion of power load in the transportation network and significantly enhanced the interactive coordination between the distribution network and the transportation network. However, large-scale centralized charging of electric vehicles will have adverse effects such as overloading on the power distribution network, threatening the safe operation of the power grid and increasing additional costs such as excessive waiting time.

[0003] The power-transportation coupled network is an inevitable trend in the development of electric vehicles. As an advanced extension of single-agent deep reinforcement learning, multi-agent deep reinforcement learning enables a single agent to maximize its reward through distributed decision-making, enhanced adaptability, and collaborative intelligence, thus effectively solving complex and large-scale problems. At present, the application of charging navigation technology based on multi-agent deep reinforcement learning in the power-transportation coupled network still needs to be developed, especially in optimizing potential future costs. Therefore, there is an urgent need for an efficient and accurate method to provide electric vehicle users with the selection of the optimal path and the optimal charging station. Summary of the Invention

[0004] To solve at least one of the technical problems existing in the prior art to a certain extent, an object of the present invention is to provide an electric vehicle charging navigation method, an electronic device, and a medium based on electricity price prediction.

[0005] The first technical solution adopted by the present invention is as follows:

[0006] An electric vehicle charging navigation method based on electricity price prediction, comprising the following steps:

[0007] Establish a power-transportation coupled network model according to the distribution network, the transportation network, the charging station, and the electric vehicle model;

[0008] Establish an economic optimization model for electric vehicle charging navigation according to the power-transportation coupled network model, and determine the variables and constraints of the model; design a state variable set X=(X1, X2,..., X m ), an action variable set A=(a1, a2,..., a m ), and a reward function R=(R1,..., R m ) in multi-agent deep reinforcement learning, where m represents the number of agents;

[0009] Construct a multi-agent SAC training network structure considering electricity price prediction information, and set the parameters of the action network and the evaluation network of the SAC training network structure;

[0010] Through the interactive power-transportation coupling network model, train the agent to learn to make optimal decisions under actual traffic conditions and dynamic electricity prices to maximize the reward function, thereby realizing the cost optimization of electric vehicle users during driving and charging and achieving an economic operation level.

[0011] Furthermore, based on the distribution network, transportation network, charging stations, and electric vehicle models, establish a power-transportation coupling network model, including:

[0012] The optimal power flow model of the distribution network is shown in the following formula (1):

[0013]

[0014] Among them, the constraints for configuring the optimal power flow model of the distribution network include power balance constraints and voltage constraints;

[0015] The power balance constraints are shown in formulas (2)-(10):

[0016]

[0017] In the formula, c g represents the electricity purchase cost, represents the active power output by the node generator, represents the power injected from other networks, a 0,i , a 1,i , ρ represent their correlation coefficients; V i is the AC voltage of node i, V i , respectively represent the upper and lower bounds of V i , respectively represent the upper and lower bounds of the generator output power , S ij represents the branch transmission power between node i and node j, represents the load of node i, Y ij is the line admittance matrix, represents the maximum value of the magnitude of the branch transmission power, represents the maximum value of the phase angle difference between node i and node j, represents the imaginary part of U ij ;

[0018] Since constraint (2) is a non-convex non-linear constraint, use the second-order cone relaxation method to transform constraint (2) into the standard second-order cone inequality form:

[0019]

[0020] |U ij | 2 ≤U ii U jj (14)

[0021]

[0022] By extracting the dual variables of the active power balance equation, the local marginal price λ is calculated i ;

[0023] Traffic network model As shown in the following formula (16):

[0024]

[0025] In the formula, ε respectively represent the sets of traffic nodes and edges; w l , w2 respectively represent the edge weight values regarding distance and branch travel time between nodes v and u; Considering the actual traffic congestion situation, the actual branch travel time can be calculated through the BPR model, and the BPR model is shown in the following formula (17):

[0026]

[0027] In the formula, x (v,u) and c (v,u) respectively represent the traffic flow and traffic capacity of the edge (v, u);

[0028] The charging station model is shown in the following formula (18):

[0029]

[0030] In the formula, L represents the set of charging stations; S(k, t) represents the state matrix of the k-th charging station; and represent that the k-th charging station corresponds to the nodes and positions in the power distribution network and the traffic network respectively; represents the charging demand of the k-th charging station; represents the real-time charging price; represents the waiting time required for electric vehicle charging;

[0031] The electric vehicle model is shown in the following formula (19):

[0032]

[0033] In the formula, t represents the time step of the charging process; Indicates the state of charge of the \(i\)-th electric vehicle; \(v\) represents the node where the electric vehicle is located in the transportation network; \(E\) ini and \(E\) max represent the initial state of charge and the electrical energy capacity respectively;

[0034] The long short-term memory network prediction network model is:

[0035] \(f\) t =\(\sigma(W\) fx x\) t +W\) fh h\) t-1 +b\) f )(20)

[0036] \(i\) t =\(\sigma(W\) ix x\) t +W\) ih h\) t-1 +b\) i )(21)

[0037] \(g\) t =\(\varphi(W\) gx x\) t +W\) gh h\) t-1 +b\) g )(22)

[0038] \(o\) t =\(\sigma(W\) ox x\) t +W\) oh h\) t-1 +b\) o )(23)

[0039] \(s\) t =\(g\) t \odot i\) t +s\) t-1 \odot f\) t (24)

[0040] \(h\) t =\(\varphi(s\) t )\odot o\) t (25)

[0041] In the formula, \(\{x_1, x_2, \ldots, x\) T-1 , x\) T \} is the sequence input; \(x\) t represents the \(k\)-dimensional subsequence input at time step \(t\); \(g\) t , \(i\) t , \(f\) t , \(o\) trespectively represent candidate values, input gate, forget gate, and output gate; σ and φ respectively represent sigmoid and tanh activation functions; ⊙ represents element-wise multiplication operation; s t represents the updated state of the memory cell, h t represents the hidden state at the current time step.

[0042] Furthermore, the expression of the economic optimization model for the electric vehicle charging navigation is:

[0043]

[0044] The constraints of the economic optimization model for the electric vehicle charging navigation include line operation constraints and charging station selection constraints;

[0045] Among them, the constraints for configuring the economic optimization model of the electric vehicle charging navigation are shown in the following equations (27)-(32):

[0046]

[0047] In the formula, and d a are the actual driving time and distance of the edge respectively; and are the charging power and waiting time respectively; c tr and c ch are the driving and charging costs respectively; if the electric vehicle drives through edge a, the binary variable x a is 1, otherwise it is 0; if charging station k is selected, the binary variable y k is 1, otherwise it is 0; p tr and q t are the cost conversion weighting factors for distance and time respectively; E end represents the remaining power after arriving at the charging station; o ini represents the initial position of the electric vehicle.

[0048] Furthermore, the state variables are designed as follows:

[0049] In the electric vehicle charging navigation task of the power-transportation coupling network, the state should select the environmental indicators that can best reflect the current operating status of the system and are directly related to the actions, and select the road node v, the power state of the electric vehicle the electricity price information of each charging station received by the electric vehicle the electricity price prediction information of each charging station output by the prediction network time step t,

[0050] The observation of the state variable is expressed as equation (33):

[0051]

[0052] Furthermore, the action variable is selected as follows:

[0053] The action variable should be selected as the one that directly affects the reward and state. Therefore, inputting the current state, the charging station k selected by the agent is used as the action variable taken:

[0054] a i = k, k ∈ L (34)

[0055] In the formula, L represents the set of charging stations.

[0056] Furthermore, the reward function is specifically:

[0057] The optimization goal of the agent is to find the economically optimal solution over time steps in the feasible region. Therefore, the reward setting is divided into the reward for traveling to the end charging station and the reward for traveling on the road:

[0058]

[0059] In the formula, E max represents the maximum battery capacity of the electric vehicle.

[0060] Furthermore, the action network selects the action variable based on the state variable, and the policy function of the action network is updated through the stochastic policy gradient algorithm:

[0061]

[0062] In the formula, is the output value of the Q network; logπ(a i |o i ; θ i ) represents an entropy term, and α is the weight value of the entropy term; is for gradient solving of the action network.

[0063] Furthermore, the target value Q target (o i , a i ) of the Q network is composed of the Critic target network V Tar (s′ i ). According to the Bellman equation, the target value of the Q network is shown as the following formula:

[0064]

[0065] Among them, γ represents the discount factor.

[0066] Furthermore, the multi-agent SAC training network structure considering electricity price prediction information simultaneously introduces a Q network and an evaluation network. The target value of the Q network is calculated through the target value of the evaluation network, and then Q target (o i ,a i ) is used to update the parameters φ of the Q network i , so as to minimize the loss function Use V Tar to update the parameters of the evaluation network so as to minimize the loss function and as shown in the following formula:

[0067]

[0068] The system first obtains the action variable based on the state variable through the policy network. Subsequently, the system state X i and the system action variable a i are used as inputs and mapped through the Q network to obtain Its network update parameters are obtained through the backpropagation algorithm of the action network loss function in Equation (36):

[0069]

[0070] In the formula, is the action network loss function.

[0071] The second technical solution adopted by the present invention is:

[0072] An electronic device, the electronic device includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned method for navigating electric vehicle charging based on electricity price prediction.

[0073] The third technical solution adopted by the present invention is:

[0074] A computer-readable storage medium, at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above-mentioned method for navigating electric vehicle charging based on electricity price prediction.

[0075] The fourth technical solution adopted by the present invention is:

[0076] A computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the above-mentioned electric vehicle charging navigation method based on electricity price prediction.

[0077] The beneficial effects of the present invention include:

[0078] (1) By establishing a power-transportation coupling network model and using an electric vehicle charging navigation technology integrating electricity price prediction and multi-agent deep reinforcement learning, the present invention provides electric vehicle users with the selection of the optimal route and the optimal charging station, which helps to improve the travel efficiency of electric vehicles and reduce the driving and charging costs.

[0079] (2) The agents in the present invention enable the system to make automated and intelligent decisions. By interacting with the power-transportation coupling network environment, the agents continuously learn and improve their decisions to adapt to different operating conditions, which reduces the need for manual intervention and improves the autonomy of the system.

[0080] (3) Through intelligent decision-making, the system can autonomously learn and optimize the route and charging station selection to ensure the efficient operation of the system and reduce the total cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the relevant technical solution drawings in the embodiments of the present invention or the prior art. It should be understood that the drawings below only facilitate the clear expression of some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0082] Figure 1 is the flowchart of the steps of the electric vehicle charging navigation method based on electricity price prediction in the embodiment of the present invention;

[0083] Figure 2 is the schematic diagram of the multi-agent SAC training network structure in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0084] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation of the present application. For the step numbers in the following embodiments, they are only set for the convenience of description and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0085] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the embodiments of the present application. The singular forms "a", "the" and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. In addition, unless otherwise clearly defined, words such as "set", "installed", "connected" should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0086] In the description of the present application, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present application.

[0087] In the description of the present application, the meaning of "several" is one or more, the meaning of "multiple" is two or more, and understandings such as "greater than", "less than", "exceeding" do not include the present number, and understandings such as "above", "below", "within" include the present number. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0088] In the description of the present application, " / and / " describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0089] Term Explanation:

[0090] SAC: Abbreviation for Soft Actor-Critic; SAC is an algorithm based on the maximum entropy reinforcement learning framework. By introducing entropy maximization in the policy optimization process, it balances exploration and exploitation, enabling the agent to explore and learn more effectively in complex environments.

[0091] Aiming at the existing technical problems, the present invention provides an electric vehicle charging navigation method for a power-traffic coupled network integrating electricity price prediction and multi-agent deep reinforcement learning, aiming to reduce the economic driving and charging costs of electric vehicles by means of data intelligence methods. The key points of its technical solution are to reasonably construct a multi-agent deep reinforcement learning scheduling framework suitable for electric vehicle charging navigation in the power-traffic coupled network, including selecting state variables, action variables, designing constraint indicators, and reward functions when the electric vehicle operates in the power-traffic coupled network; by interacting with real-time data, the system can adapt to changing environmental conditions and user needs, cope with the volatility of electricity prices, traffic conditions, and charging demands, and realize the optimization of electric vehicle scheduling and improve the overall efficiency of the network. The application fields of the present invention cover multiple fields such as electric vehicle charging navigation, prediction network, multi-agent reinforcement learning, and scheduling and control of power-traffic coupled networks. It improves the stable economic operation level of the power-traffic coupled network.

[0092] Embodiment 1

[0093] As Figure 1 shown, this embodiment provides an electric vehicle charging navigation method integrating electricity price prediction and multi-agent deep reinforcement learning, including the following steps:

[0094] Step S1: Establish a mechanism model of the power-traffic network coupling, including a distribution network model, a traffic network model, a charging station model, and an electric vehicle model;

[0095] Step S2: According to the power-traffic coupling network model, establish an economic optimization model for electric vehicle charging navigation, and clarify system variables and constraints; construct a reinforcement learning training model framework based on variables, constraints, and indicators, that is, design a state variable set, an action variable set, and a reward function in multi-agent deep reinforcement learning.

[0096] Step S3: Build a multi-agent SAC training network structure considering electricity price prediction information, set the number of training rounds, the number of network layers, the network size, and the sequence length of the electricity price prediction network, and set the hidden layer size, the number of hidden layers, the activation function, the learning rate, the entropy coefficient, the batch size, the discount factor, and the replay buffer size of the action network and the evaluation network of the SAC training network structure;

[0097] Step S4: Train the agent through the power - transportation coupling network environment model so that it learns how to make the best decisions in different situations to maximize the reward function, thereby optimizing the cost of electric vehicle users during driving and charging and achieving an economic operation level.

[0098] As an implementation, in step S1, the model is established as follows:

[0099] The optimal power flow model of the distribution network is shown in equations (1)-(10):

[0100]

[0101] Among them, the constraints for configuring the optimal power flow model of the distribution network mainly include power balance constraints and voltage constraints.

[0102] The power balance constraint is shown in equation (5):

[0103]

[0104] Among them, c g [$] represents the electricity purchase cost, represents the active power output of the node generator, represents the power injected from other networks, a 0,i , a 1,i , ρ[$ / kWh] represents their correlation coefficients; V i [kV] is the AC voltage of node i, S ij [kVA] represents the branch transmission power between node i and node j, represents the load of node i, Y ij [S] is the line admittance matrix, represents the imaginary part of U ij .

[0105] Further preferably, since constraint (2) is a non - convex non - linear constraint, therefore, constraint (2) can be convexified instead of dealing with it according to the principle of semi - definite programming. By using the second - order cone relaxation method, constraint (2) can be transformed into the standard second - order cone inequality form:

[0106]

[0107] |U ij | 2 ≤U ii U jj (14)

[0108]

[0109] Further preferably, by extracting the dual variables of the active power balance equation, the local marginal price λ can be calculated. i [$ / kWh].

[0110] Traffic network model As shown in the following equation (16):

[0111]

[0112] Where ε represents the sets of traffic nodes and edges respectively; w l [km], w2[h] represent the edge weight values of distance and branch travel time between nodes v and u respectively. The actual branch travel time can be calculated through the BPR model The BPR model is as shown in the following equation (17):

[0113]

[0114] In the formula, x (v,u) and c (v,u) represent the traffic flow and traffic capacity of the edge (v, u) respectively.

[0115] The charging station model is as shown in the following equation (18):

[0116]

[0117] Where L represents the set of charging stations; S(k, t) represents the state matrix of the kth charging station; and represent the nodes and positions corresponding to the kth charging station in the power distribution network and the traffic network respectively; represents the charging demand of the kth charging station; represents the real-time charging price; represents the waiting time required for electric vehicle charging.

[0118] The electric vehicle model is as shown in the following equation (19):

[0119]

[0120] Where t[h] represents the time step of the charging process; represents the state of charge of the ith electric vehicle; v represents the node where the electric vehicle is located in the traffic network; E ini [kWh] and E max [kWh] represent the initial state of charge and the electrical energy capacity respectively.

[0121] The long short-term memory network prediction network model is:

[0122] ft = σ(W fx x t + W fh h t-1 + b f ) (20)

[0123] i t = σ(W ix x t + W ih h t-1 + b i ) (21)

[0124] g t = φ(W gx x t + W gh h t-1 + b g ) (22)

[0125] o t = σ(W ox x t + W oh h t-1 + b o ) (23)

[0126] s t = g t ⊙ i t + s t-1 ⊙ f t (24)

[0127] h t = φ(s t ) ⊙ o t (25)

[0128] where, {x1, x2, …, x T-1 , x T} is the sequence input; x t represents the k-dimensional subsequence input at time step t; g t , i t , f t , o t represent the candidate value, input gate, forget gate, and output gate respectively; σ and φ represent the sigmoid and tanh activation functions respectively; ⊙ represents the element-wise multiplication operation.

[0129] As an implementation, the economic optimization model of the electric vehicle charging navigation is:

[0130]

[0131] Among them, the constraints for configuring the economic optimization model of electric vehicle charging navigation include line operation constraints and charging station selection constraints. The constraints for configuring the economic optimization model of electric vehicle charging navigation are shown in the following equations (27)-(32):

[0132]

[0133] Among them, and d a [km] are the actual driving time and distance of the edge respectively; and are the charging power and waiting time respectively; c tr [$] and c ch [$] are the driving and charging costs respectively; p tr [km / h] and q t [$ / h] are the cost conversion weighting factors of distance and time respectively; E end [kWh] represents the remaining power after arriving at the charging station.

[0134] The design of the deep reinforcement learning framework involves four key elements: state, action, reward, and environment. The state represents the situation where the agent is located, the action is the operation that the agent can execute, the reward is the immediate feedback, and the environment is the external world. In deep reinforcement learning, the agent observes the state, selects actions, receives rewards, and interacts with the environment to learn how to formulate strategies to maximize the cumulative reward.

[0135] 1) State variable design. Under the electric vehicle charging navigation task in the power-traffic coupled network, the state should select the environmental indicators that can best reflect the current operating conditions of the system and have a direct relationship with the actions. In this paper, the road node v, the power state of the electric vehicle the electricity price information of each charging station received by the electric vehicle the electricity price prediction information of each charging station output by the prediction network the time step t

[0136] The observation of the state variable can be expressed as Equation (33):

[0137]

[0138] 2) Action variable design. The action variable should select the variable that directly affects the reward and the state. Therefore, input the current state, and take the charging station k selected by the agent as the action variable:

[0139] a i = k, k ∈ L (34)

[0140] 3) Reward function design. The optimization goal of the agent is to find the economically optimal solution in the feasible region over time steps. Therefore, the reward is set as follows. Equation (13) is divided into the reward for driving to the end charging station and the reward for driving on the road.

[0141]

[0142] As a specific implementation, in step S3, a multi-agent SAC training network structure considering electricity price prediction information is built as Figure 2 shown. The training parameters of the multi-agent SAC training network structure are shown in Table 1.

[0143] In Figure 2 , the set of power-transportation coupling network states at time step t is described. These states are mapped to the energy system decision variable a through the policy function i . The agent interacts with the environment at the next moment, generating a new state s′ i , and the economic cost -r required for the coupling network at the next moment i . This information is stored in the experience pool for random sampling during network training as training samples.

[0144] Figure 2 The role of the action network adopted in is to select the action variable based on the state variable. The policy function of the action network is updated through the stochastic policy gradient algorithm:

[0145]

[0146] where is the output value of the Q network; logπ(a i ∣o i ; θ i ) represents an entropy term, and α is the weight value of the entropy term; is for the action network to solve the gradient.

[0147] The target value Q of the Q network target (o i , a i ) is composed of the Critic target network V Tar (s′ i ). According to the Bellman equation, the target value of the Q network is shown as follows:

[0148]

[0149]

[0150] where γ represents the discount factor.

[0151] The multi-agent SAC training network structure considering electricity price prediction information introduces both a Q network and a critic network. The target value of the Q network is calculated through the target value of the critic network, and then Q target (o i ,a i ) is used to update the parameters φ of the Q network i to minimize the loss function V is used Tar to update the parameters of the critic network to minimize the loss function and as shown in the following formula:

[0152]

[0153] First, the system obtains the action variable based on the state variable through the policy network. Subsequently, the system state X i and the system action variable a i are used as inputs and, after being mapped by the Q network, obtain Its network update parameters are obtained through the backpropagation algorithm of the action network loss function in Equation (36):

[0154]

[0155] where is the action network loss function

[0156] The objective function adopts the soft update method to gradually update the parameters of the policy network to the target policy network and, at the same time, gradually update the parameters of the critic network to the target critic network. The learning rate τ is introduced, and its update strategy is:

[0157]

[0158] Table 1 Training parameter settings for the multi-agent SAC training network structure

[0159]

[0160]

[0161] Exemplarily, agent training is carried out based on the Python platform. The number of training rounds is set to 3000, and the condition for saving the agent is set to save the agent every 1000 rounds of training. The final condition for stopping training is set to reach the maximum number of training episodes. The experimental process is divided into two steps. In the first step, the agent is trained to avoid various network constraints and achieve optimal decision-making. After convergence, the updated parameters of the network are saved. In the second step, the decision-making performance of the agent is tested, and the test is carried out on the test dataset after loading the network parameters.

[0162] Example 2

[0163] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a method for electric vehicle charging navigation based on electricity price prediction as Figure 1 shown.

[0164] It can be understood that the memory may include a random access memory (RAM), or may also include a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.

[0165] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts within the entire server. By running or executing instructions, programs, code sets or instruction sets stored in the memory, and by calling data stored in the memory, the processor executes various functions of the server and processes data. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate a central processing unit (CPU) and a modem, etc. in one or several combinations. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor and may be implemented separately by a single chip.

[0166] Since this electronic device is an electronic device corresponding to a method for electric vehicle charging navigation based on electricity price prediction in an embodiment of the present invention, and the principle of the electronic device for solving problems is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0167] Example 3

[0168] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement Figure 1 a method for electric vehicle charging navigation based on electricity price prediction as shown in

[0169] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, tape memories, or any other computer-readable medium capable of carrying or storing data.

[0170] Since this storage medium is the storage medium corresponding to the method for electric vehicle charging navigation based on electricity price prediction in the embodiment of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0171] Example 4

[0172] In some possible embodiments, aspects of the method of the embodiments of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a method for electric vehicle charging navigation based on electricity price prediction according to various exemplary embodiments described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in high-level programming languages such as C, C++, Python, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0173] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0174] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0175] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and their purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly. It should not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.

Claims

1. An electric vehicle charging navigation method based on electricity price prediction, characterized in that, It includes the following steps: Establish a power - transportation coupling network model according to the distribution network, transportation network, charging stations and electric vehicle model; Establish an economic optimization model for the charging navigation of electric vehicles according to the power - transportation coupling network model, and determine the variables and constraints of the model; Design the state variable set X, action variable set A and reward function R in multi - agent deep reinforcement learning according to the variables and constraints; Construct a multi - agent SAC training network structure considering electricity price prediction information, and set the parameters of the action network and evaluation network of the SAC training network structure; Through the interactive power - transportation coupling network model, train the agent to learn to make optimal decisions under actual traffic conditions and dynamic electricity prices to maximize the reward function, thereby realizing the cost optimization of electric vehicle users during driving and charging.

2. The electric vehicle charging navigation method based on electricity price prediction according to claim 1, characterized in that The establishment of the power - transportation coupling network model according to the distribution network, transportation network, charging stations and electric vehicle model includes: The optimal power flow model of the distribution network is shown in the following formula (1): Among them, the constraints for configuring the optimal power flow model of the distribution network include power balance constraints and voltage constraints; The power balance constraints are shown in formulas (2)-(10): where c g represents the electricity purchase cost, represents the active power output of the node generator, represents the power injected from other networks, a 0,i , a 1,i , ρ represents their correlation coefficients; V i is the AC voltage of node i, V i , respectively represent the upper and lower bounds of V i , respectively represent the upper and lower bounds of the generator output power , S ij represents the branch transmission power between node i and node j, represents the load of node i, Y ij is the line admittance matrix, represents the maximum value of the magnitude of the branch transmission power, represents the maximum value of the phase angle difference between node i and node j, represents the imaginary part of U ij ; Since constraint (2) is a non - convex non - linear constraint, use the second - order cone relaxation method to transform constraint (2) into the standard second - order cone inequality form: By extracting the dual variables of the active power balance equation, the local marginal price λ is calculated i ; Transportation network model As shown in the following formula (16): In the formula, ε respectively represents the sets of traffic nodes and edges; w l , w2 respectively represent the edge weight values regarding distance and branch travel time between nodes v and u; considering the actual traffic congestion situation, the actual branch travel time can be calculated through the BPR model, and the BPR model is shown in the following formula (17): where x (v,u) and c (v,u) represent the traffic flow and traffic capacity of edge (v, u), respectively; The charging station model is shown in the following formula (18): Wherein, L represents the set of charging stations; S(k,t) represents the state matrix of the k-th charging station; and represents that the k-th charging station corresponds to the nodes and positions in the power distribution network and the transportation network respectively; represents the charging demand of the k-th charging station; represents the real-time charging price; represents the waiting time required for electric vehicle charging; The electric vehicle model is shown in the following formula (19): where \(t\) represents the time step of the charging process; represents the state of charge of the \(i\)-th electric vehicle; \(v\) represents the location node of the electric vehicle in the transportation network; \(E\) ini and \(E\) max represent the initial state of charge and the electrical energy capacity respectively; The long - short - term memory network prediction network model is: f t = σ(W fx x t + W fh h t-1 + b f ) (20) i t = σ(W ix x t + W ih h t-1 + b i ) (21) g t = φ(W gx x t + W gh h t-1 + b g ) (22) o t = σ(W ox x t + W oh h t-1 + b o ) (23) s t = g t ⊙ i t + s t-1 ⊙ f t (24) h t = φ(s t ) ⊙ o t (25) wherein, {x1, x2, …, x T-1 , x T} is the sequence input; x t represents the k-dimensional subsequence input at time step t; g t , i t , f t , o t respectively represent the candidate value, input gate, forget gate, and output gate; σ and φ respectively represent the sigmoid and tanh activation functions; ⊙ represents the element-wise multiplication operation; s t represents the updated state of the memory cell, and h t represents the hidden state at the current moment.

3. A method for electric vehicle charging navigation based on electricity price prediction according to claim 1, characterized in that The expression of the economic optimization model for the charging navigation of electric vehicles is: The constraints of the economic optimization model for the charging navigation of electric vehicles include line operation constraints and charging station selection constraints; Among them, the constraints for configuring the economic optimization model for the charging navigation of electric vehicles are shown in the following formulas (27)-(32): wherein, and d a are the actual driving time and distance of the edge, respectively; and are the charging power and waiting time, respectively; c tr and c ch are the driving and charging costs, respectively; if the electric vehicle drives through edge a, the binary variable x a is 1, otherwise it is 0; if charging station k is selected, the binary variable y k is 1, otherwise it is 0; p tr and q t are the cost conversion weighting factors of distance and time, respectively; E end represents the remaining power after arriving at the charging station; o ini represents the initial position of the electric vehicle.

4. A method for electric vehicle charging navigation based on electricity price prediction according to claim 3, characterized in that, The state variables are designed as follows: Under the electric vehicle charging navigation task in the power-transportation coupled network, the state should select the environmental indicators that can best reflect the current operating conditions of the system and are directly related to the actions, and select the road node v and the power state of the electric vehicle The electricity price information of each charging station received by the electric vehicle The electricity price prediction information of each charging station output by the prediction network Time step t The observation of the state variables is expressed as formula (33):

5. A method for electric vehicle charging navigation based on electricity price prediction according to claim 3, characterized in that The action variables are selected as follows: The action variables should be selected as variables that directly affect the reward and state. Therefore, input the current state, and take the charging station k selected by the agent as the action variable: a i = k, k ∈ L (34) In the formula, L represents the set of charging stations.

6. The method for predicting the charging navigation of an electric vehicle based on electricity price prediction according to claim 3, wherein The reward function is specifically: The optimization goal of the agent is to find the economic optimal solution in the feasible region as time steps progress. Therefore, the reward settings are divided into the reward for driving to the end charging station and the reward for driving on the road: where E max represents the maximum battery capacity of the electric vehicle.

7. A method for electric vehicle charging navigation based on electricity price prediction according to claim 1, characterized in that, The action network selects action variables based on state variables, and the policy function of the action network is updated through the stochastic policy gradient algorithm: wherein is the output value of the Q network; logπ(a i |o i ; θ i ) represents an entropy term, and β is the weight value of the entropy term; is for gradient solving of the action network.

8. A method for electric vehicle charging navigation based on electricity price prediction according to claim 7, characterized in that The target value Q of the Q-network target (o i ,a i ) is composed of the Critic target network V Tar (s′ i ). According to the Bellman equation, the target value of the Q-network is shown as follows: Among them, γ represents the discount factor.

9. The method for electric vehicle charging navigation based on electricity price prediction according to claim 1, characterized in that, The multi-agent SAC training network structure considering electricity price prediction information simultaneously introduces a Q network and a critic network. The target value of the Q network is calculated through the target value of the critic network, and then Q target (o i ,a i ) is used to update the parameters φ of the Q network i , so as to minimize the loss function V is used Tar to update the parameters of the critic network to minimize the loss function and as shown in the following formula: The system first obtains the action variable based on the state variable through the policy network. Subsequently, the system state X i and the system action variable a i are used as inputs. After being mapped by the Q network, its network update parameters are obtained through the backpropagation algorithm of the action network loss function in Equation (36): In the formula, is the action network loss function.

10. An electronic device, characterized in that, The electronic device includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Electric vehicle charging guide optimization method based on graph neural network reinforcement learning

    CN114444802A

  • Power distribution network operator electricity selling pricing method and system for charging station considering traffic network

    CN117934025A