Electric vehicle path planning method based on evolutionary reinforcement learning in dual-network fusion scene
By adopting the method of double-layer replay buffer structure and timing difference value in electric vehicle path planning, the problem of excessive training time in the existing technology is solved, and more efficient path planning is achieved, saving simulation analysis time.
Patent Information
- Application Number
- CN202510143309.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, due to the large search space for construction and the long training time, the optimal solution efficiency of the early search in electric vehicle path planning is low, and the optimization results are unsatisfactory.
The double-layer replay buffer structure is adopted, and the timing difference value is used to ensure that the value of early selection of data used to train neural networks is good enough, thereby improving the overall training efficiency, and the optimization algorithm can find satisfactory solutions faster.
By improving training efficiency and saving performance simulation analysis time, we can find the optimal path for electric vehicle path planning faster, solving the time-consuming problems in the existing technology.
Smart Images

Figure CN120146335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to an electric vehicle path planning method based on evolutionary reinforcement learning in a dual-network fusion scenario. Background Art
[0002] In recent years, with the continuous development of science and technology, more and more cities in China have begun to promote the "dual-network fusion" of the power grid and the transportation network to improve the living experience of the people. The power grid is an overall system that includes power generation stations, transmission lines, charging stations, etc., and can realize functions such as power transformation and power transmission. The transportation network is the basis of the national comprehensive transportation system, which is related to various transportation modes such as railways, highways, waterways, civil aviation, and pipelines. An important application carrier of the dual-network fusion is electric vehicles. This new type of transportation method that uses clean energy has received increasing attention in recent years and has become one of the hot topics in the transportation field. Compared with traditional fuel vehicles, when applying electric vehicles to practice, an issue that cannot be ignored is the arrangement of charging stations - currently, such energy replenishment sites are not widespread enough, which has led to a variant type of the traditional vehicle path planning problem - the electric vehicle path planning problem. The ultimate goal of this type of problem is to enable the vehicle to complete tasks such as delivering goods and picking up goods at the lowest cost on the premise that there are charging stations on the map, and to ensure that the vehicle can maintain an appropriate battery level during the task completion.
[0003] Evolutionary reinforcement learning is a new type of method that combines the principles of traditional reinforcement learning, deep learning, and evolutionary algorithms. It is a more effective means than traditional intelligent algorithms in solving complex optimization problems in reality and has currently been applied to various intelligent optimization problems. For the electric vehicle path planning problem, this intelligent algorithm can very stably find an optimal optimization solution from a global perspective. However, evolutionary reinforcement learning requires constructing a large number of neural networks and performing tens of thousands or even hundreds of thousands or millions of fitness value evaluations and accumulations on each candidate solution to achieve the optimization process of survival of the fittest in the biological population. In the electric vehicle path planning problem, every time the vehicle arrives at a new location, the cost consumed by the background resources needs to be calculated and stored in the replay buffer in the algorithm structure together with the previous state, the previous action, and the current state for use as data for training the neural network parameters in the future. Since it will affect the training time of the entire algorithm framework, the quality of the data selected from the replay buffer is crucial. If the data selected in the initial stage of training cannot effectively train the neural network in evolutionary reinforcement learning, then the time required to simulate and analyze to obtain a feasible solution may reach dozens of minutes or even hours. Such a time consumption for solving the optimization problem is too large to be completed within an acceptable range. Therefore, how to use this new type of intelligent algorithm framework of evolutionary reinforcement learning to complete the path planning problem of electric vehicles in transportation-related applications in a relatively short time is still a difficult problem to be solved. Summary of the Invention
[0004] To overcome the above-mentioned drawbacks and deficiencies of the prior art, the purpose of the present invention is to provide an electric vehicle path planning method based on evolutionary reinforcement learning in a dual-network fusion scenario, so as to solve the problems in the prior art of the electric vehicle path planning field, such as the low efficiency of searching for the optimal solution in the early stage and the unsatisfactory optimization results due to the too large constructed search space and the too long training time.
[0005] The double-layer replay buffer structure adopted by this method can use the temporal difference value to ensure that the value of the data selected for training the neural network in the early stage is good enough, so as to improve the overall training efficiency, enable the optimization algorithm to find a solution satisfactory to the requester faster, save the performance simulation analysis time, and drive the evolution of the entire population with the training results, and iteratively optimize the path traveled by the electric vehicle to complete the transportation task.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] An electric vehicle path planning method based on evolutionary reinforcement learning in a dual-network fusion scenario, comprising:
[0008] Determine whether the type of task to be performed by the electric vehicle is delivery or pick-up;
[0009] Obtain the basic data of the vehicle dispatching center, transportation task points, and charging stations in the current task. The basic data specifically refers to their locations and the quantity of goods.
[0010] Initialize the reinforcement learning population RL and the evolutionary population P, where the number of individuals included in RL is approximately 1 / 10 to 1 / 5 of the number of individuals in P. The specific value is determined by the scale of the problem model to be solved. Each individual in the reinforcement learning population RL is a double-layer value network framework, including a value neural network Q and a target value neural network Q'. Each individual in the evolutionary population P only represents a value neural network Q, and each individual represents a path planning;
[0011] Define two replay buffers R1 and R2 to store experience tuples in the form of (s t ,a t ,s t+1 ,r t ), where the four items of data represent the real-time state, action taken, future state, and obtained reward value at time point t in sequence;
[0012] Taking reducing the energy consumed by the electric vehicle and shortening the time to complete the transportation task as the optimization goal, and on the premise of successfully completing the transportation task, iteratively update the weight parameters of the neural network in the population, so as to obtain the optimal path for the electric vehicle to execute a transportation task.
[0013] Furthermore,
[0014] the vehicle dispatching center: controls the driving trajectory of the electric vehicle;
[0015] The transportation task point: represents a transportation task to be completed, picking up or delivering goods, corresponding to a numerical value, and the numerical value represents the quantity of goods;
[0016] The charging station: The single - charging duration is in a direct proportional relationship with the consumed power when the vehicle starts charging this time.
[0017] Furthermore, in the electric vehicle path planning, the following conditions are set:
[0018]
[0019] 0 ≤ u 0 ≤ C (9)
[0020]
[0021] where the optimization objective (1) is to minimize the driving energy consumption and the total time, λ 1 and λ 2 are two balancing weights to ensure that two numbers with different magnitudes can be added; Constraint (2) ensures that each transportation task point can be visited only once; Constraint (3) stipulates that each charging station can be visited at most once to avoid frequent charging in a short period; Constraint (4) represents the conservation of vehicle flow; Constraints (5) and (6) define the balance relationship of the quantity of goods for continuously visiting two points i and j, where M is a sufficiently large positive number; Constraint (7) is the time - window constraint, e j and l j represent the earliest and latest required visiting times of transportation task point j respectively, and it can cooperate with Constraints (5) and (6) to prevent the occurrence of sub - circuits; Constraints (8) and (10) respectively ensure that the quantity of goods carried by the vehicle and the vehicle's power can be updated normally after completing any transportation task; Constraints (9) and (11) respectively ensure that the quantity of goods carried by the vehicle and the vehicle's power are within the limits that conform to the real - world logic.
[0022] Furthermore, the specific training process for each generation in the reinforcement learning population RL is as follows:
[0023] According to the optimal fitness value f returned by the previous generation p calculate the ratio p of selecting experience tuples from the double - layer replay buffer rb , and accordingly select T experience tuples from the two replay buffers;
[0024] Take out the form (s t ,a t ,s t+1 ,rt ) Calculate the loss function L for the T empirical tuples to update the weight parameters of the value neural network Q in the population. The formula for the loss function is:
[0025] L = ∑ t (y t - Q(s t , a t |θ Q )) 2 / T
[0026] Where y t = r t + γQ′(s t+1 , a′ t+1 |θ Q′ ), γ is the discount factor for the rewards obtained in the past for the current state, and the value can be set from 0.9 to 0.99. a′ t+1 is the action that should be selected under the state s t+1 judged by q′, and θ Q refers to all the parameters in the neural network Q;
[0027] Update the weight parameters of the target value neural network Q′ through the value neural network Q and the soft update hyperparameter τ. The update formula is:
[0028] q′ = (1 - τ)q′ + τq
[0029] Where q is a certain parameter in Q, q′ is the parameter with the same meaning as q in Q′, and τ is set to 0.01 to prevent missing the global optimum due to overly large updates.
[0030] Furthermore, the evolutionary population P iteratively improves the population quality by taking population optimization as the core and through the actions of selection, crossover, and mutation operators. Specifically, it includes:
[0031] Evolve the population and simulate the problem-solving model; starting from the initial state s 0 , in each individual, continuously select the action with the highest value in the simulation to update the state according to the judgment of the value neural network Q, and return the cumulative reward obtained throughout the process as the fitness value of this individual; during this process, the tuples storing information such as the sealed current state s t , the action taken a t , the future state s t+1 and the obtained reward value r t are stored as empirical data in the double-layer replay buffer, and which replay buffer to finally enter is determined by the temporal difference value TD corresponding to this tuple. The formula for this value is:
[0032] TD(s t , a t , s t+1,r t ) = r t +γQ′(s t+1 ,a′ t+1 |θ Q′ ) - Q(s t ,a t |θ Q );
[0033] Sort the individuals in the evolutionary population P from the best to the worst according to the fitness value, and select the top 10% as elite individuals for subsequent operations;
[0034] Use the selection operator to select the parent individuals for the crossover operator. The higher the fitness value of the individuals in the population, the higher the probability of being selected;
[0035] Execute the crossover operator to generate new offspring individuals and replace some non-elite individuals;
[0036] Execute the mutation operator on non-elite individuals to change the weight parameter value of Q in the individuals, ensure that the diversity exploration always exists, and prevent falling into local optima;
[0037] Replace an equal number of the worst individuals in the evolutionary algorithm part with all the individuals in the reinforcement learning population RL.
[0038] Furthermore, in each iteration, a large amount of experience generated by the evolutionary population P after population evolution is used to train the reinforcement population RL, and then the training results of RL are also injected into P. And with the iterative mechanism of "survival of the fittest" in the evolutionary population training method, if the RL individuals entering P for training are not excellent enough, they will be eliminated and will not have a negative impact on the training results of P.
[0039] Furthermore, the ratio p of the selected experience tuples rb , specifically:
[0040] p rb is a parameter that is adaptively updated during training, with a maximum value set to 0.75 and a minimum value set to 0.25, and linearly decreases as the optimal fitness value f p of the previous generation increases. Select T experience tuples from the two replay buffers according to the ratio p rb , where T can be set according to the problem scale and is generally defaulted to 32.
[0041] Furthermore, the capacities of the two replay buffers R1 and R2 are set to 2000 and 6000 respectively.
[0042] Furthermore, in the form of (s t ,a t ,s t+1 ,r tThe tuples of ( ) are stored as empirical data in R1 and R2, and which replay buffer the tuples finally enter is determined by the temporal difference value TD corresponding to the tuples. The quarter with smaller TD values is stored in R1, and the other three - quarters are saved in R2.
[0043] Furthermore, the mutation operator is executed on non - elite individuals to change the weight parameter values of Q in the individuals. Specifically:
[0044] First, the hyperparameter mut frac determines the proportion of the number of parameters for mutating the value neural network. In the present invention, it is set to 0.1. Then, for each non - elite individual, it is judged whether to perform the mutation operation by randomly generating a number. In the present invention, the mutation probability is set to 0.3. Finally, randomly select mut frac parameters from the weight matrix of the determined - to - mutate individuals, and generate a corresponding random number for each selected parameter to determine its mutation type.
[0045] The random number is between 0 and 1. When the random number is between 0 and 0.9, a small perturbation is added to the weight parameter;
[0046] When the random number is between 0.9 and 0.95, a larger perturbation is added;
[0047] When the random number is between 0.95 and 1, the parameter is directly regenerated.
[0048] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0049] The electric vehicle path planning method based on the evolutionary reinforcement learning algorithm provided by the present invention uses the obtained electric vehicle states, actions and other data as the experience for the training vehicle to find the optimal path, quantitatively calculates the path consumption, takes the neural network as the core of the algorithm framework, takes the electric vehicle energy consumption and the time required to complete the transportation task as the objective function, and uses a variety of genetic operators of the evolutionary algorithm to improve the global search ability for the variable space, solving the problem of poor globality in the prior art when solving the electric vehicle path planning problem.
[0050] The present invention also uses the interaction mechanism of reinforcement learning to achieve better performance. Furthermore, the double - layer replay buffer structure provided based on the reinforcement learning mechanism uses the temporal difference value as an aid, selects more empirical tuples that are easy to make the loss function L for updating the value neural network converge in the early stage of training to improve the early - stage training effect of the algorithm, thereby further improving its performance in solving the electric vehicle path planning problem, saving the simulation analysis time in practical applications, and solving the problem of huge time consumption in the prior art when solving the electric vehicle path planning problem.
[0051] The present invention proposes an adaptive parameter p rbTo standardize the proportion of empirical data selected by the double-layer replay buffer under different training conditions. This is because relatively good optimization results have been achieved in the later stage of training. At this time, exploring the unknown areas of parameter values is more important. Excessive focus on the temporal difference value will instead weaken the exploration and lead to local optimality. With the control of the adaptive parameter p rb the algorithm can achieve a balance between exploration and exploitation throughout the training process, improve the global exploration and development capabilities for the electric vehicle path planning problem, and solve the problems of low exploration efficiency and poor optimization results in the prior art when solving this problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 FIG. is an example diagram of an electric vehicle path planning problem provided by an embodiment of the present invention;
[0053] Figure 2 FIG. is a structural diagram of an algorithm framework of an electric vehicle path planning method based on improved evolutionary reinforcement learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The present invention will be further described in detail below with reference to the embodiments, but the embodiments of the present invention are not limited to this embodiment.
[0055] As Figure 1 shown, in a double-network fusion scenario of the present method, an electric vehicle path planning method based on evolutionary reinforcement learning, the electric vehicle management system includes:
[0056] Vehicle dispatching center: Controls the driving trajectory of the electric vehicle. Each time a transportation task starts to be executed, the vehicle departs from here and also returns here after completing the transportation task. It may also temporarily return here during the transportation task to replenish or unload goods;
[0057] Transportation task point: Represents a transportation task to be completed, either picking up or delivering goods, corresponding to a value, and the value represents the quantity of goods;
[0058] Charging station: The single charging duration is in a proportional relationship with the consumed power when the vehicle starts charging this time.
[0059] In an embodiment of the present invention, V = {1, 2,..., n} is used to represent the transportation task points, where n is the number of transportation task points; the transportation tasks are uniformly set as one of picking up or delivering goods, the former is set as P, and the latter is set as D, which is used as one of the parameters when importing the problem model; F is used to represent the set of charging stations; vertices 0 and n + 1 represent the starting point and return point of the transportation task points. Each complete route for the electric vehicle to complete the transportation task starts from 0 and ends at n + 1. For the convenience of writing the constraint formula, in Embodiment 1 of the present invention, V' = V U F, F 0 = F U {0}, V0 ' = V' ∪ {0}, V 0,n+1 ' = V' ∪ {0, n + 1}; For every two different optional locations i and j in the two problem models (vehicle dispatching center, transportation task point, and charging station, i.e., V 0,n+1 ) there is an electricity consumption e for driving i,j and a driving time t i,j ; The loading capacity of each electric vehicle is C, the battery capacity is B, and the vehicle starts from the dispatching center with a full battery each time; For each vertex i ∈ V 0,n+1 ', there exists a service time s i , that is, the time the vehicle needs to stay after arrival, and there is also a required cargo volume q i , it is defined that when i is not in V, q i is 0; The charging efficiency of each charging station is g, and the charging amount is the difference between B and the current battery power of the vehicle; Each transportation task point does not accept partial pick-up and delivery, and can only be visited once. The time for the vehicle to reach vertex i ∈ V 0,n+1 ' is defined as τ i , the cargo volume carried on the vehicle when arriving is defined as u i , and the remaining battery power when arriving is defined as y i ; Embodiment 1 of the present invention also defines a binary decision variable x ij (i ∈ V 0 ', j ∈ V n+1 ', i ≠ j), its value is 1 when visited from i to j, otherwise 0. Based on reality, the electric vehicle path planning problem of the embodiment of the present invention has the following objective and constraint formulas:
[0060]
[0061] 0 ≤ u 0 ≤ C (9)
[0062]
[0063] Among them, the optimization objective (1) is to minimize the driving energy consumption and total time, λ 1 and λ 2 are two balancing weights to ensure that two numbers with different magnitudes can be added; Constraint (2) ensures that each transportation task point can only be visited once; Constraint (3) stipulates that each charging station can be visited at most once to avoid frequent charging in a short period; Constraint (4) represents the conservation of vehicle flow; Constraints (5) and (6) define the balance relationship of the cargo volume for continuously visiting two points i and j, where M is a sufficiently large positive number; Constraint (7) is the time window constraint, e j and l jrespectively represent the earliest and latest required visit times of transportation task point j, which can cooperate with constraints (5) and (6) to prevent the occurrence of sub-circuits; constraints (8) and (10) respectively ensure that the cargo carried by the vehicle and the vehicle's battery power can be updated normally after completing any transportation task; constraints (9) and (11) respectively ensure that the cargo carried by the vehicle and the vehicle's battery power are within the limits that conform to the real logic.
[0064] The algorithm framework used in the embodiments of the present invention is proposed based on evolutionary reinforcement learning. Essentially, it is an improved reinforcement learning. Therefore, the problem model to be solved, that is, the path planning problem of electric vehicles, also needs to be established based on the Markov decision process (MDP) commonly used in reinforcement learning algorithms. The state s at time t t includes the current location j of the vehicle t , the amount of cargo carried by the current vehicle and the remaining battery power of the current vehicle The action a selected at time t t is judged by the value neural network in the population, and the optional action with the highest value is selected. Essentially, it is the location that the electric vehicle will go to in the next step at the current time. The optional items are each transportation task point and charging station that the vehicle has not been to yet; when the electric vehicle makes an action selection at time t, the next state is the location, the amount of cargo, and the remaining battery power after reaching the selected location and ending the relevant service; the corresponding reward r(s t , a t ) is determined by the objective function (1) and is directly set to the opposite of the weighted sum of the driving energy consumption and the total time consumption. Thus, the minimization problem is converted into a maximization problem. In addition, every time an action is executed in a specific state, it is necessary to detect whether there is a violation of constraints (2) to (11). If so, a large penalty P 0 will be given; when initializing the problem model each time, it is necessary to confirm whether the task at this time is pick-up or delivery through a specific parameter Type. The former is P and the latter is D. The corresponding initial states s 0 are (0, 0, C) and (0, Q, C) respectively.
[0065] In the embodiments of the present invention, an optimization algorithm for electric vehicle path planning based on evolutionary reinforcement learning is implemented. It finds a path that is relatively satisfactory in terms of both energy consumption and time consumption and can complete all transportation tasks through gradient training, population evolution, and iterative search.
[0066] As Figure 2 shown, the specific path planning process includes:
[0067] According to the dual-network integrated electric vehicle management system, determine whether the type of task to be executed by the vehicle is delivery or pick-up, import the problem model, and enter the path optimal search process as follows:
[0068] S1: Initialize the population, including the reinforcement learning population RL and the evolutionary population P. The number of individuals in RL is approximately 1 / 10 to 1 / 5 of the number of individuals in P, and the specific value is determined by the scale of the problem model to be solved. Since the action space of the problem to be solved is discrete, each individual in RL is not a conventional actor-critic neural network system, but a two-layer value network framework, which includes a value neural network Q and a target value neural network Q'. Each individual in P only contains a value neural network Q, and each individual represents a path planning;
[0069] S2: Define two empty replay buffers R1 and R2 to store experience data, and their capacities are set to r 1 = 2000 and r 2 = 6000 to store experience tuples in the form of (s t , a t , s t+1 , r t ), where the four items of data represent the real-time state, the action taken, the future state, and the reward value obtained at time point t in sequence;
[0070] S3: Before reaching the termination condition, with the goal of minimizing the energy E consumed by the electric vehicle and the shortest time T experienced, and on the premise of successfully completing the transportation task, iteratively update the weight parameters of the neural network in the population, so as to calculate an optimal path for a specific transportation task.
[0071] Further, each generation in the iterative process of step S3 specifically includes:
[0072] S31: When the scales of the two replay buffers are large enough, refer to the optimal individual fitness value f p returned by the previous generation, and randomly select a set of experience data from each of them at a specific ratio to update the neural network parameters in each group in RL by the gradient descent method;
[0073] S32: Evolve the two populations to simulate the problem model; starting from the initial state s 0 , in each individual, continuously select the action with the highest value in the simulation according to the judgment of the value neural network Q to update the state, and return the total cumulative reward obtained throughout the process as the fitness value of the individual; the tuples of information such as the sealed current state s t , the action taken a t , the future state s t+1 and the reward value r t generated during this process are stored in R1 and R2 as experience data, and which replay buffer to finally enter is determined by the temporal difference value TD corresponding to the tuple. The calculation formula of this value is:
[0074] TD(s t ,a t ,s t+1 ,r t ) = r t +γQ′(s t+1 ,a′ t+1 |θ Q′ ) - Q(s t,t |θ Q )
[0075] The TD value is highly related to the loss function L for updating Q, and it can ensure that the empirical tuples with smaller TD values are more likely to achieve good results when training L. The present invention classifies the empirical tuples obtained in each evolution according to the TD value. The quarter with smaller TD values is stored in R1, and the other three - quarters are saved in R2;
[0076] S33: Store the elites, that is, sort the individuals in the population P according to the fitness value from excellent to poor, and select the best part as elite individuals for subsequent operations. The present invention sets the proportion of the elite part to 10%;
[0077] S34: Perform crossover operations on the individuals in the population P according to a specific crossover strategy to generate new individuals and update the population on a large scale;
[0078] S35: Perform mutation operations on the individuals in the population P according to a specific mutation strategy to generate new individuals for a more comprehensive exploration of the optimal solution and prevent falling into local optima;
[0079] S36: Sort the individuals in the current P according to the fitness value, and then replace the worst - performing part of the same amount of individuals in P with all the individuals in RL, so as to inject the training results of the current generation of RL into the evolutionary algorithm part. With the iterative mechanism of the "survival of the fittest" evolutionary population training method, the RL individuals that enter the training in P will be eliminated if they are not excellent enough, and will not have a negative impact on the training results of P. The same as training RL individuals with the empirical tuples generated by evolving the population P, this can enable the two populations to communicate with each other and promote the training effect.
[0080] Further, step S31 specifically includes:
[0081] Calculate the sum of the number of empirical tuples contained in the current two replay buffers. If it is greater than the specific parameter l set by the present invention to be 500 start , then perform the following steps:
[0082] Fetch the empirical tuples: According to the optimal fitness value f returned by the previous generation pCalculate the ratio of experience tuples selected from R1 and R2. In the present invention, the capacities of R1 and R2 are set to 2000 and 6000 respectively. The process of calculating the ratio is: if f p If it is small, it means that the training results are not good enough and more experience data with small time series difference values need to be selected. Therefore, the proportion of experience tuples selected in R1 is p rb It should be a larger value that does not exceed the maximum value of 0.75; otherwise, when f p When it is larger, the algorithm will explore more instead of sticking to the empirical tuples with small time difference values. rb is set to 0.25. In short, p rb is an adaptive update during the training process, from the upper limit of 0.75 to f p The parameter increases and decreases linearly to the lower limit of 0.25, based on which the agent selects T experience tuples from the two replay buffers. The parameter T can be set according to the problem scale, and the default value is 32;
[0083] Update the value neural network Q in the reinforcement learning group: Use these T experience tuples to update each individual in RL, that is, each group of two-layer value network frameworks. The present invention updates the weight parameters of the value neural network Q by minimizing the loss function L related to the expected value of the value neural network Q and the target value neural network Q'. The loss function formula is:
[0084] L=∑ t (y t -Q(s t ,a t |θ Q )) 2 / T
[0085] where y t =r t +γQ′(s t+1 ,a′ t+1 |θ Q′ ), γ is the discount factor for the past rewards of the current state, and the value can be set to 0.9 to 0.99, a′ t+1 is the state s judged by q′ t+1 The action that should be chosen next is θ Q Refers to all parameters in the neural network Q, the current state s t , take action a t 、Future State t+1 and obtain the reward value r t All of them are information sealed in tuples, and there are T groups in total;
[0086] Update the target value neural network Q' in the reinforcement learning population: Update the weight parameters of the target value neural network Q' through the updated value neural network Q and the soft update hyperparameter τ in this iteration. The update method is as follows:
[0087] q′ = q * τ + q′ * (1 - τ)
[0088] where q and q′ are the corresponding parameters in any Q and Q' respectively. The significance of this soft update mode is to prevent the neural network parameters from changing too much, so as to avoid missing the global optimal solution due to frequent updates. τ is set to 0.01 in the present invention.
[0089] Further, step S34 specifically includes:
[0090] Select parent individuals in P: Calculate the sum p of the fitness values of the individuals in P, then randomly generate a real number r from (0, p), and then initialize the parameter a = 0 and add the fitness values of each individual in turn until a plus the fitness value of a certain individual is greater than or equal to r for the first time, then select this individual as one of the parents for crossover. In this way, the higher the individual fitness value, that is, the more excellent, the higher the probability of being selected as a parent.
[0091] Generate offspring individuals by crossover parameters: After selecting two parents, perform individual parameter exchange, that is, exchange the parameters of the value neural network Q. The exchange range is a subset of parameters randomly selected from the same neural network structure of the two, so as to generate two offspring neural networks, and select the better one to be retained in the temporary individual set I.
[0092] When the number of individuals in I reaches the hyperparameter k, stop selecting parents, randomly select k non-elite individuals from the original population, replace them with the individuals in I, and then empty I. In the present invention, k = 3 is set.
[0093] Further, step S35 specifically includes:
[0094] Judgment of mutation: Determine the proportion of the parameters for mutating the value neural network by the relevant hyperparameter mut frac In the present invention, it is set to 0.1.
[0095] Select mutant individuals: For each non-elite individual, determine whether to perform a mutation operation by randomly generating a number. In the present invention, the mutation probability is set to 0.3.
[0096] Execute the mutation operator: For the determined mutant individuals, randomly select mut from their weight matrices fracFor some parameters, for each selected parameter, a randomly generated number is used to determine its mutation type. The setting of the present invention is as follows: there is a 0.9 probability of adding a small perturbation to the parameter, a 0.05 probability of adding a large perturbation, and a 0.05 probability of regenerating the parameter within the corresponding range. It is not difficult to see that the degrees of change of the above three mutation methods for the parameter are arranged in ascending order.
[0097] Further, the termination condition in step S3 is specifically:
[0098] Algorithm termination determination: If, after the end of the current iteration, the number of interactions between the agent representing the electric vehicle and the electric vehicle path planning problem model reaches the set upper limit (one interaction is an action of generating an action according to the current state) or the current training result is already excellent enough, the algorithm terminates and outputs the individual with the optimal fitness value in the two populations trained as the path for the finally optimized electric vehicle to complete the transportation task; if the algorithm does not reach the termination condition, that is, the number of interactions between the agent and the problem model does not reach the upper limit and the current training result is not excellent enough, the algorithm jumps to step S31 to continue the iteration.
[0099] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the described embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for electric vehicle path planning based on evolutionary reinforcement learning in a dual-network fusion scenario, characterized in that: include: Determine the type of mission the electric vehicle will perform, whether it is delivery or pickup; Obtain basic data of the vehicle dispatch center, transportation task points and charging stations in the current task; Initialize the reinforcement learning population RL and the evolution population P, where the number of individuals in RL is approximately 1 / 10 to 1 / 5 of the number of individuals in P. The specific value is determined by the scale of the problem model to be solved. Each individual in RL is a two-layer value network framework, including a value neural network Q and a target value neural network Q'. The individuals in P only have a single value neural network Q, and each individual represents a path planning; Define two replay buffers R1 and R2 to store (s t ,a t ,s t+1 ,r t ) of the experience tuple, where the four data items represent the real-time state at time point t, the action taken, the future state, and the reward value obtained; With the optimization goal of reducing the energy consumed by electric vehicles and shortening the time to complete the transportation task, the weight parameters of the neural network in the population are iteratively updated on the premise of successfully completing the transportation task, so as to obtain the optimal path for the electric vehicle to perform a transportation task.
2. The electric vehicle path planning method according to claim 1, characterized in that: The vehicle dispatch center controls the driving trajectory of electric vehicles. Each time a transport task is started, the vehicle starts from here and returns to here after completing the transport task. During the transport task, the vehicle may also return to here temporarily to replenish or unload goods. Transport task point: represents a transport task to be completed, such as picking up or delivering goods, and corresponds to a numerical value, which represents the quantity of goods; Charging station: The duration of a single charge is directly proportional to the amount of power consumed by the vehicle when charging begins.
3. The electric vehicle path planning method according to claim 2, characterized in that: In the electric vehicle path planning, the following conditions are set: 0≤u0≤C (9) The optimization goal (1) is to minimize the driving energy consumption and the total time. λ1 and λ2 are two balancing weights to ensure that two numbers of different magnitudes can be added together. Constraint (2) ensures that each transportation task point can only be visited once. Constraint (3) stipulates that each charging station can be visited at most once to avoid frequent charging in a short period of time. Constraint (4) represents the conservation of vehicle flow. Constraints (5) and (6) define the cargo volume balance relationship of consecutive visits to two points i and j, where M is a sufficiently large positive number. Constraint (7) is a time window constraint, e j and l j They represent the earliest and latest required visit times of transport task point j respectively. They can cooperate with constraints (5) and (6) to prevent the occurrence of sub-loops. Constraints (8) and (10) respectively ensure that the amount of cargo carried by the vehicle and the amount of battery power of the vehicle can be updated normally after completing any transport task. Constraints (9) and (11) respectively ensure that the amount of cargo carried by the vehicle and the amount of battery power of the vehicle are within a limited range that conforms to realistic logic.
4. The electric vehicle path planning method according to claim 1, characterized in that: The specific training process of reinforcement learning population RL is: According to the optimal fitness value f returned by the previous generation p Calculate the ratio p of selecting experience tuples from the double-layer replay buffer rb , based on which T experience tuples are selected from the two replay buffers; The take-out form is (s t ,a t ,s t+1 ,r t ) to calculate the loss function L, which is used to update the weight parameters of the value neural network Q. The loss function formula is: L=∑ t (y t -Q(s t ,a t |θ Q )) 2 / T where y t =r t +γQ′(s t+1 ,a′ t+1 |θ Q′ ), γ is the discount factor for the past rewards of the current state, and the value can be set to 0.9 to 0.99, a′ t+1 is the state s judged by Q′ t+1 The action that should be chosen next is θ Q Refers to all parameters in the neural network Q; The weight parameters of the target value neural network Q' are updated through the value neural network Q and the soft update hyperparameter τ. The update formula is: q′=(1-τ)q′+τq Where q is a parameter in Q, q′ is a parameter in Q′ with the same meaning as q, and τ is set to 0.01 to prevent missing the global optimal situation due to excessive updates.
5. The electric vehicle path planning method according to claim 1, characterized in that: Each individual in the evolutionary population P represents only one value neural network Q, and each individual represents a path planning. Specifically, the evolutionary algorithm is based on population optimization and iteratively improves the quality of the population through selection, crossover, and mutation operators, including: Evolving population, simulating the problem-solving model; starting from the initial state s0, each individual continuously selects the action with the highest value to update the state according to the judgment of the value neural network Q in the simulation, and returns the accumulated reward obtained throughout the process as the fitness value of the individual; the current state s generated in this process is sealed t , take action a t 、Future State t+1 and obtain the reward value r t The tuples of information are stored as experience data in the double-layer replay buffer, and which replay buffer is finally entered is determined by the time difference value TD corresponding to the tuple, and the calculation formula of this value is: TD(s t ,a t ,s t+1 ,r t )=r t +γQ′(s t+1 ,a′ t+1 |θ Q′ )-Q(s t ,a t |θ Q ); Sort the individuals in the evolutionary population P from best to worst according to their fitness values, and select the best 10% as elite individuals for subsequent operations; The parent individuals of the crossover operator are selected using the selection operator. The higher the fitness value of the individual in the population, the higher the probability of being selected. Execute the crossover operator to generate new offspring individuals and replace some non-elite individuals; Execute mutation operators on non-elite individuals to change the weight parameter value of Q in the individuals to ensure that diversity exploration always exists and prevent falling into local optimality; Replace the worst individuals in the evolutionary algorithm with all individuals in the reinforcement learning population RL.
6. The electric vehicle path planning method according to claim 5, characterized in that: In each iteration, the large amount of experience generated by the evolving population P after population evolution will be used to train the reinforcement population RL, and the training results of RL will then be injected into P. With the iterative mechanism of "survival of the fittest" of the evolutionary population training method, if the RL individuals entering P for training are not good enough, they will be eliminated, which will not have a negative impact on the training results of P.
7. The electric vehicle path planning method according to claim 4, characterized in that: The ratio p of selected experience tuples rb , specifically: p rb It is a parameter that is adaptively updated during the training process. The maximum value is set to 0.75 and the minimum value is set to 0.
25. As the optimal fitness value f returned by the previous generation p increases and decreases linearly, according to the ratio p rb T experience tuples are selected from the two replay buffers, where T can be set according to the problem scale and is generally 32 by default.
8. The electric vehicle path planning method according to claim 1, characterized in that: The capacities of the two replay buffers R1 and R2 are set to 2000 and 6000 respectively.
9. The electric vehicle path planning method according to claim 5, characterized in that: The form is (s t ,a t ,s t+1 ,r t ) is stored in R1 and R2 as experience data, and which replay buffer it finally enters is determined by the timing difference value TD corresponding to the tuple. The quarter with the smaller TD value is stored in R1, and the other three quarters are stored in R2.
10. The electric vehicle path planning method according to claim 5, characterized in that: Execute the mutation operator on the non-elite individuals to change the weight parameter value of Q in the individuals, specifically: First, by the hyperparameter mut frac Determine the ratio of the number of parameters of the value neural network variation; Then, for each non-elite individual, a random number is generated to determine whether to perform a mutation operation on it; Finally, mut is randomly selected from the weight matrix of the individual that determines the mutation frac Part of the parameters, and generate a corresponding random number for each selected parameter to determine its mutation type.
Citation Information
Cited By
Vehicle path planning method and device based on reinforcement learning
CN120800422A
Vehicle path planning method and device based on reinforcement learning
CN120800422B