A Method and System for Urban Electric Vehicle Scheduling Based on Deep Reinforcement Learning
By modeling the vehicle path problem as a directed complete graph and using graph neural network encoding and decoder, combining soft constraints and hard constraint training methods, the problems of asymmetric and complex constraints in the existing technology are solved, and the rapid solution and generalization capabilities are improved.
Patent Information
- Application Number
- CN202210056967.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-01-18
AI Technical Summary
When solving vehicle path problems, the existing deep reinforcement learning optimization algorithm is difficult to effectively deal with asymmetric and complex constraints, and the solution time is long, so it cannot adapt to vehicle path problems in actual scenarios.
Using a method based on deep reinforcement learning, the path problem of electric vehicles with time window is modeled as a directed complete graph, the path structure is performed using graph neural network encoding and decoder, and the parameters are updated through the REINFORCE algorithm, combining a two-stage training method of soft constraints and hard constraints to solve asymmetric and complex constraints.
On the premise of obtaining better solution effects, the solution time is greatly reduced. The model has the ability to solve quickly and generalize, is suitable for asymmetric vehicle path problems, and can effectively deal with complex constraints.
Smart Images

Figure CN114418213B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle routing problems, and more specifically, to an urban electric vehicle scheduling method and system based on deep reinforcement learning. Background Art
[0002] The Vehicle Routing Problem (VRP) refers to a situation where a certain number of customers have different demands for goods. A distribution center supplies goods to the customers, and a fleet of vehicles is responsible for delivering the goods. The goal is to organize an appropriate driving route to meet the customers' demands and achieve objectives such as the shortest distance, the minimum cost, and the least time consumption under certain constraints. The vehicle routing problem is a classic combinatorial optimization problem and belongs to the NP-hard problem. Due to its wide applicability and great economic value, it has been widely studied by scholars at home and abroad. Practical problems of the vehicle routing problem include distribution center delivery, buses, industrial waste collection, etc.
[0003] Based on the basic vehicle routing problem, different types of vehicle routing problems have emerged according to different problem settings. In recent years, new energy electric vehicles have been widely used. Compared with traditional vehicles, new energy electric vehicles use renewable and clean energy, have great advantages in environmental protection, and their market share has been increasing year by year. Under the pressure of energy and environmental protection, new energy vehicles will undoubtedly become the future development direction of vehicles. Therefore, a large number of studies on electric vehicle routing problems have also emerged. The Electric Vehicle Routing Problem with Time Windows (EVRPTW) adds driving range constraints and time window constraints to the basic vehicle routing problem. Specifically, given a certain number of customers, each customer has its own demand for goods and a time window during which it can be served. Given a fleet of electric vehicles, each electric vehicle has a limited loading capacity and a limited driving range. Starting from a warehouse, it provides goods to customers within the specified time window along the way, and can visit charging stations to recharge to increase the driving range. Finally, it returns to the warehouse before the specified latest time. It is required to organize an appropriate driving route for the fleet of electric vehicles to make the total route length the shortest under the premise of meeting the customers' demands and time, capacity, and driving range constraints.
[0004] At present, the methods for solving the vehicle routing problem can be mainly divided into exact algorithms, heuristic / meta-heuristic algorithms and deep reinforcement learning optimization algorithms. Exact algorithms are algorithms that can solve the global optimal solution, including branch and bound methods, dynamic programming methods, etc. Since the vehicle routing problem is an NP-hard problem, the computational complexity of the exact algorithm will increase exponentially with the scale of the problem, and it is difficult to expand to large-scale problems. Heuristic / meta-heuristic algorithms are algorithms based on intuitive or empirical construction, which can find a feasible solution within an acceptable computing time, but cannot guarantee the quality of the solution. Specifically, they include simulated annealing, taboo search, genetic algorithms, etc. Heuristic / meta-heuristic algorithms are generally iterative optimization algorithms. When the scale of the problem is large, a large number of iterative searches will still lead to a large amount of computation, and once the problem changes, it is necessary to search and solve again. In addition, the design of heuristic rules usually requires an in-depth understanding and research of the problem, which leads to difficulties in algorithm design.
[0005] Deep reinforcement learning optimization algorithm is a solution method that has emerged in recent years. Compared with traditional methods, deep reinforcement learning optimization algorithm has the advantages of fast solution speed and strong generalization ability. It can be divided into two categories: one is the constructive method, which adopts an end-to-end approach. Given a problem instance as input, it uses a trained deep neural network to directly output the solution to the problem. The parameters of the neural network are obtained through deep reinforcement learning training. Compared with the traditional iterative optimization algorithm, the constructive method directly outputs the solution to the problem without searching, and has the advantage of fast solution speed. Once the model is trained, it can solve all problem instances with the same distribution characteristics, and has a certain generalization ability. The traditional algorithm needs to search and solve each new problem instance from scratch, which is very time-consuming. The other type is the lifting method, which uses deep reinforcement learning to learn and select heuristic rules under an iterative search framework, and performs iterative search for solutions based on the learned rules. This type of method replaces manual design with a neural network model, thereby reducing the difficulty of algorithm design. Since it is still an iterative optimization algorithm in essence, although this type of method has a good optimization effect, its solution speed is far slower than the constructive end-to-end method.
[0006] In the existing research on the optimization algorithm of deep reinforcement learning for solving the vehicle routing problem, there are two deficiencies: one is that the problem is divorced from the real scenario. Currently, most research focuses on the symmetric vehicle routing problem, where the distance between nodes is the Euclidean distance calculated through coordinates and is symmetric. However, in the real vehicle routing problem, the distance between nodes cannot be simply the Euclidean distance and is almost impossible to be symmetric. Therefore, it is necessary to extend the deep reinforcement learning optimization algorithm to the asymmetric vehicle routing problem. The other is the lack of an effective constraint handling mechanism to solve the complex constraints in the vehicle routing problem. Currently, during the training process of the constructive deep reinforcement learning optimization algorithm, the constraints are usually handled by directly masking illegal actions. Although this hard constraint handling method can ensure the generation of feasible solutions, it affects the solution quality of the model to a certain extent.
[0007] In the prior art, a method for solving the vehicle routing problem of logistics transportation with soft time windows is disclosed. For the vehicle routing problem of logistics transportation with soft time windows based on real-time traffic information, a time window penalty mechanism is adopted to establish its mathematical model; an adaptive chaotic ant colony algorithm is used to solve the model, and the optimization ability of the algorithm is improved through the adaptive update of algorithm pheromones and the chaotic adaptive adjustment of algorithm parameters. This method takes a long time and cannot be well applied to actual cases. Summary of the Invention
[0008] The primary object of the present invention is to provide a method for scheduling urban electric vehicles based on deep reinforcement learning, which can significantly reduce the solution time while obtaining better solution effects.
[0009] A further object of the present invention is to provide a system for scheduling urban electric vehicles based on deep reinforcement learning.
[0010] To solve the above technical problems, the technical solution of the present invention is as follows:
[0011] A method for scheduling urban electric vehicles based on deep reinforcement learning, characterized by comprising the following steps:
[0012] S1: Model the electric vehicle routing problem with time windows into a directed complete graph, where the warehouse, charging stations, and customers are nodes in the graph, and any two nodes are connected by edges. Normalize the demand, distance, and time data respectively.
[0013] S2: Use an encoder to encode the point information and edge information in the directed complete graph respectively to obtain corresponding feature representations.
[0014] S3: Use a decoder for decoding. In each step of decoding, based on the feature representations of points and edges obtained in step S2, as well as the current vehicle state information and historical path information, gradually construct a path in an autoregressive manner to obtain the solution to the problem.
[0015] S4: Calculate the total return according to the solution of the problem, and use the REINFORCE algorithm to update the parameters of the encoder and decoder;
[0016] S5: Use the trained encoder and decoder to solve the electric vehicle routing problem with time windows.
[0017] Further, in step S1, the node information is v i =(d i , e i , l i , t i ), where d i represents the customer demand, e i represents the earliest service time, l i represents the latest service time, t i represents the node type, and there are:
[0018]
[0019] Among them, V d , V s , V c represent the warehouse node set, the charging station node set and the customer node set respectively.
[0020] Further, in step S1, the edge information is e ij =(dis ij , time ij , a ij ), where dis ij represents the distance, time ij represents the time, a ij represents the nearest neighbor, and there are:
[0021]
[0022] Further, step S2 specifically includes the following steps:
[0023] S2.1: Use two embedding layers to map the node information v i and the edge information e ij into high-dimensional feature vectors to obtain the first-layer input of the graph neural network and
[0024]
[0025]
[0026] In the formula, W V , bV , W E , b E are all trainable parameters;
[0027] S2.2: Use a graph neural network to and to obtain the final feature vector representation through N layers of the graph neural network. In each layer of the graph neural network, each node and edge will aggregate the information of adjacent nodes and edges to update itself. The update method of the node feature representation is:
[0028]
[0029]
[0030]
[0031] The update method of the edge feature representation is:
[0032]
[0033]
[0034]
[0035] where MHA is the multi-head attention sub-layer, FF is the fully connected sub-layer, BN is the batch normalization sub-layer, ; represents the concatenation operation, and σ is the activation function Relu, are all trainable parameters. The output of the last layer of the graph neural network is the feature vector representation obtained by encoding all node information and edge information through the encoder.
[0036] Furthermore, the specific steps of step S3 include the following steps:
[0037] S3.1: According to the feature vector representations of the nodes and edges obtained by encoding through the encoder, as well as the vehicle state information and historical path information at the current decoding step, first use the glimpse mechanism to calculate a query vector. Specifically, assume that the vehicle is currently at node i, then calculate the query vector:
[0038] c t = W C C t + b C
[0039]
[0040] h t = GRU t (h i )
[0041] where MHA represents the multi-head attention layer, and W C , b C are all trainable parameters, and C t =(T t , D t , B t ) represents the current vehicle state information, where T t is the current time, D t is the remaining capacity, and B t is the remaining driving range, and h j and represent the feature vector representations of the corresponding points and edges;
[0042] S3.2: Adopt the attention mechanism to calculate the weight of each node, that is, the probability distribution p t , according to the query vector q t and the hidden vectors of the points and edges adjacent to node i:
[0043]
[0044]
[0045] p t =softmax(u t )
[0046] where W Q , W K are trainable parameters, C is a constant, and d h is the dimension of Q t , indicates that node j can be selected at the
[0047] t-step decoding, otherwise it means it cannot be selected. In the soft constraint processing method,
[0048] when one of the following situations occurs, there is
[0049] · i = j;
[0050] · Node i is a warehouse or a charging station and node j is a charging station;
[0051] · Node j is a customer and has been visited;
[0052] In the hard constraint processing method, when one of the following situations occurs, there is
[0053] · i = j;
[0054] · Node i is a warehouse or a charging station and node j is a charging station;
[0055] · Node j is a customer and has been visited;
[0056] · The remaining capacity of the vehicle is less than the demand of node j, i.e., D t <d j ;
[0057] · The arrival time at node j will be later than the latest service time of node j, i.e., T t +time ij >l j ;
[0058] · The remaining driving mileage does not support reaching node j, i.e., B t <dis ij ;
[0059] · The remaining driving mileage after reaching node j does not support reaching any warehouse or charging station;
[0060] S3.3: According to the probability distribution p t , select a node j for visit, that is, execute an action, add this node j to the historical path π, and update the vehicle status information. The current time is updated to:
[0061]
[0062] where s is the service time and c is the charging time;
[0063] The current remaining capacity is updated to:
[0064]
[0065] where D max is the maximum loading capacity of the vehicle;
[0066] The current remaining driving mileage is updated to:
[0067]
[0068] where B max is the maximum driving mileage of the vehicle;
[0069] S3.4: Repeat steps S3.1 - S3.3 until the vehicle has served all customer nodes and returned to the warehouse. The sequence of nodes selected during this process is the solution to the problem.
[0070] Furthermore, in step S3.3, when selecting a node j for visit, there are two selection methods. One is the greedy strategy, which selects the node with the highest probability at each step; the other is the random strategy, that is, the probability of a node being selected is the probability output by the decoder.
[0071] Further, in step S4, the total reward is calculated according to the solution of the problem, specifically as follows:
[0072]
[0073] In the formula, π = {i0, i1, …, i T} represents the node sequence, i.e., the solution of the problem, and α, β, and γ are all constant coefficients.
[0074] Further, in step S4, the REINFORCE algorithm is used to update the parameters of the encoder and decoder, specifically as follows:
[0075]
[0076]
[0077]
[0078] Where s represents the problem instance, and b(s) is the total reward of the solution obtained by the greedy decoding method of the current policy network. The purpose of introducing it is to reduce the variance of the policy gradient and make the training stable. Adam is the Adam optimizer.
[0079] Further, in step S5, the trained encoder and decoder are as follows:
[0080] Randomly generate a simulation example set, and divide all problem instances into a training set, a validation set, and a test set. Use the training set to train the encoder and decoder multiple times. In the previous stage of training, the soft constraint processing method is adopted, and in the latter stage of training, the hard constraint processing method is adopted. After each batch of training is completed, a solution evaluation is performed on the validation set, and the encoder and decoder with the best performance on the validation set are used to solve the electric vehicle routing problem with time windows.
[0081] An urban electric vehicle scheduling system based on deep reinforcement learning includes:
[0082] A graph modeling module that models the electric vehicle routing problem with time windows into a directed complete graph. The warehouse, charging station, and customers are nodes in the graph, and any two nodes are connected by an edge. The demand, distance, and time data are respectively normalized.
[0083] An encoding module that uses an encoder to encode the point information and edge information in the directed complete graph respectively to obtain corresponding feature representations.
[0084] A decoding module, which uses a decoder for decoding. In each step of decoding, according to the feature representations of the points and edges obtained in the encoding module, as well as the current vehicle state information and historical path information, it gradually constructs a path in an autoregressive manner to obtain the solution to the problem;
[0085] A parameter update module, which calculates the total reward according to the solution to the problem and uses the REINFORCE algorithm to update the parameters of the encoder and decoder;
[0086] A solution module, which uses the trained encoder and decoder to solve the electric vehicle routing problem with time windows.
[0087] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0088] 1. The present invention designs a deep reinforcement learning optimization algorithm for solving the asymmetric electric vehicle routing problem with time windows. Compared with traditional methods, it can significantly reduce the solution time on the premise of obtaining comparable or better solution effects, and the trained model can solve problem instances with the same distribution characteristics, having the advantages of fast solution speed and strong generalization ability.
[0089] 2. The graph neural network designed in the present invention for capturing and extracting edge information can effectively solve the asymmetric vehicle routing problem, making the algorithm have wide applicability and practical significance.
[0090] 3. The two-stage training method of soft constraint + hard constraint proposed in the present invention enables the model to better handle complex constraints and obtain better solution effects, and this method is also easy to be extended to other combinatorial optimization problems with complex constraints. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 It is a schematic flow chart of the method of the present invention.
[0092] Figure 2 It is a schematic diagram of the model structure of the present invention.
[0093] Figure 3 It is a schematic diagram of the system module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0094] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0095] For better illustration of this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the dimensions of the actual product;
[0096] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0097] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0098] Embodiment 1
[0099] This embodiment provides an urban electric vehicle scheduling method based on deep reinforcement learning. As Figures 1 to 2 shown, it includes the following steps:
[0100] S1: Model the electric vehicle routing problem with time windows into a directed complete graph, where the warehouse, charging stations, and customers are nodes in the graph, and any two nodes are connected by an edge. Normalize the demand, distance, and time data respectively.
[0101] S2: Use an encoder to encode the point information and edge information in the directed complete graph respectively to obtain corresponding feature representations.
[0102] S3: Use a decoder for decoding. In each step of decoding, according to the feature representations of the points and edges obtained in step S2, as well as the current vehicle state information and historical path information, gradually construct the path in an autoregressive manner to obtain the solution to the problem.
[0103] S4: Calculate the total reward according to the solution to the problem, and use the REINFORCE algorithm to update the parameters of the encoder and decoder.
[0104] S5: Use the trained encoder and decoder to solve the electric vehicle routing problem with time windows.
[0105] This embodiment is a method for solving the electric vehicle problem with time windows based on deep reinforcement learning, which is an end-to-end method. Given a problem instance as input, the trained deep neural network can directly output the solution to the problem. Once the model is trained, it can solve all problem instances with the same distribution characteristics. Therefore, it has the advantages of fast solving speed and strong generalization ability. First, organize and obtain the point information and edge information of the problem instance and perform data preprocessing. Then, input the point information and edge information of the problem instance into the encoder for encoding to obtain corresponding feature vector representations. Next, use the decoder to perform sequence decoding on the feature vector representations of the points and edges, as well as the vehicle state information and historical path information to obtain the node sequence, that is, the solution to the problem. Finally, calculate the total reward according to the solution and update the model parameters. Repeat the above steps several times to obtain a trained model that can be used to solve the electric vehicle routing problem with time windows.
[0106] The node information in step S1 is v i =(d i ,e i ,l i ,t i ), where d iRepresents customer requirements, e i Represents the earliest service time, l i Represents the latest service time, t i Represents the node type, and there are:
[0107]
[0108] Among them, V d , V s , V c respectively represent the set of warehouse nodes, the set of charging station nodes, and the set of customer nodes.
[0109] The edge information in step S1 is e ij =(dis ij , time ij , a ij ), where dis ij represents the distance, time ij represents the time, a ij represents the nearest neighbor, and there are:
[0110]
[0111] Then, according to the maximum loading capacity of the vehicle, the maximum driving range of the vehicle, and the earliest departure time and the latest return time of the vehicle, the requirements, distances, and times of all point information and edge information are normalized respectively.
[0112] Step S2 specifically includes the following steps:
[0113] S2.1: Use two embedding layers to map the node information v i and the edge information e ij into high-dimensional feature vectors to obtain the first-layer input of the graph neural network and
[0114]
[0115]
[0116] In the formula, W V , b V , W E , b E are all trainable parameters;
[0117] S2.2: Use the graph neural network to pass and through N layers of the graph neural network to obtain the final feature vector representation. In each layer of the graph neural network, each point and edge will aggregate the information of adjacent points and edges to update itself. The update method of the point feature representation is:
[0118]
[0119]
[0120]
[0121] The update method of the edge feature representation is as follows:
[0122]
[0123]
[0124]
[0125] Among them, MHA is the multi-head attention sub-layer, FF is the fully connected sub-layer, BN is the batch normalization sub-layer; the semicolon represents the concatenation operation, and σ is the activation function Relu. All are trainable parameters, and the output of the last graph neural network is the feature vector representation obtained by encoding all the point information and edge information through the encoder.
[0126] The specific steps of step S3 are as follows:
[0127] S3.1: According to the feature vector representations of the points and edges obtained by encoding with the encoder, as well as the vehicle state information and historical path information at the current decoding step, first use the glimpse mechanism to calculate a query vector. Specifically, assuming that the vehicle is currently at node i, the query vector is calculated as follows:
[0128] c t = W C C t + b C
[0129]
[0130] h t = DRU t (h i )
[0131] In the formula, MHA represents the multi-head attention layer, W C , b C are all trainable parameters, C t =(T t , D t , B t ) represents the current vehicle state information, T t is the current time, D t is the remaining capacity, B t is the remaining driving mileage, h j and The eigenvector representation corresponding to points and edges;
[0132] S3.2: Adopt the attention mechanism, and calculate the weight of each node, that is, the probability distribution p, according to the query vector q t and the hidden vectors of the points and edges adjacent to node i t :
[0133]
[0134]
[0135] p t = softmax(u t )
[0136] where W Q , W K are trainable parameters, C is a constant, d h is the dimension of Q t . indicates that node j can be selected at the t-step decoding, otherwise it means that it cannot be selected. The purpose of introducing the mask is to ensure the generation of feasible solutions. Here, two constraint handling methods, soft constraint and hard constraint, are designed. In the soft constraint handling method, when one of the following situations occurs
[0137] · i = j;
[0138] · Node i is a warehouse or a charging station and node j is a charging station;
[0139] · Node j is a customer and has been visited;
[0140] In the hard constraint handling method, when one of the following situations occurs
[0141] · i = j;
[0142] · Node i is a warehouse or a charging station and node j is a charging station;
[0143] · Node j is a customer and has been visited;
[0144] · The remaining capacity of the vehicle is less than the demand of node j, that is, D t <d j ;
[0145] · The time to reach node j will be later than the latest service time of node j, that is, T t + time ij > l j ;
[0146] · The remaining driving mileage does not support reaching node j, that is, Bt <dis ij ;
[0147] · The remaining driving range after reaching node j does not support reaching any warehouse or charging station;
[0148] S3.3: According to the probability distribution p t , select a node j to visit, that is, execute an action, add this node j to the historical path π, and update the vehicle status information. The current time is updated to:
[0149]
[0150] where s is the service time and c is the charging time;
[0151] The current remaining capacity is updated to:
[0152]
[0153] where D max is the maximum loading capacity of the vehicle;
[0154] The current remaining driving range is updated to:
[0155]
[0156] where B max is the maximum driving range of the vehicle;
[0157] S3.4: Repeat steps S3.1 - S3.3 until the vehicle has served all customer nodes and returned to the warehouse. The sequence of nodes selected during this process is the solution to the problem.
[0158] In step S3.3, when selecting a node j to visit, there are two selection methods. One is the greedy strategy, which selects the node with the highest probability at each step; the other is the random strategy, that is, the probability of a node being selected is the probability output by the decoder.
[0159] In step S4, the total reward is calculated according to the solution to the problem, specifically:
[0160]
[0161] In the formula, π = {i0, i1,..., i T} represents the node sequence, that is, the solution to the problem, and α, β, and γ are all constant coefficients.
[0162] In step S4, the REINFORCE algorithm is used to update the parameters of the encoder and decoder, specifically:
[0163]
[0164]
[0165]
[0166] Among them, s represents the problem instance, and b(s) is the total reward of the solution obtained by the greedy decoding method of the current policy network. The purpose of introducing it is to reduce the variance of the policy gradient and make the training stable. Adam is the Adam optimizer.
[0167] The trained encoder and decoder in step S5 are specifically as follows:
[0168] Randomly generate a simulation example set, and divide all problem instances into a training set, a validation set, and a test set. Use the training set to train the encoder and decoder multiple times. Among them, the soft constraint processing method is adopted in the previous stage of training, and the hard constraint processing method is adopted in the latter stage of training. After each batch of training is completed, a solution evaluation is performed on the validation set, and the encoder and decoder with the best performance on the validation set are used to solve the electric vehicle routing problem with time windows.
[0169] Embodiment 2
[0170] This embodiment provides a specific embodiment of Embodiment 1, specifically as follows:
[0171] Evaluate through a randomly generated simulation example set, and divide it into a training set, a validation set, and a test set. Among them, the training set has 32,000 examples, each example contains S = 2 charging station nodes and C = 20 customer nodes, the validation set has 1,000 examples, and each example also contains S = 2 charging station nodes and C = 20 customer nodes. The test set has three types of examples, each with 1,000. The three types of examples respectively contain S = 2 charging station nodes and C = 20 customer nodes (S2-C20), S = 5 charging station nodes and C = 50 customer nodes (S5-C50), and S = 10 charging station nodes and C = 100 customer nodes (S10-C100). Use the test set to test the trained model and record the experimental results. The model adopts two decoding methods, greedy and sample, during testing. The sample decoding method collects 1,280 paths for each example and selects the best result among them.
[0172] The present invention is measured using two evaluation indicators:
[0173] 1. Solution quality: It represents the total path length of the solution obtained for each example on average.
[0174] 2. Solution time: It represents the time used to solve each example on average.
[0175] Table 1 Experimental results of the solution quality of the present invention and other comparative methods on the test set (unit: m, the true result divided by 1e5)
[0176] Method S2-C20 S5-C50 S10-C100 OR-Tools 5.9124 16.0137 - SA 5.7714 11.6925 20.4695 RL (greedy) 6.5543 13.1467 23.1973 RL (sample) 6.1120 12.1550 21.5154 The present invention (greedy) 6.2472 12.6422 22.0075 The present invention (sample) 5.9028 11.6041 20.8789
[0177] Table 2 Experimental results of the solution time of the present invention and other comparative methods on the test set (unit: s)
[0178] Method S2-C20 S5-C50 S10-C100 OR-Tools 54.26 56.38 - SA 27.79 49.87 105.22 RL 0.82 1.44 2.17 The present invention 0.53 0.78 1.13
[0179] From the above experimental results, it can be seen that the present invention can achieve better solution effects while significantly reducing the solution time compared with other methods.
[0180] Example 3
[0181] This embodiment provides an urban electric vehicle scheduling system based on deep reinforcement learning, as Figure 3 shown, including:
[0182] A graph modeling module, which models the electric vehicle routing problem with time windows into a directed complete graph, where the warehouse, charging stations, and customers are nodes in the graph, and any two nodes are connected by edges, and normalizes the demand, distance, and time data respectively;
[0183] An encoding module, which uses an encoder to encode the point information and edge information in the directed complete graph respectively to obtain corresponding feature representations;
[0184] A decoding module, which uses a decoder for decoding, and in each step of decoding, constructs a path step by step in an autoregressive manner according to the feature representations of the points and edges obtained in the encoding module, the current vehicle state information, and the historical path information, to obtain the solution to the problem;
[0185] A parameter update module, which calculates the total reward according to the solution to the problem, and uses the REINFORCE algorithm to update the parameters of the encoder and decoder;
[0186] A solving module, which uses the trained encoder and decoder to solve the electric vehicle routing problem with time windows.
[0187] The same or similar reference numerals correspond to the same or similar components;
[0188] The terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0189] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. A method for scheduling urban electric vehicles based on deep reinforcement learning, characterized in that, It includes the following steps: S1: Model the electric vehicle routing problem with time windows into a directed complete graph, where the warehouse, charging stations, and customers are nodes in the graph, and any two nodes are connected by edges. Normalize the demand, distance, and time data respectively. S2: Use an encoder to encode the point information and edge information in the directed complete graph respectively to obtain corresponding feature representations. S3: Use a decoder for decoding. In each decoding step, based on the feature representations of the points and edges obtained in step S2, as well as the current vehicle state information and historical path information, gradually construct a path in an autoregressive manner to obtain the solution to the problem. S4: Calculate the total reward based on the solution to the problem, and use the REINFORCE algorithm to update the parameters of the encoder and decoder. S5: Use the trained encoder and decoder to solve the electric vehicle routing problem with time windows. The node information in step S1 is v i =(d i , e i , l i , t i ), where d i represents customer requirements, e i represents the earliest service time, l i represents the latest service time, t i represents the node type, and there is: Among them, V d , V s , V c respectively represent the set of warehouse nodes, the set of charging station nodes, and the set of customer nodes; In step S1, the edge information is e ij =(dis ij , time ij , a ij ), where dis ij represents distance, time ij represents time, and a ij represents the nearest neighbor, and there is: The specific steps of step S2 include the following steps: S2.1: Use two embedding layers to map the node information v i and the edge information e ij into high-dimensional feature vectors to obtain the input of the first layer of the graph neural network and where W V , b V , W E , b E are all trainable parameters; S2.2: Using a graph neural network, and After passing through N layers of the graph neural network, the final feature vector representation is obtained. In each layer of the graph neural network, each node and edge will aggregate the information of adjacent nodes and edges to update itself. The update method of the node feature representation is as follows: The update method of the edge feature representation is: Among them, MHA is the multi-head attention sub-layer, FF is the fully connected sub-layer, BN is the batch normalization sub-layer, [;] represents the concatenation operation, σ is the activation function Relu, All are trainable parameters, and the output of the last layer of the graph neural network is the feature vector representation obtained by encoding all point information and edge information through the encoder; The specific steps of step S3 include the following steps: S3.1: Based on the feature vector representations of the points and edges encoded by the encoder, as well as the current vehicle state information and historical path information at the current decoding step, first use the glimpse mechanism to calculate a query vector. Specifically, assume that the vehicle is currently at node i, then calculate the query vector: c t = W C C t + b C h t = GRU t (h i ) In the formula, MHA represents the multi-head attention layer, W C , b C are both trainable parameters, C t =(T t , D t , B t ) represents the current vehicle state information, T t is the current time, D t is the remaining capacity, B t is the remaining driving range, h j and represent the feature vector representations of the corresponding points and edges; S3.2: Adopt the attention mechanism and calculate the weight of each node, that is, the probability distribution p, according to the query vector q t and the hidden vectors of the adjacent points and edges of node i t : p t = softmax(u t ) Among which W Q , W K is a trainable parameter, C is a constant, d h is the dimension of Q t . indicates that node j can be selected at the t-step decoding, otherwise it indicates that it cannot be selected. In the soft constraint handling method, when one of the following situations occurs, there is ●i = j; ● Node i is a warehouse or a charging station and node j is a charging station; ● Node j is a customer and has been visited; In the hard constraint handling method, there is when one of the following situations is encountered ●i = j; ● Node i is a warehouse or a charging station and node j is a charging station; ● Node j is a customer and has been visited; ● The remaining capacity of the vehicle is less than the demand of node j, i.e., D t <d j ; ● The time to reach node j is later than the latest service time of node j, i.e., T t +time ij >l j ; ● The remaining driving range does not support reaching node j, i.e., B t <dis ij ; ● The remaining driving mileage after reaching node j does not support reaching any warehouse or charging station; S3.3: Select a node j for access, i.e., execute an action, according to the probability distribution p t , add this node j to the historical path π, and update the vehicle state information. The current time is updated to: where s is the service time and c is the charging time; The current remaining capacity is updated to: Among them, D max is the maximum loading capacity of the vehicle; The current remaining driving mileage is updated to: where B max is the maximum driving range of the vehicle; S3.4: Repeat steps S3.1 - S3.3 until the vehicle has served all customer nodes and returns to the warehouse. The sequence of nodes selected during the repetition of steps S3.1 - S3.3 is the solution to the problem.
2. The method for scheduling urban electric vehicles based on deep reinforcement learning according to claim 1, characterized in that, In step S3.3, when selecting a node j for access, there are two selection methods. One is the greedy strategy, which selects the node with the highest probability at each step; the other is the random strategy, that is, the probability of a node being selected is the probability output by the decoder.
3. The method for scheduling urban electric vehicles based on deep reinforcement learning according to claim 1, wherein, In step S4, the total reward is calculated based on the solution to the problem, specifically: where π = {i0, i1, …, i T} represents the node sequence, i.e., the solution to the problem, and α, β, and γ are all constant coefficients.
4. The urban electric vehicle scheduling method based on deep reinforcement learning according to claim 1, characterized in that, In step S4, the REINFORCE algorithm is used to update the parameters of the encoder and decoder, specifically: where s represents the problem instance, b(s) is the total reward of the solution obtained by the greedy decoding method of the current policy network. The purpose of introducing it is to reduce the variance of the policy gradient and make the training stable. Adam is the Adam optimizer.
5. The method for scheduling urban electric vehicles based on deep reinforcement learning according to claim 1, wherein In step S5, the trained encoder and decoder are specifically: Randomly generate a set of simulation examples, and divide all problem instances into a training set, a validation set, and a test set. Use the training set to train the encoder and decoder multiple times. In the previous stage of training, a soft constraint processing method is adopted, and in the latter stage of training, a hard constraint processing method is adopted. After each batch of training is completed, a solution evaluation is performed on the validation set, and the encoder and decoder with the best performance on the validation set are used to solve the electric vehicle routing problem with time windows.
6. An urban electric vehicle scheduling system based on deep reinforcement learning, characterized in that, The system applies the urban electric vehicle scheduling method based on deep reinforcement learning according to any one of claims 1 to 5, including: A graph modeling module, which models the electric vehicle routing problem with time windows into a directed complete graph. The warehouse, charging stations, and customers are nodes in the graph, and any two nodes are connected by edges. Normalize the demand, distance, and time data respectively. An encoding module, which uses an encoder to encode the point information and edge information in the directed complete graph respectively to obtain corresponding feature representations. A decoding module, which uses a decoder for decoding. In each step of decoding, according to the feature representations of the points and edges obtained in the encoding module, as well as the current vehicle state information and historical path information, construct the path step by step in an autoregressive manner to obtain the solution to the problem. A parameter update module, which calculates the total reward according to the solution to the problem and uses the REINFORCE algorithm to update the parameters of the encoder and decoder. A solution module, which uses the trained encoder and decoder to solve the electric vehicle routing problem with time windows.
Citation Information
Patent Citations
Intelligent logistics distribution and delivery based on discrete particle swarm optimization algorithm
CN102117441A
Deep learning-based vehicle path optimization method and system
CN106548645A