VRPSDPTW problem solving method based on deep reinforcement learning
Through a method based on deep reinforcement learning, a mathematical model for the path problem of simultaneous delivery and pick-up vehicles is established and solved, and the problem of difficult to quickly and high-quality planning of vehicle paths in the existing technology is solved, and the effect of reducing energy consumption and distribution costs is achieved, and the efficiency and quality of logistics distribution is improved.
Patent Information
- Application Number
- CN202510060602.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-06
AI Technical Summary
The problem of the existing technology is difficult to quickly and with high quality in planning the path of the simultaneous delivery and pick-up vehicles, resulting in high energy consumption and high distribution costs, making it difficult to achieve cost reduction and efficiency improvement of logistics companies and improve service levels.
A method based on deep reinforcement learning is adopted to establish a mathematical model of the vehicle path problem that considers the time window and the simultaneous delivery of goods. The neural network model based on attention mechanism and the Reinforce training strategy with rollback baseline is solved to generate a high-quality delivery solution.
Through deep reinforcement learning algorithms, it can quickly solve the problem of the path of the simultaneous delivery and pick-up vehicle, reduce energy consumption and distribution costs, improve the efficiency and quality of logistics distribution, and achieve cost reduction and efficiency improvement and service level improvement of logistics companies.
Smart Images

Figure CN119941104A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle routing problems, and in particular relates to a method for solving a vehicle routing problem with time windows and simultaneous delivery and pickup (VRPSDPTW) based on deep reinforcement learning. Background Art
[0002] The integrated delivery and pickup logistics distribution model is becoming the direction of development for logistics enterprises, which can bring about the increase of economic, social and environmental benefits. Optimizing the distribution link can help the logistics industry reduce costs and increase efficiency, and also help enterprises enhance their market competitiveness.
[0003] The vehicle routing problem (VRP) was first proposed by Dantzig and Ramser in 1959 and has become a research hotspot in the fields of operations research and combinatorial optimization. The study of logistics scheduling and distribution problems is essentially a further study of the optimization path planning solutions and scheduling problems of more complex NP-Hard problems.
[0004] With the advancement of computer performance, the introduced deep reinforcement learning method has shown good performance in solving the vehicle routing problem. The deep reinforcement learning method based on the attention mechanism can quickly solve the traveling salesman problem (TSP), vehicle routing problem and other path planning problems, and can ensure a high quality of solution. The training is driven by data, and the algorithm design greatly reduces the human factor. Once the reinforcement learning training is completed, the learned strategy does not require additional training and search, and only a parameterized forward propagation is required to obtain a high-quality solution. Summary of the invention
[0005] The technical problem to be solved by the present invention is to quickly and high-quality plan the vehicle distribution routes in the simultaneous delivery and pickup vehicle routing problem, so as to reduce energy consumption and distribution costs, thereby realizing cost reduction and efficiency improvement and improving service levels for logistics companies. The present invention provides an optimization method for the simultaneous delivery and pickup vehicle routing problem based on deep reinforcement learning, establishes a mathematical model for the simultaneous delivery and pickup vehicle routing problem, and solves the problem model based on a deep reinforcement learning algorithm.
[0006] The present invention solves the above technical problems through the following technical solutions, and the present invention comprises the following steps:
[0007] 1) Establish the objective function, construct a mathematical model for the routing problem that takes into account the time window and simultaneous delivery and pickup, and determine the model constraints;
[0008] 2) building a reinforcement learning environment based on the mathematical model;
[0009] 3) Building a neural network model based on the attention mechanism according to the reinforcement learning environment;
[0010] 4) Use the data set and deep reinforcement learning algorithm to train the constructed neural network model to obtain an application model that meets the requirements, and use it to solve logistics distribution scenarios and output distribution plans.
[0011] In the above scheme, the objective function constructed is specifically vehicle transportation cost, delivery time window cost and customer satisfaction cost, and its specific form is:
[0012] minZ(x)=p1z1(x)+p2z2(x)+p3z3(x)
[0013] Among them, the weighted coefficient relationship is
[0014]
[0015] The components of the objective function are:
[0016] Vehicle transportation costs where c start and c0 are the vehicle startup cost and the vehicle unit distance travel cost respectively;
[0017] Delivery time window cost where c wait and c late are the vehicle unit waiting time cost and unit delay time penalty respectively;
[0018] Customer Satisfaction Cost where c ust and c sat They are unit customer satisfaction cost and unit customer dissatisfaction cost respectively.
[0019] In the above scheme, the mathematical model of the problem is specifically:
[0020] The distribution network consists of a directed graph G =<C,E> The set of Nc customer nodes is represented by C = {0, 1, ..., Nc}, where node 0 represents the distribution warehouse. The feasible routes connecting vehicles in the network are represented by E = {(i, j) | i, j ∈ C, i ≠ j}. The set of delivery vehicles is V = {1, ..., Nv}, where Nv represents the maximum number of available vehicles.
[0021] The model follows the following assumptions: there are enough vehicles with the same capacity available at the warehouse; a customer node can only be served by one vehicle; a vehicle should be returned to the warehouse if it cannot meet customer demand or its capacity is exhausted; the load during vehicle transportation cannot exceed the vehicle capacity.
[0022] The three binary decision variables designed in the mathematical model are:
[0023]
[0024]
[0025]
[0026] In the above scheme, the reinforcement learning quadruplets are
[0027] state t = <X t ,V t >(s t ∈S), which contains two parts of elements, namely X t , used to represent node information and existing partial solutions (paths), V t , which is used to represent the vehicle status at time t.
[0028] Action Space, which is action a t A collection of t ~π(a|s t ), after each step is executed, the visited nodes will be blocked to avoid repeated visits.
[0029] Reward is based on the optimized objective function. The purpose is to minimize the objective function, so the reward function is set to R(x) = -Z(x). The reward setting will make the action selection tend to reduce the value of the objective function, so that the strategy of producing high-quality solutions can be trained;
[0030] Policy uses the following chain rule to express the probability relationship in building a complete path:
[0031]
[0032] The strategy π has trainable parameters θ. At each step of building the solution path, the strategy outputs the probability distribution of node selection. Nodes are selected one by one until the complete path is built. For example, a = {a0, a1, ..., a T}, where T is the maximum number of steps. The strategy π is parameterized as a neural network π with an encoder-decoder structure θ , which is used to solve the routing problem of simultaneous delivery and pickup vehicles;
[0033] In the above scheme, the deep neural network model is a codec structure, and the specific design is described as follows:
[0034] The encoder uses the Transformer module to obtain node embeddings, while the decoder uses the node embeddings and access information masks to generate context embeddings, and then uses a multi-head attention mechanism to obtain the probability distribution of node selection. To select actions, an attention aggregation module is designed to obtain excellent context embeddings to capture dynamic state transitions.
[0035] In the above scheme, the specific structure of the designed attention aggregation module is:
[0036] Modules use node embedding node and access information mask As input, Set to 1 if the node has been visited and 0 otherwise.
[0037] The Projection operation is used to calculate the attention weight, where P m is a learnable operator. The masked node embedding can be calculated using the following formula:
[0038]
[0039] After the Readout layer, we get the node graph embedding after access:
[0040]
[0041] Similar to the visited node graph embedding process, the present invention uses the same mechanism to generate the visited node graph embedding By obtaining the graph embeddings of visited nodes and unvisited nodes, the model can capture and utilize information from multiple relationships to achieve better solution results;
[0042] In the above scheme, the reinforcement learning training algorithm is a Reinforce training strategy with a rolling baseline, and its steps include:
[0043] Initialize the parameters used in deep reinforcement learning model training, actor network parameters θ and baseline b(s), and input training data; interact with the environment, sample trajectories, record rewards and state sequences; for each time step t, calculate the state s t Next select action a t The probability π(a t |s t ) and the baseline b(s t ) value; Calculate the policy gradient Among them G t is the cumulative reward at time step t; update the strategy parameters Repeat the above steps until the stopping condition is reached;
[0044] Compared with the prior art, the present invention has the following advantages and technical effects:
[0045] 1. In the construction of problem models, current logistics distribution usually ignores factors such as customer satisfaction and environmental issues. The present invention aims at the simultaneous delivery and pickup vehicle routing problem, fully considers the delivery cost, delivery time, customer satisfaction and energy consumption issues in the scenario, and constructs an optimization model for the simultaneous delivery and pickup vehicle routing problem.
[0046] 2. The current delivery and pickup vehicle routing problem is mostly solved by heuristic methods. For large-scale complex vehicle routing problem sets, the heuristic algorithm does not have prior knowledge in each solution process because each instance is independent of each other, resulting in a long solution time and easy to fall into the local optimal solution during the solution process. An improved reinforcement learning algorithm based on the attention mechanism is designed, using an enhanced encoder module and an attention aggregation mechanism, so that the model obtains richer context information and improves the solution quality of the model.
[0047] 3. The present invention adopts the Reinforce training strategy with rolling baseline, which can effectively accelerate the convergence of the model. Using the trained model, the learned strategy does not require additional training and search, and only needs to perform a parameterized forward propagation once to obtain a high-quality solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 This is a flow chart of a method for solving the VRPSDPTW problem based on deep reinforcement learning in Example 1 of the present invention.
[0049] Figure 2 Schematic diagram of the strategy network encoding and decoding structure of the reinforcement learning algorithm MARL used in the present invention. DETAILED DESCRIPTION
[0050] In order to make the technical solutions and advantages of the present invention more clearly understood, further detailed description is given below in conjunction with the accompanying drawings and embodiments, but the implementation and protection of the present invention are not limited thereto.
[0051] Example 1
[0052] This embodiment discloses Figure 1 An effective method for solving the simultaneous delivery and pickup vehicle routing problem based on deep reinforcement learning is shown, including the following steps:
[0053] Based on the simultaneous delivery and pickup vehicle routing problem, the objective function is established with the goal of minimizing vehicle usage cost, time window cost and customer satisfaction cost, that is, minimizing the total cost in logistics distribution;
[0054] According to the characteristics of the routing problem of simultaneous delivery and pickup vehicles, a mathematical model is constructed, customer information, vehicle information and distance information are represented by mathematical parameters, and model constraints are determined according to actual conditions;
[0055] The process of gradually selecting service customer nodes and interacting with the environment in the simultaneous delivery and pickup vehicle routing problem is constructed as a Markov decision process, and represented as a four-tuple, namely, state, action, state transition, and reward elements. Sequential decision making is performed according to the criteria of the Markov decision process.
[0056] Build a neural network model to model the selection strategy of vehicles and service nodes;
[0057] Generate a data set according to the sampling rules and preprocess it;
[0058] Use the data set and deep reinforcement learning algorithm to train the constructed neural network model to obtain an application model that meets the requirements;
[0059] The trained deep reinforcement learning model is used to solve logistics distribution scenarios, generate optimal routes, calculate the lowest distribution cost, and output distribution plans.
[0060] In the above scheme, the objective function constructed is specifically vehicle transportation cost, delivery time window cost and customer satisfaction cost, and its specific form is:
[0061] minZ(x)=p1z1(x)+p2z2(x)+p3z3(x)
[0062] Among them, the weighted coefficient relationship is
[0063]
[0064] The components of the objective function are:
[0065] Vehicle transportation costs where c start and c0 are the vehicle startup cost and the vehicle unit distance travel cost respectively;
[0066] Delivery time window cost where c wait and c late are the vehicle unit waiting time cost and unit delay time penalty respectively;
[0067] Customer Satisfaction Cost where c ust and c sat They are unit customer satisfaction cost and unit customer dissatisfaction cost respectively.
[0068] In the above scheme, the mathematical model of the problem is specifically:
[0069] The distribution network consists of a directed graph G =<C,E> The set of Nc customer nodes is represented by C = {0, 1, ..., Nc}, where node 0 represents the distribution warehouse. The feasible routes connecting vehicles in the network are represented by E = {(i, j) | i, j ∈ C, i ≠ j}. The set of delivery vehicles is V = {1, ..., Nv}, where Nv represents the maximum number of available vehicles.
[0070] The model follows the following assumptions: there are enough vehicles with the same capacity available at the warehouse; a customer node can only be served by one vehicle; a vehicle should be returned to the warehouse if it cannot meet customer demand or its capacity is exhausted; the load during vehicle transportation cannot exceed the vehicle capacity.
[0071] The three binary decision variables designed in the mathematical model are:
[0072]
[0073]
[0074]
[0075] In the above scheme, the reinforcement learning quadruplets are
[0076] state t = <X t ,V t >(s t ∈S), which contains two parts of elements, namely X t , used to represent node information and existing partial solutions (paths), V t , which is used to represent the vehicle status at time t.
[0077] Action Space, which is action a t A collection of t ~π(a|s t ), after each step is executed, the visited nodes will be blocked to avoid repeated visits.
[0078] Reward is based on the optimized objective function. The purpose is to minimize the objective function, so the reward function is set to R(x) = -Z(x). The reward setting will make the action selection tend to reduce the value of the objective function, so that the strategy of producing high-quality solutions can be trained;
[0079] Policy uses the following chain rule to express the probability relationship in building a complete path:
[0080]
[0081] The strategy π has trainable parameters θ. At each step of building the solution path, the strategy outputs the probability distribution of node selection. Nodes are selected one by one until the complete path is built. For example, a = {a0, a1, ..., a T}, where T is the maximum number of steps. The strategy π is parameterized as a neural network π with an encoder-decoder structure θ , which is used to solve the routing problem of simultaneous delivery and pickup vehicles;
[0082] In the above scheme, the deep neural network model is a codec structure, and the specific design is described as follows:
[0083] The encoder uses the Transformer module to obtain node embeddings, while the decoder uses the node embeddings and access information masks to generate context embeddings, and then uses a multi-head attention mechanism to obtain the probability distribution of node selection. To select actions, an attention aggregation module is designed to obtain excellent context embeddings to capture dynamic state transitions.
[0084] In the above scheme, the specific structure of the designed attention aggregation module is:
[0085] Modules use node embedding node and access information mask As input, It is set to 1 when the node has been visited and 0 otherwise. The Projection operation is used to calculate the attention weight, where P m is a learnable operator. The masked node embedding can be calculated using the following formula:
[0086]
[0087] After the Readout layer, we get the node graph embedding after access:
[0088]
[0089] Similar to the visited node graph embedding process, the present invention uses the same mechanism to generate the visited node graph embedding By obtaining the graph embeddings of visited nodes and unvisited nodes, the model can capture and utilize information from multiple relationships to achieve better solution results;
[0090] In the above scheme, the reinforcement learning training algorithm is a Reinforce training strategy with a rolling baseline, and its steps include:
[0091] Initialize the parameters used in deep reinforcement learning model training, actor network parameters θ and baseline b(s), and input training data; interact with the environment, sample trajectories, record rewards and state sequences; for each time step t, calculate the state s t Next select action a t The probability π(a t |s t ) and the baseline b(s t ) value; Calculate the policy gradient b(s t )), where G t is the cumulative reward at time step t; update the strategy parameters Repeat the above steps until the stopping condition is reached;
[0092] Example 2
[0093] This example discusses the problem model and related work. The problem model is as follows:
[0094] The distribution network consists of a directed graph G =<C,E> The set of Nc customer nodes is represented by C = {0, 1, ..., Nc}, where node 0 represents the distribution warehouse. The feasible routes connecting vehicles in the network are represented by E = {(i, j) | i, j ∈ C, i ≠ j}. The set of delivery vehicles is V = {1, ..., Nv}, where Nv represents the maximum number of available vehicles.
[0095] The model follows the following assumptions: there are enough vehicles with the same capacity available at the warehouse; a customer node can only be served by one vehicle; a vehicle should be returned to the warehouse if it cannot meet customer demand or its capacity is exhausted; the load during vehicle transportation cannot exceed the vehicle capacity.
[0096] The mathematical symbols and definitions of the design are shown in Table 1:
[0097] The three binary decision variables designed in the mathematical model are:
[0098]
[0099]
[0100]
[0101] The constraints of the model are:
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114] Among the above constraints, 0102-0103 ensure that each customer point is served by only one vehicle, 0104-0105 are vehicle set and vehicle load constraints, 0106-0109 are service time window constraints, 0110-0111 are service and waiting time constraints, and 0112-0113 are the time parameters of the distribution center and route elimination conditions, respectively.
[0115] Example 3
[0116] This example conducts experimental testing and analysis on a designed narrow time window and simultaneous delivery and pickup vehicle routing problem dataset to verify its solution quality and fast solution capability. Excellent heuristic algorithms and reinforcement learning algorithms are selected for testing as comparison.
[0117] This experiment selects four test sets with different numbers of customer points, the number of customer points is N c =10,N c =25,N c =50,N c =100, customer point distribution includes uniform distribution, cluster distribution, and mixed distribution. The warehouse is known and fixed in the same coordinate scenario for each instance, and the effective time window accounts for 1%-5% of the total time.
[0118] This example algorithm MARL and the comparison algorithm heuristic ABSO and reinforcement learning AM algorithm are all programmed in Python, the operating system is Win10, and the hardware is AMD Ryzen 7 5800H with Radeon Graphics 3.20GHz) and RAM (16.0GB), RTX 3070GPU.
[0119] Fixed parameters in the vehicle dispatching scheme are uniformly set in the test. The following parameters have the same description units: unit vehicle load 500, vehicle startup cost 30 / vehicle, unit distance cost 5, unit time waiting cost 10, and unit time delay acceptable cost 15. When a customer is served within the specified time, customer satisfaction increases by 5 units, and when a customer is not served within the specified time, customer satisfaction decreases by 3 units. The target weight values for model optimization are set to 0.5, 0.3, and 0.2, which can be adjusted according to actual conditions.
[0120] Test evaluation parameters
[0121] In the test, ObjectVal and CPU Time are used to judge the solution quality and solution speed. At the same time, in order to compare the solution quality gap between different algorithms, the evaluation index is introduced
[0122] Test results and analysis
[0123] The relevant test parameters in this test are set according to the test scheme described above, and the test results are shown in Table 2-Table 3
[0124] Table 2 Small-scale customer point test results
[0125] Table 3 Large-scale customer point test results
[0126] When the number of client nodes is 10, 25, and 50, the AM and MARL algorithms are significantly better than the ABSO algorithm in terms of solution quality. The MARL algorithm optimizes 7 out of 9 examples. The solution result of MARL is 4.90% higher than that of the ABSO algorithm and 3.2% higher than that of the reinforcement learning algorithm AM. At the same time, MARL still maintains the advantage of fast solution of the reinforcement learning algorithm.
[0127] When the scale of customer nodes increases (100 customer points), the solution time of the heuristic method ABSO increases rapidly. Compared with the ABSO algorithm, the reinforcement learning method AM and MARL algorithm can still quickly solve the simultaneous delivery and pickup vehicle routing problem. At the same time, the MARL algorithm also has a 2.76%-6.88% improvement in solution quality compared to the AM algorithm.
[0128] Based on the above test results and analysis, the example algorithm MARL can take advantage of the fast solution of reinforcement learning for the simultaneous delivery and pickup vehicle routing problem with time windows, and has better solution quality than existing algorithms.
Claims
1. A method for solving the VRPSDPTW problem based on deep reinforcement learning, characterized in that: The following steps are involved: 1) Establish the objective function, construct a mathematical model for the routing problem that takes into account the time window and simultaneous delivery and pickup, and determine the model constraints; 2) building a reinforcement learning environment based on the mathematical model; 3) Building a neural network model based on the attention mechanism according to the reinforcement learning environment; 4) Use the data set and deep reinforcement learning algorithm to train the constructed neural network model to obtain an application model that meets the requirements, and use it to solve logistics distribution scenarios and output distribution plans.
2. The effective method for solving the VRPSDPTW problem based on deep reinforcement learning according to claim 1, characterized in that: The objective function constructed in step 1) is specifically vehicle transportation cost, delivery time window cost and customer satisfaction cost, and its specific form is: minZ(x)=p1z1(x)+p2z2(x)+p3z3(x) Among them, the weighted coefficient relationship is The components of the objective function are: Vehicle driving cost in Delivery time window cost Customer Satisfaction Cost 3. The effective method for solving the VRPSDPTW problem based on deep reinforcement learning according to claim 2 is characterized in that: The mathematical model of the problem constructed in step 1) is specifically: The distribution network consists of a directed graph G =<C,E> The set of Nc customer nodes is represented by C = {0, 1, ..., Nc}, where node 0 represents the distribution warehouse. The feasible routes connecting vehicles in the network are represented by E = {(i, j) | i, j ∈ C, i ≠ j}. The set of delivery vehicles is V = {1, ..., Nv}, where Nv represents the maximum number of available vehicles. The model follows the following assumptions: there are enough vehicles with the same capacity available at the warehouse; a customer node can only be served by one vehicle; a vehicle should be returned to the warehouse if it cannot meet customer demand or its capacity is exhausted; the load during vehicle transportation cannot exceed the vehicle capacity. The three binary decision variables designed in the mathematical model are:
4. The effective method for solving the VRPSDPTW problem based on deep reinforcement learning according to claim 3 is characterized in that: The specific elements of step 2) building a reinforcement learning environment based on logistics scenarios are: state t = <X t , V t >(s t ∈S), which contains two parts of elements, namely X t , used to represent node information and existing partial solutions (paths), V t , which is used to represent the vehicle status at time t. Action Space, which is action a t A collection of t ~π(a|s t ), after each step is executed, the visited nodes will be blocked to avoid repeated visits. Reward is based on the optimized objective function. The purpose is to minimize the objective function, so the reward function is set to R(x) = -Z(x). The reward setting will make the action selection tend to reduce the value of the objective function, so that the strategy of producing high-quality solutions can be trained; Policy uses the following chain rule to express the probability relationship in building a complete path: The strategy π has a trainable parameter θ. At each step of building the solution path, the strategy outputs the probability distribution of node selection. Nodes are selected one by one until the complete path is built. For example, a = {a0, a1, ..., a T }, where T is the maximum number of steps. The strategy π is parameterized as a neural network π with an encoder-decoder structure θ , which is used to solve the routing problem of simultaneous delivery and pickup vehicles.
5. The effective method for solving the VRPSDPTW problem based on deep reinforcement learning according to claim 4 is characterized in that: The neural network model based on the attention mechanism built in step 3) is a codec structure, and the specific design is described as follows: The encoder uses the Transformer module to obtain node embeddings, while the decoder uses the node embeddings and access information masks to generate context embeddings, and then uses a multi-head attention mechanism to obtain the probability distribution of node selection. To select actions, an attention aggregation module is designed to obtain excellent context embeddings to capture dynamic state transitions. In the above scheme, the specific structure of the designed attention aggregation module is: Modules use node embedding node and access information mask As input, When the node is visited, it is set to 1 and otherwise set to 0; the Projection operation is used to calculate the attention weight, where P m is a learnable operator. The masked node embedding can be calculated using the following formula: After the Readout layer, we get the node graph embedding after access: Similar to the visited node graph embedding process, the present invention uses the same mechanism to generate the visited node graph embedding By obtaining the graph embeddings of visited nodes and unvisited nodes, the model can capture and utilize information from multiple relationships to achieve better solution results.
6. The effective method for solving the VRPSDPTW problem based on deep reinforcement learning according to claim 5, characterized in that: The reinforcement learning training algorithm used in step 4) is a Reinforce training strategy with a rolling baseline, and the steps include: Initialize the parameters used in deep reinforcement learning model training, actor network parameters θ and baseline b(s), and input training data; interact with the environment, sample trajectories, record rewards and state sequences; for each time step t, calculate the state s t Next select action a t The probability π(a t |s t ) and the baseline b(s t ) value; Calculate the policy gradient Among them G t is the cumulative reward at time step t; update the strategy parameters Repeat the above steps until the stopping condition is reached.
Citation Information
Patent Citations
Dynamic adjustment-based express delivery distribution optimization method
CN107145971A
Vehicle path planning method and device based on deep reinforcement learning
CN114462687A
Method and system for determining paths of vehicles for simultaneously taking and delivering goods in common delivery mode
CN117933513A
Vehicle path planning method with time window based on deep reinforcement learning
CN118350732A
Heterogeneous vehicle type vehicle path planning method and system based on deep reinforcement learning
CN118608021A