Vehicle Route Planning Method and Device Based on Deep Reinforcement Learning

Through a method based on deep reinforcement learning, a framework for solving vehicle path planning problems is built, and neural network models are used as a destructive strategy to solve the problem of exploring the large randomness and relying on expert knowledge in the existing technology, achieving more efficient and automated vehicle path planning.

CN114462687BActive Publication Date: 2025-05-30SUN YAT SEN UNIV

Patent Information

Application Number
CN202210043667.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-05-30
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

When solving complex constraints and large-scale vehicle path planning problems, the prior art explores artificial heuristics that rely on expert knowledge to design, which is difficult to effectively solve.

Method used

Using a method based on deep reinforcement learning, a framework for solving vehicle path planning problems is built. Through the neural network model as a destructive strategy, the large neighborhood search process is fitted into the Markov decision-making process, and the neural network model is trained through reinforcement learning to optimize vehicle path planning.

Benefits of technology

It improves the solution efficiency and quality of vehicle path planning, shortens the solution time, and can automatically learn adaptive damage strategies in new problem scenarios to avoid relying on expert knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114462687B_ABST
    Figure CN114462687B_ABST
Patent Text Reader

Abstract

The present invention discloses a vehicle path planning method and device based on deep reinforcement learning. The method includes: building a solution framework for the vehicle path planning problem and determining initial parameter information; building a neural network model as a disruption strategy; fitting the large neighborhood search process into a Markov decision process according to the initial parameter information and the disruption strategy; training the neural network model by a reinforcement learning method according to the Markov decision process; and solving the vehicle path planning problem through the trained neural network model to obtain a vehicle path planning result. The present invention can shorten the solution time and ensure the solution quality, and can be widely applied to the field of artificial intelligence technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, in particular to a vehicle path planning method and device based on deep reinforcement learning. Background Technique

[0002] In combinatorial optimization problems, the vehicle routing problem (VRP) is a classic and widely studied problem: given a set of vehicle fleets and a certain number of customers, under the condition of meeting constraints, how to arrange the driving routes of the vehicle fleets (i.e., the order of serving customers) to optimize the set goals (such as the total vehicle distance, the total vehicle time, etc.). In the real environment, problems such as express delivery problems and food delivery problems can be abstracted as VRP problems, but there are often problems such as a large number of customers and complex constraint conditions (time window constraints, delivery order constraints, cargo capacity constraints, etc.).

[0003] In the VRPSDPTW problem, iterative search is a classic method for solving such problems. Currently, the relatively excellent one is the solution framework based on memetic search by Liu et al., and large neighborhood search (LNS) is one of the key components. Large neighborhood search has characteristics such as a large neighborhood range and strong exploration ability, and is a key component to avoid iterative search falling into local optimality, and is also widely used in other problems or other solution frameworks. However, large neighborhood search still has two major problems: First, the exploration randomness is relatively large, and it does not fit large-scale problems and complex constraint scenarios; Second, when facing new problem scenarios, it still relies on expert knowledge to design artificial heuristics.

[0004] In recent years, some scholars have proposed to learn the heuristic rules of local search algorithms through the method of deep reinforcement learning, so as to have better search ability than the search rules designed manually. However, in the VRP problem with complex constraints and large scale, there is currently no method proposed to improve large neighborhood search using deep reinforcement learning. Summary of the Invention

[0005] In view of this, the embodiments of the present invention provide a vehicle path planning method and device based on deep reinforcement learning.

[0006] One aspect of the present invention provides a vehicle path planning method based on deep reinforcement learning, including:

[0007] Build a solution framework for the vehicle path planning problem and determine the initial parameter information;

[0008] Build a neural network model as a destruction strategy;

[0009] According to the initial parameter information and the destruction strategy, fit the large neighborhood search process into a Markov decision process;

[0010] According to the Markov decision process, train a neural network model by a reinforcement learning method;

[0011] Solve the vehicle routing problem through the trained neural network model to obtain a vehicle routing result.

[0012] Optionally, the building of the solution framework for the vehicle routing problem and determining the initial parameter information includes:

[0013] Configure the position features and node features of the target solution in the problem-solving framework;

[0014] Configure the calculation function of the quality of the target solution.

[0015] Optionally, the building of the solution framework for the vehicle routing problem and determining the initial parameter information further includes:

[0016] Perform position encoding on the node sequence to obtain the position features of each node;

[0017] Divide the individual features of the nodes into static features and dynamic features;

[0018] Among them, the static features include two-dimensional coordinates, the quantity of goods received, the quantity of goods delivered, and the service time window; the dynamic features include the waiting time, the maximum cargo capacity of the path where it is located, the current cargo capacity, the distance from the current node to the previous and next nodes on the path where it is located, and the distance between the previous and next nodes.

[0019] Optionally, building the neural network model as a destruction strategy includes:

[0020] Input the node sequence and the individual features of the nodes into the encoder, and the encoder interacts the node position features and the individual features of the nodes to obtain a sequence of node individual feature vectors and a sequence of node position feature vectors;

[0021] Input the node individual feature vectors and node position feature vectors obtained by the encoder into the decoder, and calculate the probability matrix between nodes through the decoder;

[0022] The decoder selects several nodes as the set of nodes to be destroyed according to the probability matrix to obtain a large neighborhood destruction strategy for the current solution;

[0023] Output the selected node set and the action probability.

[0024] Optionally, the interacting of the node position features and the individual features of the nodes to obtain a sequence of node individual feature vectors and a sequence of node position feature vectors includes:

[0025] Linearly map the individual node features to obtain a high-dimensional individual node feature vector;

[0026] Encode the node sequence information through positional encoding to obtain a high-dimensional node position feature vector;

[0027] Extract features from the individual node feature vector and the node position feature vector through three bidirectional collaborative attention layers to obtain an embedded vector sequence of individual node features and an embedded vector sequence of node position encodings;

[0028] Among them, the calculation formula of the individual node feature vector is:

[0029]

[0030] The calculation formula of the node position feature vector is:

[0031]

[0032] Among them, represents the individual node feature vector of node i; W and B are trainable parameters; (x i , y i ) represents the two-dimensional coordinates; represents the node position feature vector of node i; pe(·) represents performing sine positional encoding.

[0033] Optionally, the obtaining of the large neighborhood destruction strategy for the current solution by selecting several nodes as the destroyed node set according to the probability matrix includes:

[0034] Randomly select a node as the initial node;

[0035] Perform a softmax operation on the row of the probability matrix where the initial node is located, set the probability of the selected node to 0, select the second node according to the probability, and then set the probability of the selected node to 0 until Q nodes are selected.

[0036] Optionally, the fitting of the large neighborhood search process into a Markov decision process according to the initial parameter information and the destruction strategy includes:

[0037] Determine the current state according to the node sequence of the current solution and the individual features of each node;

[0038] Determine the action according to the node set output by the neural network;

[0039] Determine the next state according to the node sequence of the repaired solution and the individual features of each node;

[0040] Determine the reward value according to the quality difference of the solutions between the front and back states.

[0041] Another aspect of the embodiments of the present invention further provides a vehicle path planning device based on deep reinforcement learning, including:

[0042] The first module is used to build a solution framework for the vehicle path planning problem and determine the initial parameter information;

[0043] The second module is used to build a neural network model as a disruption strategy;

[0044] The third module is used to fit the large neighborhood search process into a Markov decision process according to the initial parameter information and the disruption strategy;

[0045] The fourth module is used to train the neural network model by a reinforcement learning method according to the Markov decision process;

[0046] The fifth module is used to solve the vehicle path planning problem by the trained neural network model to obtain a vehicle path planning result.

[0047] Another aspect of the embodiments of the present invention further provides an electronic device, including a processor and a memory;

[0048] The memory is used to store a program;

[0049] The processor executes the program to implement the method as described above.

[0050] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium, where the storage medium stores a program, and the program is executed by a processor to implement the method as described above.

[0051] The embodiments of the present invention also disclose a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the method as described above.

[0052] The embodiments of the present invention first build a solution framework for the vehicle path planning problem and determine the initial parameter information; build a neural network model as a disruption strategy; fit the large neighborhood search process into a Markov decision process according to the initial parameter information and the disruption strategy; train the neural network model by a reinforcement learning method according to the Markov decision process; solve the vehicle path planning problem by the trained neural network model to obtain a vehicle path planning result. The present invention can shorten the solution time and ensure the solution quality. Description of the Drawings

[0053] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0054] Figure 1 It is a flowchart of the overall steps provided for the embodiments of the present invention. Specific embodiments

[0055] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0056] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and provide a method for improving large neighborhood search based on deep reinforcement learning. For the VRPSDPTW problem, the present invention is trained in a relatively large-scale problem scenario (the number of nodes is 200). The experimental results show that the neural network destruction strategy of the present invention is superior to the random destruction strategy in terms of solution quality. And after combining the neural network strategy with the artificial heuristic destruction strategy, it can accelerate the iterative search effect. At the same time, the neural network strategy has a certain generalization ability and can have a similar optimization effect when the number of nodes is 250 or 300.

[0057] One aspect of the present invention provides a vehicle routing planning method based on deep reinforcement learning, including:

[0058] Build a solution framework for the vehicle routing planning problem and determine the initial parameter information;

[0059] Build a neural network model as the destruction strategy;

[0060] According to the initial parameter information and the destruction strategy, fit the large neighborhood search process into a Markov decision process;

[0061] According to the Markov decision process, train the neural network model by the reinforcement learning method;

[0062] Solve the vehicle routing planning problem through the trained neural network model to obtain the vehicle routing planning result.

[0063] Optionally, the building of the solution framework for the vehicle routing planning problem and determining the initial parameter information includes:

[0064] Configure the position features and node features of the target solution in the problem solution framework;

[0065] Configure a calculation function for the quality of the target solution.

[0066] Optionally, when building a solution framework for the vehicle path planning problem and determining the initial parameter information, it further includes:

[0067] Perform position encoding on the node sequence to obtain the position features of each node;

[0068] Divide the individual features of the nodes into static features and dynamic features;

[0069] Among them, the static features include two-dimensional coordinates, cargo receiving volume, cargo delivery volume, and service time window; the dynamic features include waiting time, maximum cargo capacity of the path where it is located, current cargo capacity, distance between the current node and the previous and next nodes on the path, and distance between the previous and next nodes.

[0070] Optionally, building the neural network model as a destruction strategy includes:

[0071] Input the node sequence and node individual features into the encoder, and the encoder interacts the node position features and node individual features to obtain a sequence of node individual feature vectors and a sequence of node position feature vectors;

[0072] Input the node individual feature vectors and node position feature vectors obtained by the encoder into the decoder, and calculate the probability matrix between the nodes through the decoder;

[0073] The decoder selects several nodes as the set of destroyed nodes according to the probability matrix to obtain a large neighborhood destruction strategy for the current solution;

[0074] Output the selected node set and action probability.

[0075] Optionally, when interacting the node position features and node individual features to obtain a sequence of node individual feature vectors and a sequence of node position feature vectors, it includes:

[0076] Perform linear mapping on the node individual features to obtain high-dimensional node individual feature vectors;

[0077] Obtain high-dimensional node position feature vectors by performing position encoding on the node sequence information;

[0078] Perform feature extraction on the node individual feature vectors and the node position feature vectors through three bidirectional collaborative attention layers to obtain a sequence of embedded vectors of node individual features and a sequence of embedded vectors of node position encodings;

[0079] Among them, the calculation formula for the node individual feature vector is:

[0080]

[0081] The calculation formula of the node position feature vector is as follows:

[0082]

[0083] Wherein, represents the node individual feature vector of node i; W and B are trainable parameters; (x i , y i ) represents two-dimensional coordinates; represents the node position feature vector of node i; pe(·) represents performing sine position encoding.

[0084] Optionally, the obtaining of the large neighborhood destruction strategy for the current solution by selecting several nodes as the destroyed node set according to the probability matrix includes:

[0085] Randomly select a node as the initial node;

[0086] Perform a softmax operation on the row where the initial node is located in the probability matrix, set the probability of the selected node to 0, select the second node according to the probability, and then set the probability of the selected node to 0 until Q nodes are selected.

[0087] Optionally, the fitting of the large neighborhood search process into a Markov decision process according to the initial parameter information and the destruction strategy includes:

[0088] Determine the current state according to the node sequence of the current solution and the individual features of each node;

[0089] Determine the action according to the node set output by the neural network;

[0090] Determine the next state according to the node sequence of the repaired solution and the individual features of each node;

[0091] Determine the reward value according to the quality difference of the solutions between the previous and current states.

[0092] Another aspect of the embodiments of the present invention further provides a vehicle path planning device based on deep reinforcement learning, including:

[0093] The first module is used to build a solution framework for the vehicle path planning problem and determine the initial parameter information;

[0094] The second module is used to build a neural network model as the destruction strategy;

[0095] The third module is used to fit the large neighborhood search process into a Markov decision process according to the initial parameter information and the destruction strategy;

[0096] The fourth module is used to train a neural network model by means of reinforcement learning according to the Markov decision process;

[0097] The fifth module is used to solve the vehicle path planning problem through the trained neural network model to obtain a vehicle path planning result.

[0098] Another aspect of the embodiments of the present invention also provides an electronic device, including a processor and a memory;

[0099] The memory is used to store programs;

[0100] The processor executes the program to implement the method as described above.

[0101] Another aspect of the embodiments of the present invention also provides a computer-readable storage medium, where the storage medium stores a program, and the program is executed by a processor to implement the method as described above.

[0102] The embodiments of the present invention also disclose a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the method as described above.

[0103] The following combines the specification drawings to describe the specific implementation process of the present invention in detail:

[0104] Aiming at the problems existing in the prior art, the present invention provides a method for improving large neighborhood search based on deep reinforcement learning. The present invention uses a deep neural network and can learn a better destruction strategy through reinforcement learning training, improve the efficiency of neighborhood exploration, and accelerate the iterative search speed. In large-scale problems and scenarios with complex constraints, the destruction strategy trained by the present invention can obtain obvious optimization effects, and in new application scenarios, it can avoid relying on expert knowledge to design artificial heuristics.

[0105] The improved large neighborhood search method based on deep reinforcement learning of the present invention includes the following steps:

[0106] S1. Build a problem-solving framework; set the position features and node features of the solution; set the calculation function of the solution quality;

[0107] S2. Build a neural network model as a destruction strategy. The neural network model includes two parts: an encoder and a decoder. The encoder extracts the position encoding features and node individual features from the state information of the solution, and the decoder calculates the correlation between nodes and obtains the set of removed nodes.

[0108] S3. Fit the large neighborhood search process into a Markov Decision Process (MDP), where the current state s t is the node sequence and node features of the current solution; the action a t is to select some nodes as a destruction strategy for removal through a neural network; the next state s t+1 is the node sequence and node features of the repaired solution; the reward value r t is the difference in solution quality between the previous and current states;

[0109] S4. Use the reinforcement learning method to train the neural network model until the model converges;

[0110] S5. Apply the trained neural network destruction strategy to the Liu algorithm framework to solve the problem;

[0111] As can be seen from the above technical solutions, the present invention improves the existing large neighborhood search method for large-scale and complex constraint problem scenarios, uses a neural network as the destruction strategy for large neighborhood search, and trains the neural network strategy through the method of deep reinforcement learning. Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0112] 1. For old problems, the present invention uses a neural network to learn more comprehensive problem features, and the obtained neural network strategy can be combined with artificial heuristics to accelerate the iterative search speed and optimize the solution effect.

[0113] 2. For new problems, the present invention can obtain a destruction strategy adapted to the problem through deep reinforcement learning, avoiding relying on expert knowledge to design artificial heuristics and optimizing the solution effect for new problems.

[0114] 3. The present invention trains the neural network model through reinforcement learning, which can not only perform one-time (offline) training and multiple (online) solutions, but also has a certain generalization ability and can be directly applied to different scenarios.

[0115] For the vehicle routing problem with simultaneous pickup and delivery with time windows (i.e., the VRPSDPTW problem), the present invention improves the large neighborhood search component therein, uses a neural network strategy as the destruction strategy, speeds up the iterative search speed, and optimizes the solution process.

[0116] A solution s of the VRPSDPTW problem consists of K paths, i.e., solution s = {r 1 , r 2 , …, r K}, where the path r k is the node sequence served by vehicle k. The information of solution s can be divided into node location features and node individual features, and the node individual features can be further divided into static features sti and the dynamic feature dy i . The solution s needs to satisfy the following constraints: (1) All nodes are served and only served once, that is, ∪ r∈s = V and (2) The vehicle satisfies the cargo capacity constraint at any time; (3) All customer nodes are served within their respective service time windows; (4) The vehicle needs to satisfy the pick-up and delivery requirements of the customer nodes at the same time. The optimization objectives of the problem are similar to those of the general VRP problem and can be divided into two parts: the total travel distance of all vehicles and the number of vehicles required. The formal description is as follows:

[0117]

[0118] s.t. s ∈ S

[0119] where c k represents the travel distance of the k-th vehicle in the solution s, s.t. means subject to, and S is the solution space that satisfies the complex constraints of capacity and time window.

[0120] The present invention improves the large neighborhood search component of the existing algorithm framework, uses a neural network as a destruction strategy, and trains it through deep reinforcement learning. The neural network strategy is not only significantly better than the random strategy in terms of solution quality, can avoid relying on expert knowledge to design artificial heuristics in new problem scenarios, but also can combine the neural network strategy and the artificial heuristic strategy to accelerate the effect of iterative search. Specifically, it includes the following steps, where step S1 is the framework construction and feature definition part of the present invention, S2 - S4 constitute the training part of the present invention, and step S5 is the application part of the present invention:

[0121] S1. Build a problem-solving framework; set the feature extraction method of the solution; set the calculation function of the solution quality. Specifically, the present invention uses the VRPSDPTW problem-solving framework, replaces the artificial heuristic (Shaw heuristic) with a neural network as its large neighborhood search destruction strategy; the features of the solution s can be divided into two parts: the first part is the node position feature, and the node sequence is position-encoded (position encoding, PE) to obtain the position features of each node; the second part is the individual features of each node. The individual features of node v i ∈ V can be divided into two categories: the first category is the static feature st i , including the two-dimensional coordinates [x i , y i , the cargo reception and delivery volume [d i , p i , the service time window [a i , b i ; the second category is the dynamic feature dy i , including the waiting time wi , the maximum cargo capacity and the current cargo capacity of the location path [m i , n i , the distances between the current node and the previous and next nodes of the location path, and the distances between the previous and next nodes [o i , p i , q i ; Set the calculation function of the solution quality. Specifically, μ 1 = 100, μ 2 = 1;

[0122] S2. Build a neural network model as a disruption strategy. The neural network model can be divided into two parts: an encoder and a decoder. Extract the features of the current solution s and input them into the neural network. The neural network inputs the selected node set as the disruption strategy.

[0123] The specific process of step S2 includes:

[0124] S21. Input the node sequence and node individual features into the encoder. The encoder interacts the node position features and node individual features to obtain two parts: the embedded vector sequence of node individual features H = (h 1 , h 2 , …, h N ), and the embedded vector sequence of node position encoding G = (g 1 , g 2 , …, g N ).

[0125] Specifically, linearly map the node individual feature x i to obtain a high-dimensional node individual feature vector Pass the node sequence information y i through position encoding (PE) to obtain a high-dimensional node position feature vector The calculation formula is as follows:

[0126]

[0127]

[0128] Among them, represents the initial individual feature vector of node i, the vector dimension is 128, W and B are trainable parameters, represents the initial position feature vector of node i, the vector dimension is 128, and pe(·) represents performing sine position encoding.

[0129] Then extract features through three bidirectional collaborative attention layers, namely:

[0130]

[0131]

[0132]

[0133] Among them, LN(·) is layer normalization (LN); FFN(·) is a feed-forward network layer (FFN), and the specific formula is as follows:

[0134] FFN(h i ; W,B) = max(0, Wh i +B)

[0135] DAC_Att(·) is bidirectional collaborative attention calculation, which can enable the individual feature vector h i and the position feature vector g i to interact, where W Q ,W K ,W V , W O are all trainable parameters, and the further formula is as follows:

[0136]

[0137]

[0138]

[0139]

[0140] Among them, is respectively obtained after passing through the Softmax layer, and Concat(·) represents the vector splicing operation. In the bidirectional collaborative attention, W Q ,W K ,W V the three parameters are the same as those of the classical attention model, and an additional parameter is added to learn the interaction between the node individual features and the node position features, and W O is used to process the concatenated multi-head attention output vector. In actual operation, the present invention sets m = 4, d k = 16.

[0141] S22. Input the node individual feature vector and the node position feature vector obtained by the encoder into the decoder. The decoder first calculates the probability matrix between nodes. Specifically, the feature vectors all pass through a layer of MAX-Pooling first, and the formula is as follows:

[0142]

[0143]

[0144] Among them, are all trainable parameters. Then, through multi-head attention calculation, a node correlation vector is obtained, and then through linear layer integration, a node correlation matrix is obtained. The formula is as follows:

[0145]

[0146]

[0147]

[0148] Among them, are all trainable parameters, m is the number of multi-head attention heads, FFA(·) is the combination of four-layer FFN layers, and the final input scalar. To control the entropy value, the scalar will pass through the Tanh(·) layer and be multiplied by the coefficient C to obtain the final value. The formula is as follows:

[0149]

[0150] S23. The decoder selects Q nodes as the set of damaged nodes according to the probability matrix, and obtains a large neighborhood destruction strategy for the current solution. The specific process is as follows: First, randomly select a node node 1 as the initial node, perform a Softmax operation on the row where node 1 is located in the probability matrix, set the probability of the selected node to 0, select the second node according to the probability, and then set the probability of the selected node to 0, and so on, until Q nodes are selected. The specific formula is:

[0151]

[0152]

[0153] Among them, Sample(·) is to select according to the probability and return the probability of the corresponding index. The final action probability is obtained by multiplying the sub-probabilities. It can be simply proved that the sum of the probabilities of all actions in the action space is 1. The action probability calculation formula is as follows:

[0154]

[0155] S24. Output the selected node set and action probability.

[0156] S3. Fit the large neighborhood search process into a Markov Decision Process (MDP), where the current state is st The node sequence of the current solution and the individual characteristics of each node; action a t The set of nodes output by the neural network; the next state s t+1 The node sequence of the repaired solution and the individual characteristics of each node; the reward value r t Is the quality difference of the solution between the front and back states, and the specific formula is:

[0157] r t = cost(s t ) - cost(s t+1 ).

[0158] S4. Use the reinforcement learning method to train the neural network model until the model converges. Specifically, the present invention uses the REINFORCE algorithm based on policy gradient to train the parameters θ of the neural network model, that is:

[0159]

[0160] Among them, b(S t ) is the baseline used to reduce the gradient variance, which is the average value of the reward values obtained by 3 times of random destruction strategies. During the training process, the gradient of the parameter θ can be approximated by Monte Carlo sampling, and the calculation formula is as follows:

[0161]

[0162] Among them, B is the number of training batch data. Use the above formula to train the parameters of the neural network model until the model converges, and use the Adam optimization method to update the model parameters.

[0163] S5. Combine the neural network destruction strategy obtained by training with the artificial heuristic (Shaw heuristic), and apply it to the existing algorithm framework to solve the problem. Specifically, when entering the large neighborhood search, first use the neural network destruction strategy. If the repaired solution still cannot jump out of the local optimum after local search, the artificial heuristic (Shaw heuristic) will be used to try to jump out.

[0164] The present invention is evaluated by randomly generating test cases as the test set. The test process includes two parts. The first part is the comparison between the neural network destruction strategy and the random destruction strategy to verify that the effective strategy can be learned through the deep reinforcement learning neural network, and the optimization solution effect can be improved on new problems; the second part is the comparison between the combined strategy of the neural network and the Shaw heuristic and the Shaw heuristic only. For the fairness and rationality of the experiment, in the Shaw heuristic only method, the Shaw heuristic is used continuously twice, and the exploration times are the same as those of the combined strategy.

[0165] Only use the example with N = 200 when training the model. In the test phase, both parts contain three subsets with N = 200, N = 250, and N = 300 (N is the number of customer nodes) respectively to test the generalization ability of the model. Each subset consists of 100 randomly generated examples, and each example is run independently 10 times. Calculate the mean and variance of two metrics to measure.

[0166] Comparison experiment results (N = 200) between the neural network combination strategy of the present invention and the artificial heuristic strategy only.

[0167] It can be seen from the above experimental results that the present invention is superior to the random destruction strategy on all subsets, and the optimization effect of the present invention is more obvious as the problem scale expands. The combination effect of the present invention and the Shaw heuristic is better than only using the Shaw heuristic, and while reducing the solution time for large-scale problems, it maintains the quality of the solution. Therefore, the present invention can learn an effective large neighborhood search destruction strategy through deep reinforcement learning to optimize the existing large neighborhood search method.

[0168] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are foreseeable, where the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.

[0169] In addition, although the present invention is described in the context of functional modules, it should be understood that unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More precisely, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the routine skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation using ordinary skills. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0170] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0171] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a predefined sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0172] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber device, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as necessary, and then storing it in a computer memory.

[0173] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0174] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0175] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0176] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A vehicle path planning method based on deep reinforcement learning, characterized in that, it includes: Construct a solution framework for the vehicle path planning problem and determine the initial parameter information; Construct a neural network model as a destruction strategy; According to the initial parameter information and the destruction strategy, fit the large neighborhood search process into a Markov decision process; According to the Markov decision process, train the neural network model by a reinforcement learning method; Solve the vehicle path planning problem through the trained neural network model to obtain a vehicle path planning result; The constructing a neural network model as a destruction strategy includes: Input the node sequence and node individual features into an encoder, and the encoder interacts the node position features and node individual features to obtain a sequence of node individual feature vectors and a sequence of node position feature vectors; Input the node individual feature vector and node position feature vector obtained by the encoder into a decoder, and calculate the probability matrix between nodes through the decoder; The decoder selects several nodes as the destroyed node set according to the probability matrix to obtain a large neighborhood destruction strategy for the current solution; Output the selected node set and action probability.

2. The vehicle path planning method based on deep reinforcement learning according to claim 1, characterized in that, the constructing a solution framework for the vehicle path planning problem and determining the initial parameter information includes: Configure the position features and node features of the target solution in the problem solving framework; Configure the calculation function of the quality of the target solution.

3. The vehicle path planning method based on deep reinforcement learning according to claim 2, characterized in that, the constructing a solution framework for the vehicle path planning problem and determining the initial parameter information further includes: Perform position encoding on the node sequence to obtain the position features of each node; Divide the individual features of the nodes into static features and dynamic features; Among them, the static features include two-dimensional coordinates, cargo receiving volume, cargo delivery volume, and service time window; the dynamic features include waiting time, maximum cargo capacity of the path where it is located, current cargo capacity, distance between the current node and the front and rear nodes of the path where it is located, and distance between the front and rear nodes.

4. The vehicle path planning method based on deep reinforcement learning according to claim 1, characterized in that, the interacting the node position features and node individual features to obtain a sequence of node individual feature vectors and a sequence of node position feature vectors includes: Perform linear mapping on the node individual features to obtain high-dimensional node individual feature vectors; Obtain high-dimensional node position feature vectors through position encoding of the node sequence information; Extract features from the node individual feature vector and the node position feature vector through three bidirectional collaborative attention layers to obtain an embedded vector sequence of node individual features and an embedded vector sequence of node position encodings; Among them, the calculation formula of the node individual feature vector is: The calculation formula of the node position feature vector is: Among them, represents the node individual feature vector of node i; W and B are trainable parameters; (x i , y i ) represents two-dimensional coordinates; represents the node position feature vector of node i; pe(·) represents performing sine position encoding.

5. The vehicle path planning method based on deep reinforcement learning according to claim 1, characterized in that, Selecting several nodes as the set of damaged nodes according to the probability matrix to obtain a large neighborhood destruction strategy for the current solution, including: Randomly select a node as the initial node; Perform a softmax operation on the row of the probability matrix where the initial node is located, set the probability of the selected node to 0, select the second node according to the probability, and then set the probability of the selected node to 0 until Q nodes are selected.

6. The vehicle path planning method based on deep reinforcement learning according to claim 1, characterized in that Fitting the large neighborhood search process into a Markov decision process according to the initial parameter information and the destruction strategy, including: Determining the current state according to the node sequence of the current solution and the individual characteristics of each node; Determining the action according to the set of nodes output by the neural network; Determining the next state according to the node sequence of the repaired solution and the individual characteristics of each node; Determining the reward value according to the quality difference of the solutions between the previous and current states.

7. A vehicle path planning device based on deep reinforcement learning, characterized in that including: The first module is used to build a solution framework for the vehicle path planning problem and determine the initial parameter information; The second module is used to build a neural network model as the destruction strategy; The third module is used to fit the large neighborhood search process into a Markov decision process according to the initial parameter information and the destruction strategy; The fourth module is used to train the neural network model by a reinforcement learning method according to the Markov decision process; The fifth module is used to solve the vehicle path planning problem through the trained neural network model to obtain the vehicle path planning result; The second module is specifically used for: Inputting the node sequence and the individual characteristics of the nodes into the encoder, and the encoder interacts the node position characteristics and the individual characteristics of the nodes to obtain a sequence of node individual feature vectors and a sequence of node position feature vectors; Inputting the node individual feature vector and the node position feature vector obtained by the encoder into the decoder, and calculating the probability matrix between the nodes through the decoder; The decoder selects several nodes as the set of damaged nodes according to the probability matrix to obtain a large neighborhood destruction strategy for the current solution; Outputting the selected node set and the action probability.

8. An electronic device, characterized in that including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The storage medium stores a program, and the program is executed by the processor to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Autonomous vehicle distribution path planning method based on self-adaptive large neighborhood search algorithm

    CN111798067A

  • Path planning method and device based on reinforcement learning

    CN112507520A

Cited By

  • Mobile robot path planning algorithm combining fuzzy control and reinforcement learning

    CN115826581A

  • A mobile robot path planning algorithm combining fuzzy control and reinforcement learning

    CN115826581B