A CVRP Solving Method Based on Memory Pointer Network
Through the method based on memory pointer network, optimized vehicle and user allocation sequences are generated, and the problem of time-consuming solving of super-large-scale CVRP problems is solved in the prior art, achieving more efficient and high-quality solutions.
Patent Information
- Application Number
- CN202210472947.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The prior art takes time to solve ultra-large-scale CVRP problems, and it is difficult for heuristic algorithms and planning solvers to solve efficiently when facing large-scale problems.
The CVRP solution method based on the memory pointer network is adopted to optimize the allocation of vehicles and users by generating sequences, and the memory pointer network is trained using the policy gradient method to enable it to learn the strategies for optimizing the solution of CVRP problems.
The solution speed and quality of CVRP problems have been significantly improved. Compared with the Google operations research solver OR-Tools, the solution results of the memory pointer network are of higher quality.
Smart Images

Figure CN114896878B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for solving the CVRP based on a memory pointer network, belonging to the technical fields of reinforcement learning, deep learning, and combinatorial optimization. Background Art
[0002] The CVRP problem is a very common problem in real life. For example, problems such as the efficiency optimization of logistics distribution and the scheduling of passenger vehicles can be modeled as CVRP problems. The scale of the CVRP problem in actual production and life is extremely large. A distribution center needs to process hundreds or thousands of express deliveries in a day, and this scenario requires challenging the solution of ultra-large-scale CVRP problems in a very short time. This patent proposes a method - the memory pointer network - that can solve ultra-large-scale CVRP problems in a short time.
[0003] Currently, the main methods for solving the CVRP problem are: 1. heuristic algorithms; 2. planning solvers; 3. reinforcement learning solutions. Among them, heuristic algorithms have relatively excellent solution speeds for small-scale problems. There have been many developments in such heuristic algorithms for solving the CVRP problem. However, each iteration of the heuristic algorithm needs to calculate the optimization objective function. In addition, when the problem scale increases, more iterations are required from the initial solution to the relatively optimal solution. Therefore, heuristic algorithms often require extremely long time-consuming when solving ultra-large-scale CVRP problems.
[0004] Solving with a planning solver is also a method for solving the CVRP problem. Mainstream planning solvers show excellent performance in solving linear / nonlinear mixed integer programming and small-scale constraint programming. However, when faced with ultra-large-scale problems, such solvers require a large amount of time-consuming and are difficult to handle the situation of no feasible solution that may occur in actual production.
[0005] In recent years, with the development of machine learning, reinforcement learning, as a new method, has injected fresh blood into the solution of the CVRP problem, and many new methods have emerged. Compared with heuristic algorithms and most solvers, reinforcement learning algorithms have the characteristic of being able to break through the bottleneck of too small a sample size. Therefore, the training model for small-scale problems has better adaptability in solving large-scale problems. Among them, adding an attention mechanism is the most effective in improving the optimization ability of reinforcement learning.
[0006] In view of this, it is indeed necessary to propose a method for solving the CVRP based on a memory pointer network to solve the above problems. Summary of the Invention
[0007] The purpose of the present invention is to provide a CVRP solving method based on a memory pointer network. The role of the memory pointer network is to generate sequences. The input of the memory network has two parts: additional information and target information. The length of the sequence output by the memory pointer network is determined by the length of the target information.
[0008] The present invention provides a CVRP solving method based on a memory pointer network, comprising the following steps:
[0009] Step 1: Model the CVRP problem according to the CVRP problem in actual production to obtain a mathematical model of the CVRP problem. A CVRP problem model has two components: users and vehicles;
[0010] Step 2: Generate a plurality of virtual CVRP problems according to the CVRP problem mathematical model in Step 1, thereby forming a CVRP problem data set;
[0011] Step 3: Preprocess the data set in Step 2;
[0012] Step 4: Initialize the weights of the two memory pointer networks;
[0013] Step 5: Input the users and vehicles in the CVRP problem data set in Step 3 into the two memory pointer networks initialized in Step 4, calculate the vehicle sequence and the user sequence respectively, and train the memory pointer network using the policy gradient method, so that the memory pointer network learns the strategy for optimizing the solution of the CVRP problem. As the training progresses, the quality of the solution obtained by the memory pointer network will gradually improve until it enters a stable state. The output of the memory pointer network is a vehicle sequence and a cargo sequence;
[0014] Step 6: Solve the actual CVRP problem using the trained memory pointer network above to obtain the vehicle sequence and cargo sequence corresponding to the actual CVRP problem;
[0015] Step 7: Use an algorithm to convert the vehicle sequence and cargo sequence obtained in Step 6 into the solution result of the CVRP problem.
[0016] A further limited technical solution of the present invention is:
[0017] For the above-mentioned CVRP solving method based on a memory pointer network, the CVRP problem description and preprocessing in Step 1 and Step 3 include the following content:
[0018] The symbol convention for the CVRP problem mathematical model is as follows:
[0019]
[0020] Among them, node 0 and node n + 1 in set N are distribution centers;
[0021] The mathematical model of the CVRP problem in step 1 described above includes:
[0022] Definition 1: Decision variable The values are as follows:
[0023]
[0024]
[0025] Definition 2: l i is the maximum load of vehicle i (unit: ton);
[0026] Definition 3: User i is expressed as a four-dimensional vector c i = [q i , x i , y i , d 0i as the input of the memory pointer network, where q i is the transportation demand of user i (unit: ton), x i , y i are the coordinates of the user, and d 0i is the distance from user i to the distribution center;
[0027] In summary, the model description of the CVRP problem is as follows:
[0028] min
[0029] s.t.
[0030]
[0031]
[0032]
[0033] For the aforementioned CVRP solving method based on the memory pointer network, the preprocessing method in step 3 is as follows:
[0034]
[0035]
[0036] where c′ i , l′ i are the preprocessed user vector and vehicle load.
[0037] The aforementioned CVRP solving method based on the memory pointer network, where the vehicle sequence and the cargo sequence in step 7 are converted into the solution of the CVRP problem, includes the following content:
[0038] Let the load sequence corresponding to the vehicle sequence be and the transportation demand sequence corresponding to the user sequence be The mapping algorithm from the sequence to the solution of the planning problem is as follows:
[0039] An operation method of the memory pointer network, which is an expression of the policy in the Markov decision process. This policy is trained by the Actor-Critic algorithm and is a network structure. This operation method includes the following steps:
[0040] Step 1: Calculate the Euclidean distance between pairwise vectors in the additional information and the target information vector group as the heuristic information matrix insMat A , insMat T ;
[0041] Step 2: Perform a linear transformation on the input and map it to a high-dimensional linear space;
[0042] Step 3: Use the improved GAT to encode the additional and target information respectively. The adjacency matrices used for encoding are insMat A , insMat T ; Step 3 and Step 4 are the encoder part as Figure 4 shown; If the additional information or the target information needs to focus on the sequence situation, an LSTM encoder is used to encode the information of the sequence to be focused on before the GAT encoder.;
[0043] Step 4: Concatenate the additional information and the user information obtained in Step 2, and use the Transformer structure to encode the concatenated information again;
[0044] Step 5: Initialize the memory query vector;
[0045] Step 6: Use the encoding obtained after Step 3 as the initial memory of the memory network and initialize the memory module;
[0046] Step 7: Input the memory query vector into the memory module, and the context vector output by the memory module is denoted as q. The memory output process is described as follows:
[0047] The i-th segment of memory in the memory module is a matrix with o rows and g columns, denoted as o i is the memory length of the i-th segment, denoted as Let \(l\) be the memory length, and \(g\) be the dimension transformed to the corresponding high-dimensional space, i.e., the dimension of memory. Suppose there are \(m\) segments of memory, and the output of memory is implemented using Multi-Head Attention, and its formal description is as follows:
[0048]
[0049]
[0050] \(P = [p_1(v, M)||p_2(v, M)||\cdots||p\) θ (v, M)];\)
[0051] output = Relu(w_4 T P);
[0052] where both \(w_4\) and \(v\) are parameters of the model and participate in gradient descent. \(\theta\) represents using \(\theta\)-head attention, and \(||\) represents matrix concatenation. The concatenated matrix \(M\) is a matrix of \(o'\) rows and \(g\) columns, and \(P\) is a matrix of \(n\) rows and \(g\) columns. The parameter \(v\) is a model parameter, and this vector remains the same in multiple iterations of the decoder. Its meaning is the query vector of memory. \(w_4\) is a \(\theta\)-dimensional vector, and output is also a \(g\)-dimensional vector. The calculation process of the above three formulas is denoted as: where \(v\) is the memory query vector or a query matrix (each row in this matrix is a query vector);
[0053] Step 8: Use the Attention mechanism to calculate the selection probability of the current target information;
[0054] Step 9: Select the target information according to the selection probability calculated in Step 8 and output the selection number;
[0055] Step 10: Update the content of each memory segment in the memory module using the selected target information above. Memory update is divided into three steps. Let the new information be a \(g\)-dimensional vector \(c''\) i , and the calculation method of the first step of updating any memory segment is as follows:
[0056] part1 = Relu(W_1c'') j );
[0057]
[0058] NewMem = \(\alpha\cdot part1+(1 - \alpha)\cdot part2\);
[0059] The second step is to calculate the memory combination coefficients of the two parts:
[0060]
[0061] Denote the matrix as the extended form of the new memory vector, where o i is the length of this segment of memory, g is the dimension of the memory vector, and this matrix is the matrix formed by replicating the NewMem vector o i times, that is: The calculation method of the third step for updating any segment of memory is as follows:
[0062]
[0063] In the above formula is a o i dimensional vector, where each element is the retention coefficient of each memory vector in this segment of memory, represents the multiplication of the i-th row in and the corresponding i-th element in ;
[0064] Step 11: Mark the item in the target information selected this time as unavailable;
[0065] Step 12: Determine whether there are available target information items. If there are, jump to Step 7; otherwise, end. Finally, the memory pointer network outputs a natural number sequence with the same length as the target information.
[0066] For the operation method of the aforementioned memory pointer network, the improved GAT in Step 3 uses three stacked GAT calculation modules, and the algorithm for one GAT calculation is as follows:
[0067] Suppose there are R nodes on the undirected graph G, and a g-dimensional vector is defined on each node. Denote the vector of the i-th node as Then the attention coefficient between two nodes in the improved multi-head GAT mechanism is a ij is defined as follows:
[0068]
[0069]
[0070] where W k is each model parameter, represents concatenating two vectors into a new vector twice the original dimension, and then performing an inner product operation with the model parameter . The above process obtains the first part of the attention weight. w ij is an element in the adjacency matrix of graph G, and the edges of graph G have weights. Here, insMat is used as the adjacency matrix of graph G, and w ij is an element inside the heuristic information matrix insMat, are model parameters. In the above expression, f is the activation function Leaky ReLU. According to the definition of the above similarity coefficient, the vector obtained by improved multi-head GAT encoding is calculated as follows:
[0071]
[0072]
[0073] where || represents vector concatenation, θ represents the GAT mechanism with θ heads, means stacking the vectors row by row into a matrix, w is a model parameter whose role is to combine the matrices into a vector, and σ is the activation function Elu.
[0074] The operation method of the aforementioned memory pointer network, and the calculation method in step 8 is as follows:
[0075]
[0076]
[0077] In the above expression, Q is the matrix Q composed of replicated context encoding row vectors q ·(a) =[q T , q T , …, q T , the vector q is the output result of the memory module, a is the length of the target information, R is the matrix composed of target information vectors (each column vector is a target information vector). v, W1, and W2 are model parameters participating in gradient descent, where d is the dimension of the vector v. i is the current iteration number, γ is a model parameter, represents the i-th row of the heuristic information matrix insMat T .
[0078] Compared with the prior art, the present invention adopting the above technical solutions has the following technical effects: it can effectively improve the solving speed and quality of the CVRP problem, and has better solving speed and quality compared with the commonly used Google Operations Research Solver at present. The present invention focuses on solving large-scale CVRP problems. In the experiment, the scale of the problem is denoted as "a×b", where this notation means that the problem has a users and b vehicles. For example, the meaning of "1600×400" is that the problem scale is 1600 users and 400 vehicles, and the optimization goal is to minimize the vehicle driving distance. Compared with the Google Operations Research Solver OR-Tools commonly used in related fields, for large-scale problems, the quality of the solution obtained by the Memory Pointer Network (MemPtrN) is significantly better than that of OR-Tools under the same time consumption. The Memory Pointer Network can be accelerated by GPU, and the time-consuming advantage is more obvious under GPU acceleration. The solving time consumption and solution results (unit: second) for various problem scales are shown in the following table:
[0079]
[0080] To better compare the quality of the solutions, the superiority η is defined. Let the solution result of MemPtrN be s1 and the solution result of OR-Tools be s2, then the superiority η is defined as η = ((s2 - s1) / s2)×100%. Under the premise of the above time consumption, the Memory Pointer Network has better solution quality than OR-Tools. When the problem scale is 200×50, that is, under the problem scale used for training, the quality of the solutions obtained by the Memory Pointer Network and OR-Tools is basically the same. It can be seen from the table that when the problem scale increases, there is an obvious gap in the solution quality of OR-Tools compared with the Memory Pointer Network (MemPtrN). BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 is the flowchart of the present invention;
[0082] Figure 2 is the main working process of the present invention for solving CVRP;
[0083] Figure 3 is the main framework of the Memory Pointer Network of the present invention;
[0084] Figure 4 is the main structure of the encoder of the present invention;
[0085] Figure 5 is the GAT encoder structure of the present invention;
[0086] Figure 6 is the memory update process of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0087] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0088] The present invention provides a method for solving the CVRP based on a memory pointer network. The overall solution process is as Figure 1 shown, and the method includes the following steps:
[0089] Step 1: Model the CVRP problem according to the CVRP problem in actual production to obtain a mathematical model of the CVRP problem. A CVRP problem model has two components: users and vehicles;
[0090] Step 2: According to the mathematical model of the CVRP problem in Step 1, specifically generate multiple virtual CVRP problems to form a CVRP problem dataset;
[0091] Step 3: Preprocess the dataset in Step 2;
[0092] Step 4: Initialize the weights of two memory pointer networks;
[0093] Step 5: Input the users and vehicles in the CVRP problem dataset in Step 3 into the two memory pointer networks initialized in Step 4, calculate the vehicle sequence and the user sequence respectively, and use the policy gradient method to train the memory pointer network so that the memory pointer network learns the strategy for optimizing the solution of the CVRP problem. As the training progresses, the quality of the solution obtained by the memory pointer network will gradually improve until it enters a stable state. The output of the memory pointer network is a vehicle sequence and a cargo sequence;
[0094] Step 6: Use the trained memory pointer network above to solve the actual CVRP problem to obtain the vehicle sequence and cargo sequence corresponding to the actual CVRP problem;
[0095] Step 7: Use an algorithm to convert the vehicle sequence and cargo sequence obtained in Step 6 into the solution of the CVRP problem.
[0096] So far, the solution process ends.
[0097] As Figure 1 shown, in this training process, the sequence needs to be mapped to a solution of the CVRP problem. Because only by converting the sequence into a solution can it be considered a real simulation of solving the CVRP problem to achieve the training effect. That is, the actual solution steps need to be carried out in each training iteration. Only in this way can the reward function (driving distance) of the reinforcement learning be calculated, and then the parameters of the memory pointer network can be gradually optimized.
[0098] The symbol conventions for the mathematical model of the CVRP problem in Step 1 are as follows:
[0099]
[0100] Among them, nodes 0 and n + 1 in set N are distribution centers;
[0101] The mathematical model of the CVRP problem in Step 1 includes:
[0102] Definition 1: Decision variable The value is as follows:
[0103]
[0104]
[0105] Definition 2: l i is the maximum load of vehicle i (unit: ton);
[0106] Definition 3: User i is expressed as a four-dimensional vector c i = [q i , x i , y i , d 0i as the input of the memory pointer network, where q i is the transportation demand of user i (unit: ton), x i , y i are the coordinates of the user, and d 0i is the distance from user i to the distribution center;
[0107] In summary, the model description of the CVRP problem is as follows:
[0108] min
[0109] s.t.
[0110]
[0111]
[0112]
[0113] There are two optimization objectives, namely: minimizing the driving distance; maximizing the number of served users. The first constraint is to make the result of the CVRP problem a path, the second constraint is the vehicle load constraint, and the role of the third constraint is to prevent the formation of loops. The last constraint is some basic conditions that the decision variables need to satisfy.
[0114] The preprocessing method in Step 3 is as follows:
[0115]
[0116]
[0117] where c' i , l' i are the preprocessed user vector and vehicle load respectively.
[0118] The input of the memory network in step 5 has two parts: additional information and target information. The length of the output sequence of the memory pointer network is determined by the length of the target information. The role of the first memory pointer network is to generate the vehicle sequence. Therefore, the additional information is the user and the target information is the vehicle. The calculation process of the first memory pointer network mainly includes the following steps:
[0119] Step 5.1: Calculate the heuristic information matrix insMat of the additional information and the target information A , insMat T .
[0120] Step 5.2: Perform a linear transformation on the input and map it to a high-dimensional linear space
[0121] Step 5.3: Use the improved GAT to encode the additional and target information respectively. The adjacency matrices used for encoding are insMat A , insMat T . Steps 5.3 and 5.4 are the encoder part as Figure 4 shown. If the additional information or the target information needs to focus on the sequence situation, an LSTM encoder is used to encode the information that needs to focus on the sequence before the GAT encoder. The encoder structure of the GAT is shown in Figure Figure 5 .
[0122] Step 5.4: Concatenate the additional information and the user information obtained in step 5.2, and use the Transformer structure to encode the concatenated information again.
[0123] Step 5.5: Initialize the memory query vector
[0124] Step 5.6: Use the encoding obtained after step 5.3 as the initial memory of the memory network and initialize the memory module.
[0125] Step 5.7: Input the memory query vector into the memory module to obtain the context vector output by the memory module, denoted as q.
[0126] Step 5.8: Calculate according to the following expression to calculate the current selection probability for the target information:
[0127]
[0128]
[0129] In the above expression, Q is a matrix composed of replicated context encoding row vectors q, where Q = [q ·(a) , q T , …, q T ]. The vector q is the output result of the memory module, a is the length of the target information, and R is a matrix composed of target information vectors (each column vector is a target information vector). v, W1, and W2 are model parameters participating in gradient descent, where d is the dimension of the vector v. i is the current iteration number, and γ is a model parameter. ·(a) =[q T ,q T ,…,q T , the vector q is the output result of the memory module, a is the length of the target information, and R is a matrix composed of target information vectors (each column vector is a target information vector). v, W1, W2 are model parameters participating in gradient descent, where d is the dimension of the vector v. i is the current iteration number, and γ is a model parameter. represents the i-th row of the heuristic information matrix insMat T . T of the
[0130] Step 5.9: Select the target information according to the selection probability calculated in Step 5.8, and output the selection number.
[0131] Step 5.10: Update the content of each memory paragraph in the memory module using the selected target information above. Each memory paragraph adopts the update process as shown in Figure 6 .
[0132] Step 5.11: Mark the selected target information item as unavailable
[0133] Step 5.12: Determine whether there are available target information items. If there are, jump to Step 5.7; otherwise, end. Finally, the memory pointer network outputs a natural number sequence with the same length as the target information.
[0134] The target information of the second memory pointer network is the user additional information "vehicle". Its calculation process is basically the same as that of the first memory pointer network. The only difference in the calculation steps is that an LSTM encoding of the target information is added in Step 5.3.
[0135] The elements in insMat in Step 5.1 are the Euclidean distances between pairs of vectors in the vector group.
[0136] The improved GAT in Step 5.3 adopts a stack of three GAT calculations. The algorithm for one GAT calculation is as follows:
[0137] Suppose there are R nodes on the undirected graph G, and a g-dimensional vector is defined on each node. Denote the vector of the i-th node as Then the attention coefficient between pairs of nodes in the improved multi-head GAT mechanism is a ij ij defined as follows:
[0138]
[0139]
[0140] where W k is each model parameter, represents concatenating two vectors into a new vector with twice the original dimension, and then using the model parameter to perform an inner product operation with it. The above process obtains the first part of the attention weight. w ij is an element in the adjacency matrix of graph G, and the edges of graph G have weights. Here, insMat is used as the adjacency matrix of graph G, w ij is an element within the heuristic information matrix insMat, is the model parameter. f in the above expression is the activation function Leaky ReLU. According to the definition of the above similarity coefficient, the vector obtained through the improved multi-head GAT encoding is calculated as follows:
[0141]
[0142]
[0143] where || represents vector concatenation, θ represents the GAT mechanism with θ heads, represents stacking the vectors into a matrix by rows, w is the model parameter whose role is to combine the matrix into a vector, and σ is the activation function Elu.
[0144] The memory output process in step 5.7 is described as follows: The i-th segment of memory in the memory module is an o-row g-column matrix denoted as o i is the memory length of the i-th segment, denoted as is the memory length, and g is the dimension transformed to the corresponding high-dimensional space, i.e., the dimension of the memory. Suppose there are m segments of memory, and the output of the memory is implemented using Multi-Head Attention, and its formulaic description is as follows:
[0145]
[0146]
[0147] P = [p1(v, M) || p2(v, M) || … || p θ (v, M)];
[0148] output = Relu(w4 T P);
[0149] where Both w4 and v are parameters of the model and participate in gradient descent. θ represents the use of θ-head attention, and || represents matrix concatenation. The concatenated matrix M is a matrix with o' rows and g columns, and P is a matrix with n rows and g columns. The parameter v is a model parameter, and this vector remains the same during multiple iterations of the decoder, meaning it is the query vector for memory. w4 is a θ-dimensional vector, and output is also a g-dimensional vector. The calculation process of the above three formulas is denoted as: where v is the memory query vector or a query matrix (each row in this matrix is a query vector).
[0150] In step 5.10, memory update is divided into three steps: 1. Calculate two new components of memory; 2. Calculate the combination coefficients of the two new memories; 3. Update all memory segments using the new memory vector. Let the new information be the g-dimensional vector c″ i , and the calculation method for the first step of updating any memory segment is as follows:
[0151] part1 = Relu(W1c″ j );
[0152]
[0153] NewMem = α·part1 + (1 - α)·part2;
[0154] The new memory in the memory space consists of two parts: 1. Directly formed by transforming the new information; 2. Formed from the original memory and combined proportionally. α ∈ [0, 1] is a model parameter shared by all memory segments and participates in gradient descent. In the above formula, c″ i is the memory query vector, and NewMem is a g-dimensional vector. The calculation method for the combination coefficients of updating any memory segment is as follows:
[0155]
[0156] Denote the matrix as the extended form of the new memory vector, where o i is the length of this memory segment, g is the dimension of the memory vector, and this matrix is formed by replicating the NewMem vector o i times, that is: The calculation method for the third step of updating any memory segment is as follows:
[0157]
[0158] In the above formula is a vector with o i dimensions, where each element is the retention coefficient of each memory vector in this memory segment, represents Each i-th row in is multiplied by the corresponding i-th element in
[0159] In step 7, it is assumed that: the load sequence corresponding to the vehicle sequence is and the transportation demand sequence corresponding to the user sequence is
[0160] The mapping algorithm from the sequence to the solution of the planning problem is as follows:
[0161]
[0162] As described above, it is only the specific implementation manner in the present invention, but the protection scope of the present invention is not limited thereto. Any transformation or replacement that can be understood and conceived by those familiar with the technology within the technical scope disclosed by the present invention should be covered within the scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for solving CVRP based on a memory pointer network, characterized in that: It includes the following steps: Step 1: Model the CVRP problem according to the CVRP problem in actual production to obtain a mathematical model of the CVRP problem. A CVRP problem model has two components: users and vehicles; Step 2: According to the mathematical model of the CVRP problem in Step 1, specifically generate multiple virtual CVRP problems to form a CVRP problem dataset; Step 3: Preprocess the dataset in Step 2; Step 4: Initialize the weights of two memory pointer networks; Step 5: Input the users and vehicles in the CVRP problem dataset in Step 3 into the two memory pointer networks initialized in Step 4, calculate the vehicle sequence and user sequence respectively, and use the policy gradient method to train the memory pointer network so that the memory pointer network learns the strategy for optimizing the solution of the CVRP problem. As the training progresses, the quality of the solution obtained by the memory pointer network will gradually improve until it enters a stable state. The output of the memory pointer network is a vehicle sequence and a cargo sequence; Step 6: Use the trained memory pointer network above to solve the actual CVRP problem to obtain the vehicle sequence and cargo sequence corresponding to the actual CVRP problem; Step 7: Use an algorithm to convert the vehicle sequence and cargo sequence obtained in Step 6 into the solution result of the CVRP problem; The operation method of the memory pointer network includes the following steps: Step 1: Calculate the Euclidean distance between each pair of vectors in the additional information and the target information vector group as the heuristic information matrix insMat A , insMat T ; Step 2: Perform a linear transformation on the input and map it to a high-dimensional linear space; Step 3: Use the improved GAT to encode the additional and target information respectively, and the adjacency matrices used for encoding are insMat A , insMat T ; If the additional information or the target information needs to focus on the sequence situation, use the LSTM encoder to encode the information that needs to focus on the sequence before the GAT encoder. Step 4: Concatenate the additional information obtained through the linear transformation in Step 2 and the user information, and use the Transformer structure to encode the concatenated information again; Step 5: Initialize the memory query vector; Step 6: Use the encoding obtained in Step 3 as the initial memory of the memory network and initialize the memory module; Step 7: Input the memory query vector into the memory module, and denote the context vector output by the memory module as q. The memory output process is described as follows: The i-th segment of memory in the memory module is a matrix with o rows and g columns, denoted as o i is the memory length of the i-th segment, denoted as is the memory length, g is the dimension of the corresponding high-dimensional space, that is, the dimension of memory. Suppose there are m segments of memory in total, and the output of memory is implemented by Multi-Head Attention. Its formulaic description is as follows: P = [p1(v, M) ∥ p2(v, M) ∥ … ∥ p θ (v, M)]; output = Relu(w4 T P); Among them Both w4 and v are parameters of the model, participating in gradient descent. θ represents the use of θ-head attention, ∥ represents matrix concatenation, and the concatenated matrix M is a matrix of o′ rows and g columns. P is a matrix of n rows and g columns. The parameter v is a model parameter, and this vector remains the same during multiple iterations of the decoder. Its meaning is the query vector of the memory. w4 is a θ-dimensional vector, and output is also a g-dimensional vector. The calculation process of the above four formulas is denoted as: where v is the memory query vector, or a query matrix, and each row in this matrix is a query vector; Step 8: Use the Attention mechanism to calculate the selection probability of the current target information; Step 9: Select the target information according to the selection probability calculated in Step 8 and output the selection number; Step 10: Update the content of each memory paragraph in the memory module using the selected target information above. The memory update is divided into three steps. Let the new information be a g-dimensional vector c″ i , and the calculation method for the first step of any memory update is as follows: part1 = Relu(W1c″ j ); NewMem = α·part1+(1-α)·part2; The second step calculates the memory combination coefficients of the two parts: Denote the matrix as the extended form of the new memory vector, where o i is the length of this segment of memory, g is the dimension of the memory vector, and this matrix is composed of o i copies of the NewMem vector, that is: The calculation method for the third step of updating any segment of memory is as follows: In the above formula is an i -dimensional vector, where each element is the retention coefficient of each memory vector in this segment of memory. represents the multiplication of the i-th row in by the corresponding i-th element in Step 11: Mark the item in the target information selected this time as unavailable; Step 12: Determine whether there are available target information items. If there are, jump to Step 7 of the operation method of the memory pointer network. Otherwise, end. Finally, the memory pointer network outputs a natural number sequence with the same length as the target information.
2. The method for solving CVRP based on a memory pointer network according to claim 1, characterized in that: The CVRP problem description and preprocessing in Step 1 and Step 3 include the following content: The symbol convention of the mathematical model of the CVRP problem is as follows: Among them, node 0 and node n + 1 in set N are distribution centers; The mathematical model of the CVRP problem in Step 1 includes: Definition 1: Decision variable The values are as follows: Definition 2: l i is the maximum load of vehicle No. i, with the unit of ton; Definition 3: The user No. i is expressed as a four-dimensional vector c i = [q i , x i , y i , d 0i as the input of the memory pointer network, where q i is the transportation demand of the user No. i, in tons, x i , y i are the coordinates of the user, and d 0i is the distance from the user No. i to the distribution center; In summary, the model description of the CVRP problem is as follows:
3. The CVRP solving method based on the memory pointer network according to claim 1, characterized in that: The preprocessing method in Step 3 is as follows: where c′ i , l′ i are the user vector and vehicle load after preprocessing.
4. The CVRP solving method based on the memory pointer network according to claim 1, characterized in that: The conversion of the vehicle sequence and cargo sequence in Step 7 into the solution result of the CVRP problem includes the following content: Suppose that the load sequence corresponding to the vehicle sequence is and the transportation demand sequence corresponding to the user sequence is The mapping algorithm from the sequence to the solution of the planning problem is as follows:
5. The CVRP solving method based on the memory pointer network according to claim 1, characterized in that: The improved GAT in step 3 of the memory pointer network operation method described above uses three stacked GAT calculation modules, and the algorithm for one GAT calculation is as follows: Let there be R nodes in the undirected graph G, and a g-dimensional vector is defined on each node. Denote the vector of the i-th node as Then the attention coefficient between two nodes in the improved multi-head GAT mechanism is a ij It is defined as follows: where W k is each model parameter, represents concatenating two vectors into a new vector with twice the original dimension, and then using the model parameter to perform an inner product operation with it. The above process obtains the first part of the attention weight, w ij is an element in the adjacency matrix of graph G. The edges of graph G have weights. Here, insMat is used as the adjacency matrix of graph G, w ij is an element within the heuristic information matrix insMat, is a model parameter. The f in the above expression is the activation function Leaky ReLU, and the vector is the vector obtained through improved multi-head GAT encoding is calculated as follows: where ∥ represents vector concatenation, and θ represents the GAT mechanism with θ heads. It means stacking vectors by rows into a matrix. w is a model parameter whose role is to combine the matrix into a vector, and σ is the activation function Elu.
6. The CVRP solving method based on the memory pointer network according to claim 1, characterized in that: The calculation method in step 8 of the memory pointer network operation method described above is as follows: In the above expression, Q is the matrix Q composed of replicated context encoding row vectors q ·(a) = [q T , q T ,..., q T , the vector q is the output result of the memory module, a is the length of the target information, R is the matrix composed of target information vectors, each column vector is a target information vector, v, W1, and W2 are model parameters participating in gradient descent, where d is the dimension of the vector v, i is the current iteration number, and γ is a model parameter. represents the i-th row of the heuristic information matrix insMat T .