Efficient reinforcement learning method for solving batch traveling salesman problem in express parcel distribution
By introducing the planning information design and integration module, the diversity group exploration module and the coarse-grained three-point search method in the reinforcement learning method, the problems of inefficient exploration and insufficient generalization capabilities in the combination optimization problem are solved, and efficient solutions to the problem of batch travel merchants in express parcel delivery are achieved.
Patent Information
- Application Number
- CN202510346705.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-13
AI Technical Summary
The existing reinforcement learning methods face the problems of inexploratory efficiency and insufficient generalization ability when solving the combination optimization problem, especially in the solution to the batch travel merchant problem (TSP) in express parcel delivery.
An efficient reinforcement learning method is adopted, including the construction of a planning information design and integration module, a diversity group exploration module and a coarse-grained three-point search method. Through planning information guidance, diversification strategies and Gaussian noise perturbation, the exploration efficiency and generalization capabilities of the model are improved.
It significantly improves the batch combination optimization solution capability of neural solvers in practical scenarios such as express logistics, enhances the model's strategy robustness and generalization performance, and ensures the shortest total distribution distance.
Smart Images

Figure CN120146353A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and machine learning, and particularly to an efficient reinforcement learning method for solving the batch traveling salesman problem in express package delivery. Background Art
[0002] The combinatorial optimization problem refers to the problem of finding the optimal solution in a finite solution space, usually involving discrete variables and combinatorial structures. These problems have extensive applications in fields such as computer science, operations research, logistics, and communication. A remarkable feature of combinatorial optimization problems is that the solution space is usually very large, and as the problem scale increases, the number of solutions grows exponentially, making it extremely difficult to find the global optimal solution. Conventional combinatorial optimization problems include TSP and CVRP. The following is a brief introduction to these two problems:
[0003] Definition of the TSP (Traveling Salesman Problem):
[0004] The goal of the TSP problem is to find the shortest path such that the traveling salesman starts from one city, passes through all other cities exactly once, and then returns to the origin city.
[0005] Definition of the CVRP (Capacitated Vehicle Routing Problem):
[0006] The goal of the CVRP problem is to design the shortest route while satisfying the vehicle capacity limit, passing through all other cities exactly once, and then returning to the depot city. The depot city can be visited multiple times, and the capacity will be reset to 0 each time it is visited.
[0007] Reinforcement learning (RL) provides an effective solution. Reinforcement learning enables an agent to interact with the environment to learn the optimal policy to maximize the cumulative reward.
[0008] POMO (Policy Optimization with Multiple Optima) is an end-to-end method based on reinforcement learning for constructing a heuristic solver for solving combinatorial optimization problems (CO). The core idea of POMO is to utilize the symmetry in the CO solution representation and, through a modified REINFORCE algorithm, force diverse rollouts towards all optimal solutions. This method uses a low-variance baseline during training, making RL training fast and stable and more resistant to local minima.
[0009] POMO is based on a sequence-to-sequence architecture and uses an attention mechanism. The model is mainly divided into two parts: an encoder and a decoder.
[0010] The encoder embeds the features of each node into a high-dimensional space. Specifically, the embedding of each node \(i\) is denoted as \(h\) i , where \((1\leq i\leq n)\). The mean of all node embeddings is denoted as
[0011] The decoder uses the attention mechanism to generate the solution sequence. In POMO, the decoder uses multiple different context node embeddings The definition of each context node embedding is as follows:
[0012]
[0013] where \(t\) is the number of iterations, is the embedding of the \(t\)-th selected node. At \(t = 1\), POMO does not use context node embeddings but directly defines:
[0014]
[0015] POMO uses the modified REINFORCE algorithm for training, and the specific steps are as follows:
[0016] Initialize the neural network parameters \(\theta\). Select a set of starting nodes \(\{1, 2, \ldots, N\}\).
[0017] For each starting node \(i\), generate a trajectory \(\tau\) i . Each trajectory \(\tau\) i starts from the starting node \(i\) and gradually selects the next node through the decoder until all nodes have been visited.
[0018] Use the attention mechanism to calculate the selection probability of each node:
[0019]
[0020] where \(u\) j is the embedding of node \(j\), \(v\) t is the query vector of the decoder, is the set of unvisited nodes.
[0021] Calculate the reward: For each trajectory \(\tau\) i , calculate its total path length \(R(\tau\) i ).
[0022] Use the REINFORCE algorithm to update the neural network parameters:
[0023]
[0024] where \(b\) is the baseline, usually taking the mean of all trajectory rewards to reduce variance.
[0025]
[0026] During the inference stage, POMO uses a variety of greedy rollout and instance augmentation techniques to further reduce the optimality gap. The specific steps are as follows:
[0027] Generate multiple trajectories. For each starting node i, generate multiple greedy trajectories
[0028] Augment each instance to generate multiple variants. For example, for the TSP problem, the node coordinates can be rotated and flipped to generate multiple equivalent instances.
[0029] Select the optimal trajectory. Select the trajectory with the shortest total path length from all the generated trajectories as the final solution.
[0030] It should be noted that the information disclosed in the above background art section is only used for understanding the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0031] The main object of the present invention is to overcome the defects existing in the above background art, and provide an efficient reinforcement learning method for solving the batch traveling salesman problem in express parcel delivery.
[0032] To achieve the above object, the present invention adopts the following technical solutions:
[0033] An efficient reinforcement learning method for solving the batch traveling salesman problem (TSP) in express parcel delivery, comprising the following steps:
[0034] S1. Construct a planning information design and fusion module: Based on the graph structure environment model of the express delivery task, design a planning strategy. The graph structure environment model includes the node coordinates of the distribution center and customer points and their distance matrix, and integrates non-parametric planning information into the decision-making process of the neural solver to guide the path exploration direction and improve the exploration efficiency of batch tasks;
[0035] S2. Design a diversity population exploration module: Adopt an architecture of a shared encoder - multiple decoders, endow each decoder with diverse decoding strategies for multiple delivery tasks through different planning information intensity control parameters, and introduce a population cooperation baseline. Store all the delivery path plans generated by the decoders in a shared buffer, and calculate their average delivery distance as the baseline to promote strategy diversity and improve the overall exploration ability of batch tasks;
[0036] S3. Implement the coarse-grained ternary search method: For the approximate unimodal relationship between the planning information intensity and the generalization performance in multiple delivery tasks, utilize Gaussian noise perturbation combined with the ternary search mechanism to quickly lock in the optimal planning information intensity interval for different delivery tasks, optimize the generalization ability of the neural solver in batch tasks, and minimize the total delivery distance.
[0037] Further, in step S1, the generation of the planning strategy includes the following steps:
[0038] Based on the graph-structured environment model of the express delivery task, construct the distance information matrix between the distribution center and the customer points. The graph-structured environment model includes node coordinates and the Euclidean distance calculated based on the coordinates; linearly normalize the distance information between nodes into a probability distribution to form the planning strategy for the delivery path.
[0039] Generate the planning information matrix through iterative sampling and incorporate the planning information matrix as non-parametric guiding information into the path decision-making process of the neural solver.
[0040] In the single-step decision of the neural solver, fuse the planning information with the strategy of the neural solver through the Softmax normalization method to generate the final action selection probability distribution, so as to guide the neural solver to preferentially select the access order of customer points with shorter distances.
[0041] Further, in step S1, the fusion of the planning strategy is achieved by introducing the balance parameter λ to adjust the weights between the planning information and the strategy of the neural solver, so as to control the tendency of the neural solver between relying on the delivery path distance information (planning information) and its own learning strategy.
[0042] Further, in step S2, the construction of the diversity group exploration module includes the following steps:
[0043] Adopt the architecture of a shared encoder - multiple decoders. Each decoder has the same neural network structure, but is given diverse path generation strategies through different planning information intensity control parameters λ to adapt to the differentiated requirements of multi-distribution center parallel tasks.
[0044] Introduce the group cooperation baseline. Store all the delivery path plans generated by the decoders in a shared buffer and calculate their average delivery distance as the baseline to guide the overall exploration direction.
[0045] Through the group cooperation baseline, enable the decoders that generate longer paths to learn from the decoders that generate shorter paths, and reduce the ineffective exploration of the access order of inefficient customer points.
[0046] Further, in step S2, the calculation method of the group cooperation baseline is as follows:
[0047] Evaluate the delivery path plans generated by each decoder and calculate their total delivery distances;
[0048] Take the average of the delivery distances of all decoders as the group collaboration baseline, which is used to update the strategy of the neural solver, prompting the model to optimize the goal towards minimizing the total distance of multiple tasks.
[0049] Furthermore, in step S3, the implementation of the coarse-grained ternary search method includes the following steps:
[0050] For the customer point scale and distribution characteristics in different delivery tasks, use Gaussian noise perturbation to expand the search interval to avoid path detours caused by local optima;
[0051] Gradually narrow the search range through the ternary search mechanism to quickly lock in the optimal planning information intensity interval suitable for the current delivery task;
[0052] In each step of the search, introduce Gaussian noise perturbation to generate multiple perturbation parameters, evaluate the total delivery distance corresponding to them, reduce the dependence on the precise search of a single parameter, and accelerate the discovery of the global optimal solution.
[0053] Furthermore, in step S3, the specific steps of the ternary search mechanism include:
[0054] Divide the search interval into three equal parts and calculate the total delivery distance of the middle point;
[0055] Expand the middle point through Gaussian noise perturbation, evaluate the solution quality of the surrounding intervals, and preferentially retain the interval with a shorter total distance;
[0056] According to the comparison results of the solution quality, gradually narrow the search range until the optimal planning information intensity interval suitable for the current delivery task scale is found.
[0057] Furthermore, in step S3, the way to introduce Gaussian noise perturbation is:
[0058] In each step of the search, apply Gaussian noise perturbation to the middle point to generate multiple perturbation points;
[0059] Determine the next search direction by evaluating the total delivery distance generated by the perturbation points, avoiding the local optimal path trap caused by a fixed step size.
[0060] Furthermore, in step S3, the termination condition of the coarse-grained ternary search method is:
[0061] When the width of the search interval is less than the preset threshold, stop the search;
[0062] Randomly sample multiple planning information intensity values from the final search range, and select the value that generates the shortest total delivery distance among them as the final optimal planning information intensity, which is used to drive the neural solver to output the optimal path for the batch delivery task.
[0063] A computer program product includes a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0064] The present invention has the following beneficial effects:
[0065] By constructing a planning information design and fusion module, a diversity population exploration module, and a coarse-grained ternary search method, the present invention effectively solves the challenges of low exploration efficiency and insufficient generalization ability faced by existing reinforcement learning in solving combinatorial optimization problems. First, in the scenario of the batch traveling salesman problem (TSP) for express parcel delivery, a non-parametric planning information based on delivery node coordinates and distance matrix is introduced as a guide and integrated into the path decision-making process of the neural solver, providing an efficient exploration direction for the model, significantly improving the exploration efficiency of multi-delivery tasks during the training phase, and enhancing the model's adaptability to large-scale customer point distributions through fine-grained distance information guidance. Second, for the requirements of multi-distribution center parallel tasks, an architecture with a shared encoder and multiple decoders is adopted, combined with different planning information intensity control parameters, to endow the decoders with diverse path generation strategies. At the same time, a population cooperation baseline is introduced, and the average delivery distance of the delivery path plans generated by each decoder is calculated as an optimization benchmark through a shared buffer, promoting strategy diversity and improving the overall exploration ability of batch tasks, further enhancing the strategy robustness of the model in complex delivery networks. In addition, for the generalization problem of customer point scale and distribution in different delivery tasks, a coarse-grained ternary search method is proposed, which uses Gaussian noise perturbation combined with the ternary search mechanism to quickly lock the optimal planning information intensity interval for each delivery task, optimizing the generalization performance of the neural solver in cross-scale and cross-distribution scenarios and ensuring the shortest total delivery distance. These innovative means not only improve the batch combinatorial optimization solving ability of the neural solver in practical scenarios such as express logistics, but also promote the wide applicability of reinforcement learning in industrial-level path planning tasks, bringing significant efficiency improvement and technological innovation to the field of intelligent logistics scheduling.
[0066] Other beneficial effects in the embodiments of the present invention will be further described below. Description of the Drawings
[0067] Figure 1 Schematic diagram for visualizing planning information of combinatorial optimization problems under multiple distributions;
[0068] Figure 2 Schematic diagram of strategy exploration diversity under the action of planning information with different intensities;
[0069] Figure 3 Schematic diagram of the intelligent population framework guided by planning proposed by the present invention (left), and schematic diagram of the planning strategy (right);
[0070] Figure 4 Schematic diagram of the influence of the planning information intensity on the approximate unimodal characteristic of the generalization performance of the neural solver under zero-shot transfer;
[0071] Figure 5 Performance visualization on the traveling salesman problems with different scales of uniform distribution (Omni-POMO is only evaluated under the cross-scale experimental setting), and the metric is normalized using 1 / Gap, aiming to emphasize the gap of various solvers approaching the optimality of various problems;
[0072] Figure 6 Generalization performance visualization comparison under various distributions with a scale of 1000 nodes. The left is the traveling salesman problem (TSP), and the right is the capacitated vehicle routing problem (CVRP). The metric is normalized using 1 / Gap, aiming to emphasize the gap of various solvers approaching the optimality of various problems;
[0073] Figure 7 Performance influence curve of populations with different scales on the neural solver. MDAM can be regarded as a version without a cooperation baseline;
[0074] Figure 8 Visualization of the solution to the traveling salesman problem by the present invention, where Uniform comes from the artificial dataset and rat comes from the real-world dataset of TSPLIB;
[0075] Figure 9 Flowchart of the efficient reinforcement learning method for solving the batch traveling salesman problem in express parcel delivery by the present invention. Detailed implementation manners
[0076] The following makes a detailed description of the implementation manners of the present invention. It should be emphasized that the following description is merely exemplary and not intended to limit the scope of the present invention and its applications.
[0077] The algorithm of the present invention can address the challenges of low exploration efficiency and poor generalization ability in the field of Reinforcement Learning for Combinatorial Optimization (RL4CO). Although the application of reinforcement learning technology to achieve efficient solution of combinatorial optimization problems has been proven feasible in existing work, there are still problems of low exploration efficiency in training and poor generalization ability in the inference process. In existing reinforcement learning-based solution methods, in the face of the huge action and state spaces of combinatorial optimization problems, there is a lack of efficient guiding information, and the exploration ability plays a core role in the training process of reinforcement learning methods. In the inference process of neural solvers, due to the differences in problem distributions, performance degradation often occurs, that is, the solution effect is good on the training set, but the performance drops significantly on the test set. As a non-parametric guiding information, planning information can, to a certain extent, compensate for the perception bias of neural solvers caused by distribution deviations to improve generalization performance. Therefore, the present invention introduces exploration information into the reinforcement learning-based neural solver to construct a neural solver with higher performance.
[0078] (1) New business requirements: Specifically, the technical solution of the present invention aims to address the problems of insufficient exploration efficiency and generalization performance of neural solvers based on deep reinforcement learning in solving combinatorial optimization problems. Combinatorial optimization problems are widely applied in many scientific research and industrial production fields such as path planning, task scheduling, network optimization, protein synthesis, and drug design. Effectively solving these problems is crucial for reducing costs and improving efficiency. Traditional optimization methods highly rely on domain expert knowledge and require careful manual design when dealing with complex problems, resulting in low application efficiency. Using reinforcement learning methods to construct neural solvers is a promising research direction, which optimizes its own solution strategy by interacting with the environment to reduce the dependence on complex domain knowledge and achieve rapid solution. Therefore, the technical solution of the present invention aims to improve the existing neural solvers, combined with the guidance of planning information, to construct a neural solver with higher performance to solve combinatorial optimization problems.
[0079] (2) Problem Solving: The existing problems are that it is difficult for existing neural solvers to explore the huge action and state spaces of combinatorial optimization problems during training, and it is difficult to generalize to instances outside the training distribution during inference. The reason for the low exploration efficiency is that combinatorial optimization problems have a huge action space, and the existing exploration strategies of reinforcement learning methods are often designed for narrow action spaces and are difficult to be directly applied to combinatorial optimization problems; the reason for the low generalization performance is that deep reinforcement learning models are prone to overfitting to the training distribution during training, and when directly applied to out-of-distribution problem instances, due to their different scale differences, the model cannot adapt to the fine-grained decision-making process. This makes it difficult for existing neural solvers to be applied to large-scale tasks in the real world. This performance gap limits the practical application and popularization of neural solvers. The new technical solution aims to solve the performance problems of traditional existing neural solvers by introducing planning guidance information, a multi-population strategy architecture, and a coarse-grained ternary search method. The planning guidance information can improve the exploration efficiency during training and provide fine-grained decision-making information guidance. The multi-population strategy architecture can improve the strategy robustness and overall performance of the model, while the coarse-grained ternary search method can improve the generalization performance of the neural solver when facing unseen distribution instances. The solution of these problems will enable the neural solver to achieve better performance in solving diverse combinatorial optimization problems, thereby promoting the expansion of its application scope in practical applications.
[0080] Specifically, the embodiment of the present invention provides an efficient reinforcement learning method for solving the capacitated vehicle routing problem (CVRP) in express package delivery, including the following steps (see Figure 9 ):
[0081] S1. Construct a planning information design and integration module: Based on the graph-structured environment model of the express delivery task, design a planning strategy. The graph-structured environment model includes the node coordinates and distance matrix of the distribution center and customer points, and integrate non-parametric planning information into the decision-making process of the neural solver to guide the path exploration direction and improve the exploration efficiency of the capacitated task;
[0082] S2. Design a diversity population exploration module: Adopt an architecture of a shared encoder - multiple decoders, endow each decoder with diverse decoding strategies for multi-delivery tasks through different planning information intensity control parameters, and introduce a population cooperation baseline. Store the delivery path plans generated by all decoders through a shared buffer, and calculate their average delivery distance as the baseline to promote strategy diversity and improve the overall exploration ability of the capacitated task;
[0083] S3. Implement the coarse-grained ternary search method: For the approximate unimodal relationship between the planning information intensity and the generalization performance in multiple delivery tasks, use Gaussian noise perturbation combined with the ternary search mechanism to quickly lock in the optimal planning information intensity interval for different delivery tasks, optimize the generalization ability of the neural solver in batch tasks, and minimize the total delivery distance.
[0084] The present invention designs Neural Plan, an architecture for efficiently solving combinatorial optimization problems using reinforcement learning. The present invention addresses the challenges of low exploration efficiency and poor generalization ability in the field of Reinforcement Learning for Combinatorial Optimization (RL4CO). Although the application of reinforcement learning techniques to achieve efficient solution of combinatorial optimization problems has been proven feasible in existing work, there are still problems of low exploration efficiency during training and poor generalization ability during the inference process. In existing reinforcement learning-based solution methods, in the face of the huge action and state spaces of combinatorial optimization problems, there is a lack of efficient guiding information, and the exploration ability plays a core role in the training process of reinforcement learning methods. During the inference process of the neural solver, due to the difference in problem distributions, there is often a performance degradation phenomenon, that is, the solution effect is good on the training set, but the performance drops significantly on the test set. Planning information, as a non-parametric guiding information, can to some extent compensate for the perceptual bias of the neural solver caused by distribution deviation to improve the generalization performance. Therefore, the goal of the present invention is to introduce exploration information into the reinforcement learning-based neural solver to construct a higher-performance neural solver.
[0085] Planning information design and fusion module: In Neural Plan, planning information is introduced as the decision guidance for the neural solver. A planning strategy is constructed through the combinatorial optimization problem environment model, and non-parametric planning information is fused into the single-step decision-making process of the neural solver, thereby improving the exploration efficiency and fine-grained information perception of the neural solver.
[0086] Diverse population exploration module: Through the diverse population exploration module, the existing neural solver is extended to a population scale, and diverse internal strategies and shared cooperation baselines are constructed according to different planning information intensities, enabling the model to better explore the huge strategy space of combinatorial optimization problems, thereby improving the robustness and exploration performance of the model.
[0087] Coarse-grained ternary search method module: To apply the pre-trained neural solver to a wider range of problems, a coarse-grained ternary search method module is introduced for the approximate unimodal relationship between the hyperparameter λ that controls the intensity of planning information and the solution quality. Gaussian noise perturbation is used to randomly expand the search interval to avoid potential local optimal traps. And the ternary search mechanism is fully utilized to efficiently narrow the search range step by step, further optimizing the generalization performance of the model.
[0088] By means of introducing planning information, constructing a diverse strategy population, and introducing the coarse-grained ternary search method, Neural Plan performs excellently in solving combinatorial optimization problems, and has strong practical application value and competitiveness. These highlights make Neural Plan an important innovation in the field of applying reinforcement learning to solve combinatorial optimization problems, and are expected to play an important role in future research and applications.
[0089] The following is a detailed description of the algorithm of the embodiments of the present invention:
[0090] Overall architecture: The architecture of Neural Plan consists of three main parts: the planning information design and fusion module, the diverse population exploration module, and the coarse-grained ternary search method.
[0091] 1. Planning information design and fusion module: The main task of the planning information design and fusion module is to design a planning strategy through the environmental model of the combinatorial optimization problem and integrate it as guiding information into the decision-making process of the neural solver. In the planning strategy, first, the environmental model of the combinatorial optimization problem is defined. Combinatorial optimization problems can often be modeled using a graph structure, and the distance information matrix on the graph can be used as the definition of the environmental model. The distances from each point in its distance matrix to other points are linearly normalized into a probability distribution, which is used as the decision-making distribution of the planning strategy. By defining the planning depth L, iterative sampling is continuously performed L times according to the planning strategy, and the sampling information is used as the planning information matrix. In the process of fusing the planning information, the neural solver adds non-parametric planning information during single-step decision-making and constructs a probability distribution through the Softmax normalization method as the strategy of the neural solver. This non-parametric planning information can be used as a correction term in the decision-making process of the neural solver to guide it to explore more promising actions, thereby constructing a more efficient solution strategy.
[0092] Neural Plan designs a non-parametric planning strategy π constructed according to the environmental transition model Plan . In the planning strategy π PlanAmong them, for each potential action a at a specific time point t, the planning strategy constructs a probability distribution model based on the current state and the reward signal that the current action will return provided by the environment model to predict the next best action direction. Such planning not only provides predictive information about future possible states but also helps the model better evaluate the potential value of each action in the current state. Taking the path planning problem as an example, specifically, the planning strategy at time t predicts the action τ t is:
[0093] τ t ~π Plan (τ t |s t-1 ),
[0094]
[0095] where π Plan is the planning strategy, τ t is the decision-making action at time t in the solution trajectory τ of the neural solver, s t-1 is the decision-making state of the combinatorial optimization problem at the previous moment, which can be represented as the set of nodes visited last time and unselected nodes in the traveling salesman problem. Among them, dist is generated according to the problem instance and can be represented as distance information in the path planning problem. a is all selectable candidate actions.
[0096] Denote the planning depth as L, then there is:
[0097]
[0098] π Plan (a t |s t ,L)=0.
[0099] To effectively integrate this predictive planning information with the strategy of the neural solver, a balance parameter λ is introduced. When used as a supplementary correction for estimating the adaptation degree of the nodes of the neural solver, this parameter adjusts the weight between the planning information and the strategy of the neural solver, thereby affecting the probability distribution of the final action selection:
[0100]
[0101] By adjusting the value of λ, the solver's tendency to rely on predictive planning information and its own strategy can be controlled to find the best balance. In practical applications, a higher λ value means that the model relies more on planning information when making decisions, which usually helps to effectively explore in unknown or complex environments; while a lower λ value means that the model relies more on its own strategy trained through reinforcement learning, which is more effective when the environment is relatively simple or the model strategy is already relatively mature.
[0102] 2. Diversity Group Exploration Module: Existing neural solvers usually adopt an encoder-decoder architecture. The present invention further promotes the integration of planning information and strategy diversity (such as Figure 3 Specifically, each decoder has the same neural network architecture, but each decoder is combined with a different planning information strength control parameter λ to ensure effective guidance of planning information while promoting strategy diversification of the decoder group. On this basis, the present invention additionally designs a group collaborative baseline, and uses the average of the solutions generated by all decoders as the evaluation baseline in the reinforcement learning training method to guide the overall exploration direction. As a means of information interaction, the group collaborative baseline can enable individuals with poor strategies to learn from individuals with better strategies, thereby alleviating the phenomenon of ineffective exploration in the group.
[0103] The present invention extends the planning-guided approach through a swarm collaboration framework, whereby a single neural solver is scaled up to the swarm scale, with each agent utilizing planning information conditioned by its unique parameters. This setup is based on the idea that varying degrees of planning influence can foster a range of intrinsic strategies among agents, thereby enhancing the overall exploration capability.
[0104] Specifically, for agent i in the group, its strategy can be expressed as:
[0105]
[0106] where λ i Sampling is done from a uniform distribution U(0,10). λ constructs different intrinsic strategies in each individual by changing the degree of influence of the planning information intensity on each agent. Due to the difference in their intrinsic prior information, when faced with the same decision state, the estimation of the same action will still be different. Such differences correspond to the different behaviors of the agents. In the exploration theory of reinforcement learning, multiple different estimates of unseen environmental states will lead to different actions. The more diverse the estimates, the higher the exploration ability of the current agent. However, the simple strategy diversity cannot guarantee the effectiveness of its strategy exploration in the huge action space of the combinatorial optimization problem. The integration of diverse planning information takes into account both strategy diversity and effective exploration.
[0107] While maintaining policy diversity to promote exploration, excessive diversity may still lead to low search efficiency. If there are significant performance differences in the population, the exploration of agents with poor performance will not contribute to the overall performance improvement, but instead cause waste of computing resources. Based on this, the present invention designs a cooperative baseline, in which all solutions of the same instance are stored in a shared buffer, and the average quality of these solutions is used as the baseline for model update. By utilizing the shared buffer, different agents with inherent preference differences can solve the same problem, thereby improving data utilization and generating more effective state encodings. This enables individuals with poor performance to benefit from the exploration information provided by individuals with better performance during the update process, helping to reduce performance differences in the population and minimize the phenomenon of ineffective exploration. Specifically, in the reinforcement learning training method, the present invention designs a shared cooperative baseline:
[0108]
[0109] Denote the population policy set Then the overall model optimization objective can be written as:
[0110]
[0111] 3. Coarse-grained ternary search method module: When directly applying a neural solver pre-trained on a fixed-distribution training set to problems of other scales or distributions, due to the inconsistency of problem scales, existing neural solvers cannot be adapted to fine-grained decision-making processes (such as Figure 2 shown). The present invention optimizes the approximate unimodal characteristic presented between the planning information intensity and the generalization ability of the neural solver (such as Figure 4 shown), and designs a coarse-grained ternary search method to quickly find the optimal λ value range. This type of search method inherits the characteristics of the ternary search for efficiently processing unimodal functions, and uses Gaussian noise as the perturbation of the middle point. Unlike traditional ternary search methods that only evaluate through individual values, it judges the area around this point, and constructs an interval with Gaussian noise as the estimated value of this point, and further uses the ternary search method to gradually narrow the search range. This type of method can avoid the fine-tuning process of the neural solver and quickly improve the generalization ability of the neural solver.
[0112] By studying the impact of planning information intensity on generalization solution performance, the present invention discovers an approximate unimodal relationship between the hyperparameter λ that controls the planning information intensity and the solution quality. However, this unimodal property is interfered by the complex combinatorial optimization environment, and misleading local optima still hinder the utilization of its property. Although the performance can be improved by finely tuning λ, this inefficient search process incurs excessive training costs. Therefore, the present invention proposes a coarse-grained ternary search method to effectively utilize its approximate unimodal characteristics and quickly lock in the approximate optimal planning intensity.
[0113] Different from the conventional ternary search method, the coarse-grained ternary search does not rely on a continuous and narrow search range and can efficiently avoid the strict unimodal property assumption required by the conventional ternary search method through perturbation. In each step of the search, by introducing perturbation information, the dependence on precise search is reduced and the interval for discovering the global optimum is accelerated. This perturbation avoids potential local optimum traps by randomly expanding the search interval to explore the global optimum solution. This search method inherits the computational efficiency of the traditional ternary search and can solve problems within the time complexity of O(logn). While maintaining effective search, it significantly reduces the computational amount required for parameter tuning. Its algorithm flow is as follows:
[0114]
[0115] Table 1 shows the algorithm flow of the Neural Plan method proposed by the present invention;
[0116] Table 2 shows the algorithm flow of the coarse-grained ternary search method proposed by the present invention;
[0117] Table 3 shows the exploration ability (Gap ID ) and cross-distribution generalization (Gap CD ) experimental performance comparison of the Neural Plan method (including Ours-AM and Ours-POMO) proposed by the present invention on small-scale (20, 50, 100) artificial datasets of TSP and CVRP with existing mainstream methods;
[0118] Table 4 shows the cross-scale generalization (Gap CS ) and simultaneous cross-scale and cross-distribution generalization (Gap CSD ) experimental performance comparison of the Neural Plan method (including Ours-AM and Ours-POMO) proposed by the present invention on large-scale (200, 500, 1000) artificial datasets of TSP and CVRP;
[0119] Table 5 shows the generalization performance comparison of the Neural Plan method (including Ours-AM and Ours-POMO) proposed by the present invention on real-world datasets of TSP and CVRP;
[0120] Table 6 shows the ablation experiments of the Neural Plan method proposed by the present invention for each functional module.
[0121] Table 1
[0122]
[0123] Table 2
[0124]
[0125] Table 3
[0126]
[0127] Table 4
[0128]
[0129] Table 5
[0130]
[0131] Table 6
[0132]
[0133] Application Example
[0134] In practical applications, the present invention can efficiently solve the batch Traveling Salesman Problem (TSP) in the express logistics field. For example, a large express company needs to handle multiple delivery tasks every day, and each task involves 50 to 100 customer points. Taking 5 distribution centers as an example, each center needs to plan the parcel delivery routes for 50 to 100 customer points. Traditional TSP methods are difficult to meet the real-time solution requirements of such large-scale batch tasks, while the neural network architecture of the present invention can optimize the routes of multiple delivery tasks simultaneously through its parallel computing ability.
[0135] In a specific scenario, the input data includes the coordinates of the nodes (distribution centers and customer points) in each distribution task and the generated distance matrix. The goal is to batch generate the optimal paths for all tasks through a neural solver to minimize the total distribution distance. For example, on a certain day, 5 distribution tasks need to be processed. The node coordinates of Task 1 include the distribution center (0, 0) and customer points (2, 3), (5, 7), etc. The node coordinates of Task 2 include the distribution center (10, 10) and customer points (12, 13), (15, 17), etc. Through the planning guidance and population exploration mechanism of the present invention, the model can parallelly output the optimal paths for each task. For example, the path of Task 1 is "0→1→2→...→0" with a total distance of 120 kilometers; the path of Task 2 is "0→1→2→...→0" with a total distance of 150 kilometers. This application verifies the efficiency and generalization ability of the present invention in large-scale and multi-task scenarios, providing reliable technical support for intelligent logistics scheduling.
[0136] In summary, the new technical solution constructs a Neural Plan architecture by introducing a planning information design and fusion module, a diversity population exploration module, and a key coarse-grained ternary search method. This model improves the policy exploration ability and generalization ability of the neural solver based on deep reinforcement learning, overcoming the challenges in the RL4CO field. By introducing planning information into the decision-making process of the neural solver, this solution is expected to bring significant progress to the RL4CO field. In the above specific application, the present invention accurately captures the spatial distribution characteristics of distribution nodes through the planning information guidance mechanism, and combines the parallel path generation ability of the diversity population exploration module, significantly improving the efficiency of multi-task collaborative solving; at the same time, the coarse-grained ternary search method uses Gaussian noise perturbation combined with the ternary search mechanism to quickly lock the optimal planning information intensity interval for each distribution task, optimizing the generalization performance of the neural solver in cross-scale and cross-distribution scenarios. Compared with traditional methods, the total distribution distance is on average shortened and the calculation time is reduced, fully verifying its efficiency, robustness and industrial application potential in complex logistics scenarios, providing implementable intelligent decision-making support for large-scale path optimization in the express delivery industry.
[0137] The embodiment of the present invention also provides a storage medium for storing a computer program, which when executed, at least executes the method as described above.
[0138] The embodiment of the present invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein, the processor is used to execute the computer program to at least execute the method as described above.
[0139] The embodiment of the present invention also provides a processor, which executes a computer program and at least executes the method as described above.
[0140] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include but not limited to these and any other suitable types of memories.
[0141] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0142] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0143] In addition, in each embodiment of the present invention, the various functional units can all be integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit; the above integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0144] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the aforementioned storage medium includes: various media such as removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0145] Alternatively, if the above integrated units are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. And the aforementioned storage medium includes: various media such as removable storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0146] The methods disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0147] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0148] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0149] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the technical field to which the present invention pertains, without departing from the concept of the present invention, several equivalent substitutions or obvious variations can be made, and as long as the performance or use is the same, they should all be regarded as falling within the protection scope of the present invention.
Claims
1. An efficient reinforcement learning method for solving batch traveling salesman problem in express parcel delivery, characterized in that: The following steps are involved: S1. Construct planning information design and fusion module: Design planning strategies based on the graph structure environment model of express delivery tasks. The graph structure environment model includes the node coordinates of the distribution center and the customer point and their distance matrix. Integrate non-parametric planning information into the decision-making process of the neural solver to guide the path exploration direction and improve the exploration efficiency of batch tasks. S2. Design a diverse group exploration module: adopt a shared encoder-multiple decoder architecture, give each decoder a diverse decoding strategy for multiple delivery tasks through different planning information strength control parameters, introduce a group collaboration baseline, store the delivery path solutions generated by all decoders through a shared buffer, and calculate their average delivery distance as a baseline to promote strategy diversity and improve the overall exploration capability of batch tasks; S3. Implement a coarse-grained three-part search method: Aiming at the approximate unimodal relationship between planning information intensity and generalization performance in multiple delivery tasks, Gaussian noise perturbation is combined with a three-part search mechanism to quickly lock in the optimal planning information intensity range for different delivery tasks, optimize the generalization ability of the neural solver in batch tasks, and minimize the total delivery distance.
2. The method according to claim 1, characterized in that In step S1, the generation of the planning strategy includes the following steps: Based on the graph structure environment model of combinatorial optimization problems, a distance information matrix is constructed, and the distance information between nodes is linearly normalized into a probability distribution to form a planning strategy; Generate a planning information matrix through iterative sampling, and incorporate the planning information matrix into the decision-making process of the neural solver as non-parametric guidance information; In the single-step decision of the neural solver, the planning information is fused with the strategy of the neural solver through Softmax normalization to generate the final action selection probability distribution to guide the decision-making process of the neural solver.
3. The method according to claim 1 or 2, characterized in that: In step S1, the fusion of the planning strategy adjusts the weight between the planning information and the neural solver strategy by introducing a balance parameter to control the tendency of the neural solver between relying on the planning information and its own strategy.
4. The method according to claim 1, characterized in that: In step S2, the construction of the diversity group exploration module includes the following steps: Adopting a shared encoder-multiple decoder architecture, each decoder has the same neural network structure, but is endowed with diverse intrinsic strategies through different planning information strength control parameters; A group collaborative baseline is introduced to store the solutions generated by all decoders through a shared buffer and calculate their average quality as a baseline to guide the overall exploration direction; Through the group collaborative baseline, the decoder with poor performance can learn from the decoder with better performance, reducing the invalid exploration phenomenon.
5. The method according to claim 1 or 4, characterized in that: In step S2, the group collaboration baseline is calculated as follows: Evaluate the solution generated by each decoder and calculate its cost; The solution costs of all decoders are averaged and used as a group collaboration baseline to update the neural solver's policy.
6. The method according to claim 1, characterized in that In step S3, the implementation of the coarse-grained three-division search method includes the following steps: Aiming at the approximate unimodal relationship between planning information strength and generalization performance, Gaussian noise perturbation is used to expand the search interval to avoid falling into the local optimal trap; The three-part search mechanism gradually narrows the search scope and quickly locks in the optimal planning information intensity interval; In each step of the search, Gaussian noise perturbation is introduced to reduce the reliance on precise search and accelerate the discovery of the global optimal solution.
7. The method according to claim 1 or 6, characterized in that: In step S3, the specific steps of the three-part search mechanism include: Divide the search interval into three equal parts and calculate the quality of the solution at the middle point; The middle point is expanded by Gaussian noise perturbation to evaluate the solution quality of the surrounding interval; According to the comparison results of solution quality, the search range is gradually narrowed until the optimal planning information intensity interval is found.
8. The method according to any one of claims 1 to 7, characterized in that: In step S3, the Gaussian noise disturbance is introduced as follows: In each search step, Gaussian noise perturbation is applied to the intermediate points to generate multiple perturbation points; By evaluating the solution quality of the disturbance point, the next search direction is determined to avoid falling into the local optimal solution.
9. The method according to any one of claims 1 to 8, characterized in that: In step S3, the termination condition of the coarse-grained three-part search method is: When the width of the search interval is less than a preset threshold, the search is stopped; Multiple planning information strength values are randomly sampled from the final search interval, and the value with the best solution quality is selected as the final optimal planning information strength.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.