Automatic algorithm design method for solving vehicle path planning problem by using partial reconstruction heuristic operator

Through the AutoSAF framework and deep reinforcement learning, combined with partial reconstruction heuristic operators and multiple sampling strategies, the problem of combining construction and perturbation neural combination optimization in the existing technology is solved, and efficient solution and quality improvement of vehicle path planning problems is achieved.

CN120046689APending Publication Date: 2025-05-27NORTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510046308.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively combine structural neural combination optimization and perturbation neural combination optimization to solve problems such as differences in design concepts, intelligent scheduling problems and method integration complexity when solving vehicle path planning problems.

Method used

Using partial reconstruction heuristic operators, through the AutoSAF framework and deep reinforcement learning, we intelligently select appropriate heuristic algorithms, combine reconstruction pools, improved pools and perturbation pools, design a variety of sampling strategies and adversarial transformer models, and optimize the reconstruction and selection of solutions.

Benefits of technology

It significantly improves the solution efficiency of vehicle path planning problems, avoids the problem of local optimal solutions, improves the quality and diversity of understanding, and can adaptively adjust strategies to adapt to different optimization tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046689A_ABST
    Figure CN120046689A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic algorithm design method for solving a vehicle path planning problem by using a partial reconstruction heuristic operator, which integrates a partial reconstruction heuristic method, is used for solving a complex combinatorial optimization problem, constructs an AutoSAF framework and defines a heuristic pool, and the heuristic pool comprises an improvement pool, a reconstruction pool and a disturbance pool. In order to realize strong learning and reasoning capabilities, a selected sub-path is subjected to solution reconstruction by using an adversarial converter model and the solution is updated according to an acceptance standard of a mountain climbing strategy, the adversarial converter model is a heavy encoder-decoder architecture, a bidirectional GAN training strategy is adopted in the optimization process, the quality and the cost effectiveness of a solution generated by a generator are improved, and the method is suitable for large-scale popularization and application. When complex optimization problems such as vehicle path planning and the like are solved, the performance is remarkably improved, and the efficiency problem of a manual design algorithm is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an automatic algorithm design method, and more particularly to an automatic algorithm design method for solving the vehicle routing problem by using a partial reconstruction heuristic operator. Background Art

[0002] In recent years, the application of automatic algorithm design in combinatorial optimization problems has achieved rapid development. COPs mainly involve finding optimal solutions in discrete solution spaces and are widely used in fields such as logistics, bioinformatics, and energy management. However, manually designing solution algorithms requires a large amount of domain knowledge and is extremely time-consuming, resulting in subjectivity and limitations in algorithm design. With the rapid development of deep reinforcement learning and neural network technologies, automatic algorithm design is gradually becoming a research direction with great potential.

[0003] In this field, neural combinatorial optimization, as a branch of automatic algorithm design, is divided into two major categories: constructive neural combinatorial optimization and perturbation neural combinatorial optimization. Although both methods have made significant progress respectively, they still have some key problems, making it so that no research has effectively combined the two.

[0004] The constructive neural combinatorial optimization (CNCO) method uses neural networks to construct solutions from scratch. Common strategies include supervised learning (SL) and reinforcement learning (RL). Although CNCO shows excellent performance in solving small-scale problems, it faces challenges such as high training data requirements, high learning costs, sparse rewards, and large memory requirements in large-scale problems. To solve these problems, researchers have proposed the heavy encoder and light decoder (HELD) and light encoder and heavy decoder (LEHD) models to improve the solution performance of large-scale VRPs.

[0005] The perturbation neural combinatorial optimization (PNCO) method improves solutions by perturbing and adjusting existing solutions. Typical methods include Learn to Improve (L2I) and the General Search Framework (GSF). Although PNCO can effectively improve the quality of solutions, they require additional computational time to predict algorithm components.

[0006] Although CNCO and PNCO each have their advantages in solving COPs, there is currently no research effectively combining the two. The following challenges highlight the difficulty of their combination:

[0007] (1) Differences in design concepts: CNCO focuses on constructing solutions from scratch, while PNCO pays more attention to optimizing existing solutions. This essential difference makes the integration of the two require a re-design of the algorithm framework;

[0008] (2) The problem of intelligent scheduling: How to intelligently switch between CNCO and PNCO according to the problem state and the solution stage is an unsolved research problem;

[0009] (3) Complexity of method fusion: The combination of two methods requires the design of an efficient unified framework, which not only avoids waste of computing resources but also gives full play to their respective advantages. Summary of the Invention

[0010] The object of the present invention is to provide an automatic algorithm design method for solving the vehicle routing problem by using a partial reconstruction heuristic operator, which can intelligently select a suitable heuristic algorithm, not only flexibly cope with different vehicle routing problems, solve the efficiency problem of manually designed algorithms, but also avoid the problem of local optimal solutions, improve the quality and diversity of solutions, and significantly improve the solution efficiency of the vehicle routing problem.

[0011] The present invention is realized through the following technical solutions:

[0012] An automatic algorithm design method for solving the vehicle routing problem by using a partial reconstruction heuristic operator, comprising the following steps:

[0013] S1: Initialize the solution space to construct the AutoSAF framework and define the heuristic pool, and the heuristic pool includes an improvement pool, a reconstruction pool, and a perturbation pool;

[0014] S2: Design a variety of sampling strategies for the reconstruction pool to generate candidate sub-paths;

[0015] S3: Define the automatic algorithm design problem as a deep reinforcement learning task and model it as a Markov decision process M = (S, A, π, r);

[0016] S4: At each time step t, select the most suitable heuristic method from the heuristic pool according to the current state information;

[0017] S5: For the selected sub-path, use the adversarial transformer model to reconstruct the solution and update the solution according to the acceptance criteria of the hill climbing strategy;

[0018] S6: After T times of iterative optimization, select the solution with the lowest cost among all iterations as the final solution.

[0019] Preferably, the sampling strategies in step S2 include:

[0020] (1) Randomly sample a specified number of nodes;

[0021] (2) Starting from the starting point of the sub-path, sequentially sample several nodes;

[0022] (3) Starting from the middle position of the sub-path, sequentially sample several nodes forward.

[0023] Furthermore, each CNCO heuristic in the reconstruction pool in step S2 consists of a sampler and a reconstructor. The former samples fragments from the complete solution, and the latter reconstructs the fragments and expects to improve them.

[0024] Preferably, the parameters of the Markov decision process M=(S, A, π, r) in step S3 are described as follows:

[0025] The state space S includes:

[0026] (1) Static states: fixed information such as node coordinates, vehicle capacity, and customer demands;

[0027] (2) Dynamic states: the current solution (such as the path arrangement) and the historical record of heuristic selections;

[0028] The action space A: the agent selects actions from the heuristic pool;

[0029] The policy network π: selects the optimal action according to the current state through a neural network model;

[0030] The reward function r is defined as follows:

[0031] (1) Quality reward: when a higher-quality solution is found, the reward is +1, otherwise -1;

[0032] (2) Baseline reward: when the current solution is better than the initial solution, the reward is +1, otherwise -1;

[0033] (3) Exploration reward: when a solution different from the previous one is found, the reward is +1, otherwise -1.

[0034] Preferably, when using the adversarial transformer model to reconstruct the solution in step S5, the adversarial transformer model includes a generator and two discriminators D1 and D2;

[0035] Among them, the generator models the relationship between nodes through a heavy encoder-decoder architecture and generates a solution π for a given problem instance s by learning a stochastic policy p(π│s);

[0036] The discriminator D1 evaluates the similarity between the new path and the high-quality solution. By comparing the generated solution with the known high-quality solutions, it guides the generator to optimize its generation strategy to ensure that the generated solution is closer to these high-quality solutions;

[0037] The discriminator D2 evaluates whether the cost of the new path is close to the cost of the true solution, ensuring that the generated solution is not only similar to the high-quality solution but also competitive in terms of cost, generating a solution with a cost close to that of the true efficient solution.

[0038] Furthermore, the specific training process of step S5 adopts a bidirectional GAN strategy, and the specific process is as follows:

[0039] (1) Forward training stage: When the cost L(π│s) of the solution generated by the generator HEHD is not equal to the historical best cost b(s), enter the forward training stage. At this time, the cost L(π│s) of the generated solution will be compared with the cost cf(s) of the solution obtained by the professional solver to guide the training of the generator and improve the generated solution.

[0040] (2) Reverse training stage: If the generated solution maintains the same cost L(π│s)=b(s) in at least T = 5 training cycles, it means that the current solution has reached a local optimum and the training has entered a stagnation stage. At this time, reverse training will be started. In the reverse training stage, the generated solution is regarded as a fake sample, and its cost is still expressed as L(π│s), while the real sample is generated by applying the 2-OPT algorithm to degrade the current solution to generate cb(s).

[0041] The present invention has the following beneficial effects compared with the prior art:

[0042] (1) In the present invention, the AutoSAF framework adopts an automatic algorithm design. Based on deep reinforcement learning and a partial reconstruction heuristic operator, by intelligently selecting appropriate heuristic algorithms, it reduces the dependence on expert knowledge, can more flexibly handle different vehicle routing problems, and achieves significant performance improvement when solving complex optimization problems such as vehicle routing. It solves the efficiency problem of manually designed algorithms. By using the partial reconstruction heuristic operator, it can not only locally optimize the existing solution, but also effectively explore in the global solution space, avoiding the problem of local optimal solutions, improving the quality and diversity of the solutions, significantly improving the solving efficiency of vehicle routing problems, and enabling the AutoSAF framework to adaptively adjust strategies and obtain better optimization effects in different optimization tasks.

[0043] (2) The CNCO method in the reconstruction pool of the AutoSAF framework can effectively handle complex solution spaces, and by intelligently selecting different heuristic methods, it avoids the computational bottleneck in solving large-scale problems, significantly improves the solving efficiency of the AutoSAF framework when dealing with large-scale problems, and reduces resource consumption, showing good scalability on large-scale data sets.

[0044] (3) Through the diverse combination of heuristic pools, including the reconstruction pool, improvement pool, and perturbation pool, the AutoSAF framework can flexibly select different heuristics, balance the quality and diversity of solutions simultaneously during the optimization process. In particular, the perturbation pool introduces large perturbations to help the algorithm jump out of local optimal solutions, expands the solution space, and improves the comprehensive performance of problem solving.

[0045] (4) In the present invention, a two-way GAN training strategy is proposed, enabling the generator to optimize the quality of the solution in both the forward training and reverse training directions, promoting the generator to improve the quality of the solution. And through the semi-supervised training method, the generator can generate solutions that are close to high-quality solutions and cost-effective, enabling the AutoSAF framework to effectively generate high-quality solutions that meet the problem constraints, while optimizing the cost and diversity of the solutions, and improving the ability and quality of solution generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a schematic structural diagram of the AutoSAF framework of the present invention;

[0047] Figure 2 is a schematic diagram of the HEHD model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Referring to Figure 1 as shown, an AutoSAF framework proposed by the present invention, in which the agent is located at the core position, responsible for selecting the most suitable heuristic method from the reconstruction pool, improvement pool and perturbation pool according to the environmental state (including problem-specific features, current solutions and heuristic usage history) by using the policy network to optimize the solution of the vehicle routing problem, and feeding back the obtained reward to the agent to update its policy network to optimize the future decision-making process.

[0049] The present invention will be further described in detail below with specific embodiments, which are explanations rather than limitations of the present invention.

[0050] Assume a CVRP instance:

[0051] Number of customers: 20 customers (nodes), numbered [1, 2,..., 20];

[0052] Number of vehicles: 2 vehicles, with a capacity of 10 units for each vehicle;

[0053] Customer demand: randomly generated from the interval [1, 9], and the demand of each node is Qi.

[0054] This embodiment provides an automatic algorithm design method for solving the vehicle routing problem by using a partial reconstruction heuristic operator, including the following steps:

[0055] S1: Define the problem and the environment

[0056] Define the CVRP instance parameters as the problem environment, define the automatic algorithm design problem as a deep reinforcement learning task, and model it as a Markov decision process M=(S, A, π, r). Construct the state space S, action space A, policy network π, and reward function r. The specific parameter descriptions are as follows:

[0057] The state space S includes:

[0058] (1) Static states: fixed information such as node coordinates, vehicle capacity, and customer demands;

[0059] (2) Dynamic states: the current solution (such as route arrangement) and the historical record of heuristic selections.

[0060] The action space A: The agent selects actions from the heuristic pool.

[0061] The policy network π: Through a neural network model, select the optimal action according to the current state.

[0062] The rules of the reward function r are as follows:

[0063] (1) Quality reward: When a higher-quality solution is found, the reward is +1, otherwise -1;

[0064] (2) Baseline reward: When the current solution is better than the initial solution, the reward is +1, otherwise -1;

[0065] (3) Exploration reward: When a solution different from the previous one is found, the reward is +1, otherwise -1.

[0066] S2: Initialize the AutoSAF framework

[0067] Set up a deep reinforcement learning (DRL) agent, initialize the policy network π, and define the heuristic pool, where the policy network is used to output the probability distribution of heuristic selections, and the heuristic pool includes an improvement pool, a reconstruction pool, and a perturbation pool.

[0068] Improvement pool: Heuristic codes h1~h27, including local search methods, mainly for locally fine-tuning the existing solutions, such as 2-opt, Relocate, Cross.

[0069] Reconstruction pool: Heuristic codes h28~h39, consisting of pure CNCO methods trained on small-scale VRPs, combined with a learned sampling strategy, used to reconstruct part or the whole of the solution, such as sub-path reconstruction based on the adversarial transformer model (ATM), and each CNCO heuristic in the reconstruction pool consists of a sampler and a reconstructor. The former samples fragments from the complete solution, and the latter reconstructs the fragments and expects to improve them.

[0070] Perturbation Pool: Heuristic codes h40 - h42, including large neighborhood search methods such as random swap and random permutation, avoid falling into local optima by introducing large - scale perturbations.

[0071] S3: The agent interacts with the CVRP environment

[0072] The agent starts to interact with the CVRP environment. First, it obtains the current state information, including dynamic states such as customer demands, current route arrangements, and heuristics used in history, and then according to the current state S t , inputs it into the policy network, and outputs the selection probability of each heuristic.

[0073] S4: Select heuristics

[0074] Different state information triggers different types of heuristic methods. According to the probability distribution output by the policy network, the ε - greedy strategy is used to select heuristics. With a probability of 1 - ε, the currently optimal heuristic is selected, and with a probability of ε, other heuristics are explored to increase the diversity of exploration.

[0075] Suppose the currently selected heuristic is a method in the reconstruction pool, such as the randomly sampled sub - route based on the adversarial transformer (ATM).

[0076] S5: Apply heuristics and receive feedback

[0077] Candidate sub - routes are selected from the current solution, and the heuristic optimization method in the reconstruction pool is applied. For the sampling strategy (the reconstruction pool supports multiple sampling methods) including randomly sampled sub - routes and sequential sampling, randomly sampled sub - routes are randomly selected from the current solution, while sequential sampling starts from the start point, mid - point, or end point of the sub - route and selects several node segments in sequence. For example, if the currently selected is a randomly sampled sub - route of 5 nodes, the sub - route is: 3→5→7→9→6.

[0078] The sampled sub-paths are then optimized using the Adversarial Transformer Model (ATM). First, the generator (HEHD model) reorders the sub-paths to generate a new node order such as 3→7→9→6→5. That is, its task is to generate new solutions, model the relationships between nodes through a heavy encoder-decoder architecture, reduce the local cost of the sub-paths, and at the same time ensure the constraints of node requirements. A solution π is generated for a given problem instance s by learning a stochastic policy p(π│s). Then, the discriminators (D1 and D2) are used. Discriminator D1 is a solution similarity discriminator, which is used to evaluate the similarity between the new path and the high-quality solution. By comparing the generated solution with the known high-quality solutions, it guides the generator to optimize its generation strategy to ensure that the generated solution is closer to these high-quality solutions. Discriminator D2 is a cost discriminator, which is used to evaluate whether the cost of the new path is close to the cost of the true solution, ensuring that the generated solution is not only similar to the high-quality solution but also competitive in terms of cost, generating a solution with a cost close to that of the true efficient solution. And the two discriminators (D1 and D2) adopt a semi-supervised training method and share the same structure. Through such a training strategy, the generator can gradually improve and generate solutions that are close to high-quality solutions and are cost-effective. For the training process of this ATM optimization, a Bidirectional GAN (BGAN) strategy is adopted to improve the quality and cost-effectiveness of the solutions generated by the generator. The specific process is as follows:

[0079] (1) Forward training phase: When the cost L(π│s) of the solution generated by the generator HEHD is not equal to the historical best cost b(s), it enters the forward training phase. At this time, the cost L(π│s) of the generated solution will be compared with the cost cf(s) of the solution obtained by the professional solver to guide the training of the generator and improve the generated solution;

[0080] (2) Reverse training phase: If the generated solution maintains the same cost L(π│s) = b(s) for at least T = 5 training cycles, it means that the current solution has reached a local optimum and the training has entered a stagnant phase. At this time, reverse training will be started. In the reverse training phase, the generated solution is regarded as a fake sample, and its cost is still expressed as L(π│s), while the real sample is generated by applying the 2-OPT algorithm to degrade the current solution to generate cb(s).

[0081] The specific training parameters of the above training process include:

[0082] α: Represents the difference coefficient between the cost L(π│s) of the current generated solution and the cost b(s) of the historical best solution, controlling the influence on the historical solution during training;

[0083] β: represents the difference coefficient between the cost L(π│s) of the currently generated solution and the cost cf(s) of the solution obtained by the professional solver, and controls the learning of the generator from the solution of the professional solver;

[0084] γ: represents the difference coefficient between the cost L(π│s) of the currently generated solution and the cost cb(s) of the solution deteriorated using the 2-OPT algorithm, and controls how the generator processes and optimizes the solution based on the 2-OPT strategy.

[0085] Then calculate the feedback reward, evaluate the difference between the total cost of the reconstructed path and the total cost of the original path. If the cost decreases, the reward is +1. Considering the computing time required to apply the heuristic, if the time is too long, the penalty is -1. And if the new path is different from the historical path, the reward is +1, otherwise the penalty is -1.

[0086] As Figure 2 shown, the above generator HEHD model is a neural network model based on the Tranformer architecture, specifically a heavy encoder-decoder architecture, that is, it consists of two parts: an encoder and a decoder, mainly used to process sequence data. The encoder is used to extract the feature vectors of nodes, and the encoder includes seven attention layers, each attention layer consists of two sub-layers, namely the multi-head attention layer (MHA) and the fully connected feed-forward layer (FF). Skip connections and batch normalization (BN) are used between these two sub-layers to update the feature vectors. The decoder then uses the embeddings, context vectors generated by the encoder, and the mask of the determined node sequence to perform the autoregressive output of the solution sequence π. For the query qi of node i and the key kj of node j, use the formula to calculate to determine the importance of node i relative to node j. During this process, use the Swish function to limit the result within the range of [-C, C] (where C = 10), and finally calculate the output probability vector p through the Softmax function.

[0087]

[0088] S6: Update the policy network

[0089] Update the parameters of the policy network according to the reward value, optimize the policy π, and use the gradient ascent method to adjust the network weights to increase the probability of selecting high-reward heuristics in the future, that is, if the new solution is better than the current solution, accept the new solution and update the current state, otherwise retain the original solution.

[0090] S7: Verification and evaluation

[0091] Test the effectiveness of the current strategy on CVRP instances of different scales, including comparing the routes generated by AutoSAF with those generated by baseline methods (such as LKH3 and NeuRewriter), recording the route cost, running time, and solution efficiency of each method, and verifying the performance of AutoSAF in terms of quality, efficiency, and cost-effectiveness.

[0092] S8: Iterative optimization

[0093] Repeat S3 to S7, continuously select heuristics for optimization. After T = 4000 iterations, the agent policy network gradually converges. Record the optimal solutions in all iterations and finally select the solution with the lowest route cost as the result. Assume that the total route cost before route optimization is 230. After optimizing the route through random sampling optimization in the reconstruction pool, the total route cost drops to 210, that is, it decreases by 20 units. This optimization process proves the effectiveness of the heuristic method in the reconstruction pool in local route reconstruction and can significantly reduce the route cost.

[0094] For example, the parameters of AutoSAF in Table 1 below, and its experimental results are shown in Tables 2 and 3.

[0095] Table 1 Parameters of AutoSAF

[0096]

[0097] Table 2 Experimental results of AutoSAF in small-scale random trials of CVRP

[0098]

[0099] Table 3 Experimental results of AutoSAF in large-scale CVRP

[0100]

[0101] The above content is a further detailed description of the present invention in combination with specific embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these embodiments. For those of ordinary skill in the technical field of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.

Claims

1. An automatic algorithm design method for solving vehicle path planning problems using a partial reconstruction heuristic operator, characterized in that: The following steps are involved: S1: Initialize the solution space to build the AutoSAF framework and define the heuristic pool, which includes the improvement pool, reconstruction pool and perturbation pool; S2: Design multiple sampling strategies for the reconstruction pool to generate candidate subpaths; S3: Define the automatic algorithm design problem as a deep reinforcement learning task and model it as a Markov decision process M = (S, A, π, r); S4: At each time step t, the most appropriate heuristic method is selected from the heuristic pool according to the current state information; S5: for the selected subpath, the adversarial transformer model is used to reconstruct the solution and the solution is updated according to the acceptance criteria of the hill climbing strategy; S6: After T iterations of optimization, the solution with the lowest cost among all iterations is selected as the final solution.

2. The automatic algorithm design method for solving vehicle path planning problems using a partial reconstruction heuristic operator according to claim 1, characterized in that: The sampling strategy in step S2 includes: (1) Randomly sample a specified number of nodes; (2) Starting from the starting point of the subpath, sequentially sample several nodes; (3) Starting from the middle of the subpath, sample several nodes sequentially.

3. The automatic algorithm design method for solving vehicle path planning problems using a partial reconstruction heuristic operator according to claim 2, characterized in that: Each CNCO heuristic in the reconstruction pool in step S2 consists of a sampler that samples fragments from the complete solution and a reconstructor that reconstructs the fragments in the hope of improving them.

4. The automatic algorithm design method for solving vehicle path planning problems using a partial reconstruction heuristic operator according to claim 1, characterized in that: The parameters of the Markov decision process M=(S, A, π, r) in step S3 are described as follows: The state space S consists of: (1) Static state: fixed information such as node coordinates, vehicle capacity, and customer demand; (2) Dynamic state: the current solution (e.g., path arrangement) and the history of heuristic choices; Action space A: The agent selects actions from a pool of heuristics; Policy network π: selects the optimal action based on the current state through a neural network model; The reward function r rule is: (1) Quality reward: When a higher quality solution is found, the reward is +1, otherwise -1; (2) Baseline reward: When the current solution is better than the initial solution, the reward is +1, otherwise -1; (3) Exploration reward: When a solution different from the previous one is found, the reward is +1, otherwise -1.

5. The automatic algorithm design method for solving vehicle path planning problems using a partial reconstruction heuristic operator according to claim 1, characterized in that: In the step S5, when the adversarial transformer model is used to reconstruct the solution, the adversarial transformer model includes a generator and two discriminators D1 and D2; The generator models the relationship between nodes through a heavy encoder-decoder architecture and generates a solution π for a given problem instance s by learning a random strategy p(π│s); The discriminator D1 evaluates the similarity of the new path with the high-quality solutions. By comparing the generated solutions with the known high-quality solutions, it guides the generator to optimize its generation strategy to ensure that the generated solutions are closer to these high-quality solutions. Discriminator D2 evaluates whether the cost of the new path is close to the cost of the true solution, ensuring that the generated solution is not only similar to the high-quality solution but also competitive in cost, generating a solution with a cost close to the true efficient solution.

6. The automatic algorithm design method for solving vehicle path planning problems using a partial reconstruction heuristic operator according to claim 5, characterized in that: The specific training process of step S5 adopts a bidirectional GAN ​​strategy, and the specific process is as follows: (1) Forward training phase: When the cost L(π│s) of the solution generated by the generator HEHD is not equal to the historical best cost b(s), the forward training phase begins. At this time, the generated solution cost L(π│s) will be compared with the solution cost cf(s) obtained by the professional solver to guide the training of the generator and improve the generated solution; (2) Reverse training phase: If the generated solution maintains the same cost L(π│s) = b(s) for at least T = 5 training cycles, it means that the current solution has reached the local optimum and the training has entered a stagnant phase. At this time, reverse training will be started. In the reverse training phase, the generated solution is regarded as a false sample, and its cost is still expressed as L(π│s). The real sample is degraded by applying the 2-OPT algorithm to the current solution to generate cb(s).