TSP path combination optimization method and system based on reinforcement learning

By integrating supervised learning and reinforcement learning in TSP path combination optimization, combined with direct preference optimization and anisotropic graph neural network, the problem of low solution efficiency in large-scale combination optimization problems is solved, and efficient and high-quality solution effects are achieved.

CN119939107APending Publication Date: 2025-05-06CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510016231.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When the prior art deals with large-scale combination optimization problems, the solution efficiency is low, making it difficult to achieve efficient and high-quality solutions.

Method used

A TSP path combination optimization method that integrates supervised learning and reinforcement learning is proposed. By introducing direct preference optimization into the diffusion model, combined with anisotropic graph neural network, the generalization ability and adaptability of the model are improved.

Benefits of technology

It significantly improves the generation quality and generalization ability of understanding, solves the problems of poor convergence and stability of traditional methods on large-scale complex problems, and improves the quality of solution efficiency and reconciliation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939107A_ABST
    Figure CN119939107A_ABST
Patent Text Reader

Abstract

The invention provides a TSP path combinatorial optimization method and system based on reinforcement learning, and the method comprises the steps: constructing an anisotropic graph neural network as a backbone network of a diffusion model through the combination of solution distribution learning and combinatorial optimization target learning, capturing a complex relation in graph structure data through the representation capability of the network, and carrying out the optimization of a TSP path. And probability distribution is modeled by using a diffusion model in a single-to-Markov forward process. In addition, two kinds of accelerated sampling methods DDIM and DPM-solver of diffusion models are introduced, the sampling process of denoising is accelerated, and the training efficiency is improved. According to the method, direct preference optimization can be introduced into a diffusion model to further provide preference-guided combinatorial optimization (PGCO), so that the generalization ability and adaptability of a traveling salesman problem (TSP) solving model are improved, and a more efficient and high-quality model for solving large-scale combinatorial optimization is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the intersection field of artificial intelligence and operations research, and relates to a TSP path combination optimization method and system based on reinforcement learning. Background Art

[0002] Combinatorial optimization problems play a key role in fields such as operations research, computer science, and transportation planning, and are crucial for improving efficiency and optimizing resource allocation. However, due to its NP (non-deterministic polynomial time)-hardness, solving large-scale problems is extremely challenging. Existing traditional solution methods include exact algorithms and heuristic algorithms. Although the former can obtain the global optimal solution, the computational complexity increases exponentially with the problem size, and although the latter can find a better solution in a limited time, it cannot guarantee optimality.

[0003] At present, machine learning is widely used to solve combinatorial optimization problems. Machine learning shows its potential to deal with complex problems through data-driven automatic learning heuristic methods. However, these methods face challenges in practical applications: supervised learning methods are limited by the diversity and complexity of training data, and it is difficult to adapt to test instances that have not appeared before, and high-quality supervised data for combinatorial optimization problems is difficult to obtain; although reinforcement learning can directly optimize the objectives of combinatorial optimization problems by learning the optimal strategy through interaction with the environment, the training time is long, the computing resources are consumed a lot, and there are convergence and stability problems in large-scale problems. Although deep learning and reinforcement learning have made certain progress, the solution efficiency of existing methods is still low when dealing with large-scale instances. Therefore, how to apply deep learning to achieve efficient and high-quality solutions to large-scale combinatorial problems is still a technical problem that needs to be solved urgently. Summary of the invention

[0004] In view of the above problems, the present invention proposes a TSP path combinatorial optimization method and system that can integrate supervised learning and reinforcement learning. By introducing direct preference optimization into the diffusion model, preference-guided combinatorial optimization (PGCO) is proposed, thereby improving the generalization ability and adaptability of the model for solving the Traveling Salesman Problem (TSP), and providing a more efficient and high-quality model for solving large-scale combinatorial optimization.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A TSP path combination optimization method and system based on reinforcement learning, the key technology of which is: by combining solution distribution learning with combinatorial optimization target learning, an anisotropic graph neural network is constructed as the backbone network of the diffusion model, and its representation ability is used to capture the complex relationship in the graph structure data, and the diffusion model single-turn Markov forward process is used to model the probability distribution. In addition, the present invention also introduces two accelerated sampling methods of the diffusion model, DDIM and DPM-solver, to accelerate the denoising sampling process and improve the training efficiency.

[0007] A TSP path combination optimization method and system based on reinforcement learning specifically includes the following steps:

[0008] S1: Problem modeling. For a TSP combinatorial optimization problem instance G = (V, E, d), where V = {v1, v2, ..., v n} represents the set of cities, E represents the set of edges between cities, d i represents the distance of the ith edge. G ={0,1} N Defined as the space of candidate solutions {x} for problem instance G, Defined as the solution x∈X G The objective function is:

[0009]

[0010] Where cost(·) is the cost function of the candidate solution, that is, the path length in the TSP problem, and valid(·) is the penalty term, which returns a penalty value of +∞ for an invalid solution and 0 for a valid solution. i is the indicator vector of the i-th edge.

[0011] S2: Conditional distribution representation problem modeling. The present invention defines the conditional distribution p θ (x|G) represents the probability of obtaining the corresponding problem solution x for the instance problem G given the neural network parameters θ, so that the combinatorial optimization problem can be solved by minimizing the expected cost through reinforcement learning. The corresponding expected cost is:

[0012]

[0013] The goal of supervised learning is transformed into the solution distribution p θ 's seeking.

[0014] S3: Input processing of anisotropic graph neural network. An anisotropic graph neural network (AGNN) is constructed as the backbone network of the diffusion model to model and embed the graph structure of the instance G. Given a node vector Weighted edge vector Diffusion model denoising time step t∈{τ1,...,τ M}, where N represents the number of nodes in the graph and E represents the number of edges. The embedding of node, edge and time features are defined as:

[0015]

[0016] Among them, T a is a fixed hyperparameter used to adjust the frequency range. is the multi-frequency encoding of the characteristic component j of node i. T is a large number (can be selected as 10000), and concat(·) represents the concatenation operation.

[0017] The linear transformation of the features is:

[0018]

[0019] in, is the temporal feature embedding dimension. The embedding input transformation vector e is a weighted adjacency matrix that represents the distance between different nodes. 0 As calculated above. are model weights respectively, SiLU(·) is the activation function, and the formula is:

[0020] SiLU(x)=x·σ(x)

[0021]

[0022] S4: graph convolution layer of anisotropic graph neural network. The feature update formula of the lth layer of graph convolution in the present invention is:

[0023]

[0024] in denote the node feature vector and edge feature vector of the lth layer respectively, GN(·) is the normalization function, μ and σ are the mean and standard deviation respectively. κ is a numerical stability constant to avoid the denominator being 0 and is set to 10-6. are the model weights, Represents a dense attention map. For the TSP problem, the present invention converts the time step feature t 0 Convolutional features with edges After aggregation, the redefined edge features are updated as follows:

[0025]

[0026] in, is the transformation matrix of time embedding.

[0027] S4: Output layer of anisotropic graph neural network. Through multi-layer convolution of node and edge features, the predicted edge probability distribution is finally obtained through normalization and Softmax operation:

[0028]

[0029] Among them, L is the number of network layers, and norm(·) is the feature normalization operation.

[0030] S5: Supervised learning of diffusion model on AGNN to solve the distribution p θ The forward diffusion takes the optimal solution of the optimization problem, represented as x0, and gradually adds noise to generate the sequence x1, x2, ..., x T , and its recursive formula is:

[0031]

[0032] where β t is the noise coefficient at each time step t. Through multi-step diffusion, the original solution gradually degenerates into random noise of standard Gaussian distribution.

[0033] S6: The reverse denoising goal of the diffusion model on AGNN is to go from the state of adding noise Gradually recover the solution x0 of the optimization problem. This process is determined by the conditional probability distribution p θ (x t-1 |x t , G) control, the formula is:

[0034]

[0035] Among them, μ θ (·) is the conditional mean function, Σ θ (·) is the conditional covariance matrix, which is usually assumed to be a diagonal matrix.

[0036] Using AGNN to capture graph structure information, the present invention converts the conditional mean μ of the denoising process into θ (·) and covariance Σ θ (·) is further defined as:

[0037] μ θ (x t ,G,t)=AGNN θ (x t ,G,t)

[0038] ∑ θ (x t ,G,t)=σ 2 (t)I

[0039] Among them, σ 2(t) is the noise intensity fixed at time step t. The solution generated by diffusion is passed through the loss function To train:

[0040]

[0041] Among them, ∈ is Gaussian noise, ∈ θ is the noise predicted by the model. This loss function is the loss function of supervised learning.

[0042] S7: This paper proposes a preference-guided combinatorial optimization (PCGO) method, which conceptualizes the denoising process of the diffusion model as a multi-step Markov decision process (MDP), uses human preference data to correspond to the combinatorial optimization problem, uses the quality of the generated solution to optimize the strategy, and uses direct preference optimization (DPO) to fine-tune the diffusion model to maximize the expected return. The state of the denoising process is defined as:

[0043]

[0044] Represents the current graph structure, time step, and solution status.

[0045] The action of the denoising process is defined as:

[0046]

[0047] Represents the next step solution generated from the current state.

[0048] Therefore, for reinforcement learning, the following optimization objectives are given:

[0049]

[0050] Among them, G is a heterogeneous graph, t is the time step, and x t is the state at time step t, σ ω ,σ l It represents the weight or category distribution between weak preference and strong preference in preference constraint. logρ(·) is the logarithm of partial score, and β is a hyperparameter that controls the strength of preference constraint and weighs the influence between two preferences. For the first part of preference score:

[0051]

[0052] represents the logarithmic probability ratio of the solution predicted by the model under weak preference. θ (·) is the predicted probability based on the model parameters θ. ref (·) is the probability of the reference distribution, which usually represents the true distribution or the distribution of the benchmark model. is the final solution based on weak preference, It is the intermediate state of weak preference at time step t, and the weak preference corresponds to the combinatorial optimization problem, that is, the solution of low quality / the solution that better satisfies the combinatorial optimization constraints under the same quality.

[0053] For the second part of the preference score:

[0054]

[0055] represents the log-probability ratio of the solution predicted by the model under strong preference. is the final solution based on strong preference, It is the strong preference intermediate state at time step t, and the strong preference corresponds to the combinatorial optimization problem, which is a high-quality solution.

[0056] S8: Joint training of supervised learning and reinforcement learning for PGCO. Through S6 and S7, the losses of supervised learning and reinforcement learning on parameter θ have been obtained respectively, which are and Using adaptive weighting based on gradient norm, the total loss is as follows:

[0057]

[0058] in, This dynamic weighting mechanism effectively avoids the dominant effect of a certain loss term during training by balancing the gradient norm of each loss term, significantly stabilizing model training.

[0059] S9: Fast reasoning of PGCO based on DDIM and DPM-solver accelerated sampling. After the model is trained through PGCO, the solution to the combinatorial optimization problem can be obtained by stepwise denoising. DDIM reduces the sampling steps of the diffusion model through a non-Markov denoising process while retaining the high quality of the generated solution. n+k to n DDIM PF-ODE solver DDIM Then:

[0060]

[0061] in, is the input diffusion noise, α and σ are diffusion coefficients, is the noise prediction model, and c is the conditional information.

[0062] For DPM-solver, this patent only considers the case of order 2, the ordinary differential equation solver Ψ DPM-Solver The detailed formula is:

[0063]

[0064] in, is the Log-SNR difference, and is the noise prediction value of the intermediate time step, and the calculation formula is:

[0065]

[0066] In this patented method, the advantages of DDIM and DPM-Solver are combined to accelerate sampling in stages. In the initial stage, DDIM is used for large-step deterministic sampling to quickly generate intermediate noise solutions; in the fine stage, DPM-Solver is used for second-order optimization sampling to further improve the accuracy and quality of the solution.

[0067] The beneficial effects of the present invention are: by introducing direct preference optimization into the diffusion model, and combining the construction of the solution with anisotropic graph neural network, the solution of TSP combinatorial optimization problem is solved by combining supervised learning with reinforcement learning. The generation quality and generalization ability of the solution are significantly improved. The present invention conceptualizes the denoising process of the diffusion model on the anisotropic graph neural network as a multi-step Markov decision process, and adopts a human preference data optimization strategy to avoid the dilemma of poor generalization of traditional methods due to the diversity and complexity of solution distribution. In addition, the present invention utilizes improved DDIM (Denoising Diffusion Implicit Models) and DPM-solver accelerated sampling technology to improve the training speed and reasoning efficiency of the model, solving the problems of weak convergence and stability of traditional reinforcement learning on large-scale complex problems, as well as the problems of low solution efficiency and excessive resource consumption of existing deep learning in this case. The present invention can efficiently utilize all data information under the same TSP problem scale, providing better solutions and more efficient solution performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0069] Figure 1 It is a forward diffusion framework diagram of the present invention;

[0070] Figure 2 It is the overall framework diagram of the present invention;

[0071] Figure 3 This is the solution process of the present invention on the TSP-50 / 100 / 500 problem;

[0072] Figure 4 It is the comparison result of the present invention with other methods on TSP-50 and TSP-100;

[0073] Figure 5 It is the result of the present invention on a large-scale TSPLIB benchmark instance of TSP200-1000 nodes;

[0074] Figure 6 is the resource requirement for training of the present invention at different problem scales; DETAILED DESCRIPTION

[0075] The specific implementation manner and working principle of the present invention are further described in detail below with reference to the accompanying drawings.

[0076] In this embodiment, only the city ID and city coordinates of the TSP problem need to be input to obtain the corresponding number of nodes. Figure 2 It can be seen that a TSP path combination optimization method and system based on reinforcement learning actually includes two parts: forward diffusion and backward denoising. In the forward diffusion part, two initial solutions will be generated. The input layer, graph convolution layer and output layer of AGNN will integrate information at all levels and learn with forward noise that supports diffusion. Under the constraint of loss, DPO is introduced to compare the advantages and disadvantages of the two solutions to obtain a neural network that can accurately learn noise. Finally, backward denoising is performed to obtain the optimal solution for TSP. This patent proposes a training method of supervised learning combined with DPO. First, supervised learning is used, that is, after the diffusion model training based on AGNN can obtain a feasible solution, DPO is used alone for fine-tuning training to optimize the quality of the solution.

[0077] Further, the following situation is used as an example to illustrate:

[0078] Assume that the data set is a TSP city list with M nodes, and the content is:

[0079] City = {(cityid i ,coord i )|i=1,2,...,M}

[0080] where cityid i is the city number, coord i is the two-dimensional coordinate vector of city i, in the form of (x i ,y i ).

[0081] First, two initial solutions are randomly generated and the forward diffusion part begins. Figure 1 A detailed framework diagram is shown, and the specific implementation steps are as follows:

[0082] S1: Problem modeling. For a TSP combinatorial optimization problem instance G = (V, E, d), where V = {v1, v2, ..., v n} represents the set of cities, E represents the set of edges between cities, d irepresents the distance of the ith edge. G ={0,1} N Defined as the space of candidate solutions {x} for problem instance G, Defined as the solution x∈X G The objective function is:

[0083]

[0084] Where cost(·) is the cost function of the candidate solution, that is, the path length in the TSP problem, and valid(·) is the penalty term, which returns a penalty value of +∞ for an invalid solution and 0 for a valid solution. i is the indicator vector of the i-th edge.

[0085] S2: Conditional distribution representation problem modeling. The present invention defines the conditional distribution p θ (x|G) represents the probability of obtaining the corresponding problem solution x for the instance problem G given the neural network parameters θ, so that the combinatorial optimization problem can be solved by minimizing the expected cost through reinforcement learning. The corresponding expected cost is:

[0086]

[0087] The goal of supervised learning is transformed into the solution distribution p θ 's seeking.

[0088] S3: Input processing of anisotropic graph neural network. An anisotropic graph neural network (AGNN) is constructed as the backbone network of the diffusion model to model and embed the graph structure of the instance G. Given a node vector Weighted edge vector Diffusion model denoising time step t∈{τ1,...,τ M}, where N represents the number of nodes in the graph and E represents the number of edges. The embedding of node, edge and time features are defined as:

[0089]

[0090] Among them, T a is a fixed hyperparameter used to adjust the frequency range. is the multi-frequency encoding of the characteristic component j of node i. T is a large number (can be selected as 10000), and concat(·) represents the concatenation operation.

[0091] The linear transformation of the features is:

[0092]

[0093] in, is the temporal feature embedding dimension. The embedding input transformation vector e is a weighted adjacency matrix that represents the distance between different nodes. 0 As calculated above. are model weights respectively, SiLU(·) is the activation function, and the formula is:

[0094] SiLU(x)=x·σ(x)

[0095]

[0096] S4: graph convolution layer of anisotropic graph neural network. The feature update formula of the lth layer of graph convolution in the present invention is:

[0097]

[0098] in denote the node feature vector and edge feature vector of the lth layer respectively, GN(·) is the normalization function, μ and σ are the mean and standard deviation respectively. κ is a numerical stability constant to avoid the denominator being 0 and is set to 10-6. are the model weights, Represents a dense attention map. For the TSP problem, the present invention combines the time step feature t0 with the edge convolution feature After aggregation, the redefined edge features are updated as follows:

[0099]

[0100] in, is the transformation matrix of time embedding.

[0101] S4: Output layer of anisotropic graph neural network. Through multi-layer convolution of node and edge features, the predicted edge probability distribution is finally obtained through normalization and Softmax operation:

[0102]

[0103] Among them, L is the number of network layers, and norm(·) is the feature normalization operation.

[0104] S5: Supervised learning of diffusion model on AGNN to solve the distribution p θ The forward diffusion takes the optimal solution of the optimization problem, represented as x0, and gradually adds noise to generate the sequence x1, x2, ..., x T , and its recursive formula is:

[0105]

[0106] where β tis the noise coefficient at each time step t. Through multi-step diffusion, the original solution gradually degenerates into random noise of standard Gaussian distribution.

[0107] S6: The reverse denoising goal of the diffusion model on AGNN is to go from the state of adding noise Gradually recover the solution x0 of the optimization problem. This process is determined by the conditional probability distribution p θ (x t-1 |x t , G) control, the formula is:

[0108]

[0109] Among them, μ θ (·) is the conditional mean function, ∑ θ (·) is the conditional covariance matrix, which is usually assumed to be a diagonal matrix.

[0110] Using AGNN to capture graph structure information, the present invention converts the conditional mean μ of the denoising process into θ (·) and covariance Σ0(·) are further defined as:

[0111] μ θ (x t ,G,t)=AGNN θ (x t ,G,t)

[0112] Σ θ (x t G,t)=σ 2 (t)I

[0113] Among them, σ 2 (t) is the noise intensity fixed at time step t. The solution generated by diffusion is passed through the loss function To train:

[0114]

[0115] Among them, ∈ is Gaussian noise, ∈ θ is the noise predicted by the model. This loss function is the loss function of supervised learning.

[0116] S7: This paper proposes a preference-guided combinatorial optimization (PCGO) method, which conceptualizes the denoising process of the diffusion model as a multi-step Markov decision process (MDP), uses human preference data to correspond to the combinatorial optimization problem, uses the quality of the generated solution to optimize the strategy, and uses direct preference optimization (DPO) to fine-tune the diffusion model to maximize the expected return. The state of the denoising process is defined as:

[0117]

[0118] Represents the current graph structure, time step, and solution status.

[0119] The action of the denoising process is defined as:

[0120]

[0121] Represents the next step solution generated from the current state.

[0122] Therefore, for reinforcement learning, the following optimization objectives are given:

[0123]

[0124] Among them, G is a heterogeneous graph, t is the time step, and x t is the state at time step t, σ ω , σ l It represents the weight or category distribution between weak preference and strong preference in preference constraint. logρ(·) is the logarithm of partial score, and β is a hyperparameter that controls the strength of preference constraint and weighs the influence between two preferences. For the first part of preference score:

[0125]

[0126] represents the logarithmic probability ratio of the solution predicted by the model under weak preference. θ (·) is the predicted probability based on the model parameters θ. ref (·) is the probability of the reference distribution, which usually represents the true distribution or the distribution of the benchmark model. is the final solution based on weak preference, It is the intermediate state of weak preference at time step t, and the weak preference corresponds to the combinatorial optimization problem, that is, the solution of low quality / the solution that better satisfies the combinatorial optimization constraints under the same quality.

[0127] For the second part of the preference score:

[0128]

[0129] represents the log-probability ratio of the solution predicted by the model under strong preference. is the final solution based on strong preference, It is the strong preference intermediate state at time step t, and the strong preference corresponds to the combinatorial optimization problem, which is a high-quality solution.

[0130] S8: Joint training of supervised learning and reinforcement learning for PGCO. Through S6 and S7, the losses of supervised learning and reinforcement learning on parameter θ have been obtained respectively, which are and Using adaptive weighting based on gradient norm, the total loss is as follows:

[0131]

[0132] in, This dynamic weighting mechanism effectively avoids the dominant effect of a certain loss term during training by balancing the gradient norm of each loss term, significantly stabilizing model training.

[0133] S9: Fast reasoning of PGCO based on DDIM and DPM-solver accelerated sampling. After the model is trained through PGCO, the solution to the combinatorial optimization problem can be obtained by stepwise denoising. DDIM reduces the sampling steps of the diffusion model through a non-Markov denoising process while retaining the high quality of the generated solution. n+k to n DDIM PF-ODE solver DDIM Then:

[0134]

[0135] in, is the input diffusion noise, α and σ are diffusion coefficients, is the noise prediction model, and c is the conditional information.

[0136] For DPM-solver, this patent only considers the case of order 2, the ordinary differential equation solver Ψ DPM-Solever The detailed formula is:

[0137]

[0138] in, is the Log-SNR difference, and is the noise prediction value of the intermediate time step, and the calculation formula is:

[0139]

[0140] In this patented method, the advantages of DDIM and DPM-Solver are combined to accelerate sampling in stages. In the initial stage, DDIM is used for large-step deterministic sampling to quickly generate intermediate noise solutions; in the fine stage, DPM-Solver is used for second-order optimization sampling to further improve the accuracy and quality of the solution.

[0141] It should be pointed out that the above implementation description is not a limitation of the present invention, and the present invention is not limited to the above examples. Any changes, modifications, additions, modifications or substitutions made by ordinary technicians in this technical field within the indicated range of the present invention should be included in the scope of the claims of the present invention.

Claims

1. A TSP path combination optimization method and system based on reinforcement learning, characterized in that: The method specifically comprises the following steps: S1: Problem modeling. For a TSP combinatorial optimization problem instance G = (V, E, d), where V = {v1, v2, ..., v n } represents the set of cities, E represents the set of edges between cities, d i represents the distance of the ith edge. G ={0,1} N Defined as the space of candidate solutions {x} for problem instance G, Defined as the solution x∈X G The objective function is: Where cost(·) is the cost function of the candidate solution, that is, the path length in the TSP problem, and valid(·) is the penalty term, which returns a penalty value of +∞ for an invalid solution and 0 for a valid solution. i is the indicator vector of the i-th edge; S2: Conditional distribution representation problem modeling. The present invention defines the conditional distribution p θ (x|G) represents the probability of obtaining the corresponding problem solution x for the instance problem G given the neural network parameters θ, so that the combinatorial optimization problem can be solved by minimizing the expected cost through reinforcement learning. The corresponding expected cost is: The goal of supervised learning is transformed into the solution distribution p θ of seeking; S3: Input processing of anisotropic graph neural networks. Construction An anisotropic graph neural network (AGNN) is used as the backbone network of the diffusion model to model and embed the graph structure of the instance G. Given a node vector Weighted edge vector Diffusion model denoising time step t∈{τ1,...,τ M }, where N represents the number of nodes in the graph and E represents the number of edges. The embedding of node, edge and time features are defined as: in, T a is a fixed hyperparameter used to adjust the frequency range. is the multi-frequency encoding of the characteristic component j of node i. T is a large number (can be selected as 10000), and concat(·) represents the concatenation operation; The linear transformation of the features is: in, d t is the temporal feature embedding dimension. The embedding input transformation vector e is a weighted adjacency matrix that represents the distance between different nodes. 0 As calculated above. are model weights respectively, SiLU(·) is the activation function, and the formula is: SiLU(x)=x·σ(x) S4: graph convolution layer of anisotropic graph neural network. The feature update formula of the lth layer of graph convolution in the present invention is: in denote the node feature vector and edge feature vector of the lth layer respectively, GN(·) is the normalization function, μ and σ are the mean and standard deviation respectively. k is a numerical stability constant to avoid the denominator being 0 and is set to 10 -6 , are the model weights, Represents a dense attention map. For the TSP problem, the present invention converts the time step feature t 0 Convolutional features with edges After aggregation, the redefined edge features are updated as follows: in, is the transformation matrix of time embedding; S4: Output layer of anisotropic graph neural network. Through multi-layer convolution of node and edge features, the predicted edge probability distribution is finally obtained through normalization and Softmax operation: Where L is the number of network layers, norm(·) is the feature normalization operation; S5: Supervised learning of the diffusion model on AGNN to solve the solution distribution p θ The forward diffusion takes the optimal solution of the optimization problem, represented as x0, and gradually adds noise to generate the sequence x1, x2, ..., x T , and its recursive formula is: where β t is the noise coefficient at each time step t. Through multi-step diffusion, the original solution gradually degenerates into random noise of standard Gaussian distribution; S6: The reverse denoising goal of the diffusion model on AGNN is to go from the state of adding noise Gradually recover the solution x0 of the optimization problem. This process is determined by the conditional probability distribution p θ (x t-1 |x t , G) control, the formula is: Among them, μ θ (·) is the conditional mean function, ∑ θ (·) is the conditional covariance matrix, which is usually assumed to be a diagonal matrix; Using AGNN to capture graph structure information, the present invention converts the conditional mean μ of the denoising process into θ (·) and covariance ∑ θ (·) is further defined as: μ θ (x t ,G,t)=AGNN θ (x t ,G,t) ∑ θ (X t G,t)=σ 2 (t)I Among them, σ 2 (t) is the noise intensity fixed at time step t. The solution generated by diffusion is passed through the loss function To train: Among them, ∈ is Gaussian noise, ∈ θ is the noise predicted by the model. This loss function is the loss function of supervised learning; S7: This paper proposes a preference-guided combinatorial optimization (PCGO) method, which conceptualizes the denoising process of the diffusion model as a multi-step Markov decision process (MDP), uses human preference data to correspond to the combinatorial optimization problem, uses the quality of the generated solution to optimize the strategy, and uses direct preference optimization (DPO) to fine-tune the diffusion model to maximize the expected return. The state of the denoising process is defined as: Represents the current graph structure, time step, and solution status; The action of the denoising process is defined as: represents the next step solution generated from the current state; Therefore, for reinforcement learning, the following optimization objectives are given: Among them, G is a heterogeneous graph, t is the time step, and x t is the state at time step t, σ ω , σ l It represents the weight or category distribution between weak preference and strong preference in preference constraint. logρ(·) is the logarithm of partial score, and β is a hyperparameter that controls the strength of preference constraint and weighs the influence between two preferences. For the first part of preference score: represents the logarithmic probability ratio of the solution predicted by the model under weak preference. θ (·) is the predicted probability based on the model parameters θ, p ref (·) is the probability of the reference distribution, which usually represents the true distribution or the distribution of the benchmark model. is the final solution based on weak preference, is the intermediate state of weak preference at time step t, and weak preference corresponds to combinatorial optimization problem, that is, the solution with low quality / the solution with the same quality that better satisfies the combinatorial optimization constraints; For the second part of the preference score: represents the log-probability ratio of the solution predicted by the model under strong preference. is the final solution based on strong preference, is the strong preference intermediate state at time step t, and the strong preference corresponds to the combinatorial optimization problem, which is a high-quality solution; S8: Joint training of supervised learning and reinforcement learning for PGCO. Through S6 and S7, the losses of supervised learning and reinforcement learning on parameter θ have been obtained respectively, which are and Using adaptive weighting based on gradient norm, the total loss is as follows: in, This dynamic weighting mechanism effectively avoids the dominant effect of a certain loss term during training by balancing the gradient norm of each loss term, significantly stabilizing model training. S9: Fast reasoning of PGCO based on DDIM and DPM-solver accelerated sampling. After the model is trained through PGCO, the solution to the combinatorial optimization problem can be obtained by stepwise denoising. DDIM reduces the sampling steps of the diffusion model through a non-Markov denoising process while retaining the high quality of the generated solution. n+k to n DDIM PF-ODE solver DDIM Then: in, is the input diffusion noise, α and σ are diffusion coefficients, is the noise prediction model, c is the conditional information; For DPM-solver, this patent only considers the case of order 2, the ordinary differential equation solver Ψ DPM-Solver The detailed formula is: in, is the Log-SNR difference, and is the noise prediction value of the intermediate time step, and the calculation formula is: In this patented method, the advantages of DDIM and DPM-Solver are combined to accelerate sampling in stages. In the initial stage, DDIM is used for large-step deterministic sampling to quickly generate intermediate noise solutions; in the fine stage, DPM-Solver is used for second-order optimization sampling to further improve the accuracy and quality of the solution.

Citation Information

Cited By

  • Hybrid optimization solution method and device based on generative adversarial network assistance, equipment and storage medium

    CN121235029A

  • Time window constraint path modeling and optimizing method based on graph reinforcement learning

    CN121457777A

  • A time window constraint path modeling and optimization method based on graph reinforcement learning

    CN121457777B