Multi-objective combinatorial optimization method and system based on generative artificial intelligence

Through the sparse graph neural network and reinforcement learning method based on generative artificial intelligence, the solution space complexity and computing complexity challenges of multi-objective combination optimization problems are solved, and faster and better Pareto optimal solution set acquisition is achieved.

CN120373350APending Publication Date: 2025-07-25SHANGHAI JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510441054.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing multi-objective combination optimization methods have challenges in solving spatial complexity and computational complexity, and it is difficult to effectively solve multi-objective combination optimization problems, especially in path planning, resource allocation and scheduling problems.

Method used

Using a generative artificial intelligence method, sparse graph neural network modeling combined optimization problems, and combining consistency model with reinforcement learning, pre-trained models are trained to generate Pareto optimal solution sets, including building sparse graph neural network, training consistency model, initializing reinforcement learning environment, sampling data and optimizing strategies, and fitting action distribution and non-dominant solutions through KL divergence.

Benefits of technology

With faster inference speed, better Pareto frontier and Pareto optimal solution sets are obtained, which improves the efficiency and effect of multi-objective combination optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373350A_ABST
    Figure CN120373350A_ABST
Patent Text Reader

Abstract

The invention provides a multi-objective combinatorial optimization method and system based on generative artificial intelligence, and the method comprises two stages: firstly, a sparse graph neural network is used for modeling a combinatorial optimization problem, a consistency model training method suitable for multi-objective optimization is designed, and parameters of a consistency model serve as a pre-training model for reinforcement learning; and secondly, adopting reinforcement learning to guide the pre-training model to carry out strategy optimization in each single target direction, and further training a fusion strategy to realize multi-target optimization. According to the method, a better Pareto leading edge and a corresponding Pareto optimal solution set can be obtained at a higher reasoning speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-objective combinatorial optimization, and specifically, to a multi-objective combinatorial optimization method and system based on generative artificial intelligence. Background Art

[0002] Multi-objective combinatorial optimization is an important research field that combines multi-objective optimization and combinatorial optimization, aiming to simultaneously optimize multiple potentially conflicting objective functions in a discrete or finite solution space. Compared with single-objective combinatorial optimization, multi-objective combinatorial optimization needs to find the Pareto optimal solution set in the solution space to achieve scientific decision-making in complex scenarios. Due to the conflicts between objective functions and the complexity of the solution space, this problem is usually more challenging and has wide applications in path planning, resource allocation, and scheduling problems.

[0003] Existing methods for solving multi-objective combinatorial optimization problems are mainly divided into the following categories: The weighted scalarization method transforms the problem into a single-objective combinatorial optimization problem for solution by assigning different weights to multiple objectives. However, this method is sensitive to weight selection and difficult to fully cover the Pareto front. Evolutionary algorithms are also used to solve multi-objective problems, and multiple solutions are explored in parallel through population evolution to gradually approach the Pareto front. However, with the increase in the number of decision variables, the computational complexity of these algorithms and the demand for training data increase significantly, resulting in a gradual decline in their performance.

[0004] Patent application document CN118036329A discloses a method and system for ultra-multi-objective optimization based on knowledge transfer, including the following steps: Based on a multi-objective optimization problem, a graph construction algorithm is used to generate an optimized objective dependency graph by defining the optimization objective as a node in the graph, and the multi-objective optimization problem is represented by the edges between nodes, reflecting the structure of the multi-objective optimization problem. However, this patent cannot completely solve the existing technical problems and cannot meet the requirements of the present invention. Summary of the Invention

[0005] Aiming at the deficiencies in the prior art, the purpose of the present invention is to provide a multi-objective combinatorial optimization method and system based on generative artificial intelligence.

[0006] The multi-objective combinatorial optimization method based on generative artificial intelligence provided by the present invention includes:

[0007] Step 1: Construct a sparse graph neural network to model the graph structure of the combinatorial optimization problem, where the graph structure includes a node set and an edge set;

[0008] Step 2: Generate an initial solution set through a combinatorial optimization algorithm on multiple single objectives, and embed the objective weights into the global features of the graph structure;

[0009] Step 3: Train the consistency model to directly map the target weights to the corresponding solutions by generating trajectories conditionally, and generate the pre-trained model parameters;

[0010] Step 4: Initialize the reinforcement learning environment with the pre-trained consistency model parameters, and generate training data through data sampling;

[0011] Step 5: Adopt the Proximal Policy Optimization algorithm (PPO) to optimize the policy on multiple single objectives respectively, and fit the multi-objective action distribution through KL divergence;

[0012] Step 6: Iteratively update the Pareto optimal solution set to achieve multi-objective optimization through non-dominated solution screening.

[0013] Preferably, the training of the consistency model includes:

[0014] Gradually add Gaussian noise to the true solution x0 through the forward process of the diffusion model q to generate the noisy solution x t , and its mathematical representation is:

[0015]

[0016] where β t is the noise scheduling parameter, N is the Gaussian distribution, I is the identity matrix, and x t-1 is the noisy solution at time step t - 1;

[0017] The consistency model directly learns a one-step mapping from the noisy solution x t to the true solution x0 through training, that is:

[0018]

[0019] where f θ is the model function, t is the time step, is the conditional input;

[0020] Concatenate the target weight vector with the global feature vector of the graph structure G to generate the conditional input:

[0021]

[0022] The global feature h G is extracted through the graph pooling layer of the sparse graph neural network, representing the overall information of the graph structure;

[0023] Using the conditional input and the noisy solution x t as inputs, directly output the true solution x0 through the consistency model f θ , and the loss function is defined as:

[0024]

[0025] Among them, E is the expectation; by minimizing the loss function, the model learns the direct mapping relationship from the noise distribution and conditional weights to the Pareto optimal solution, realizing efficient conditional generation.

[0026] Preferably, the data sampling specifically includes:

[0027] Uniformly sample in the target weight space, embed the weights into the node features of the graph structure, and generate the distribution of solutions through the consistency model;

[0028] Use the greedy algorithm to extract candidate solutions that meet the constraints from the distribution, interact with the environment to obtain multi-objective reward values, and store them in the training buffer.

[0029] Preferably, the application of the KL divergence is specifically as follows:

[0030] For each optimization objective i, define the policy network π i (α|s; θ i ), which represents the probability distribution of selecting action α in state s, where θ i are the policy parameters;

[0031] In each policy optimization iteration, record the old policy parameters θ i,old , calculate the KL divergence D i between the new policy π i,new (α|s; θ KL ) and the old policy:

[0032]

[0033] Among them, E represents the expectation, and the KL divergence is used to quantify the magnitude of the policy change;

[0034] For each objective i, design a loss function that includes the policy gradient and the KL divergence penalty term:

[0035]

[0036] Among them, A i (s,a) is the advantage function of objective i, measuring the pros and cons of action a relative to the average performance; α is the penalty coefficient, controlling the constraint strength of the KL divergence;

[0037] Simultaneously optimize the loss functions of all objectives through the gradient descent method:

[0038]

[0039] Among them, η is the learning rate;

[0040] If the actual value of the KL divergence exceeds the preset threshold, adaptively adjust α or the update amplitude of the clipping strategy to prevent sudden changes in the strategy and maintain the balance of multi-objective optimization.

[0041] Preferably, the method for updating the Pareto optimal solution set includes:

[0042] In each generation of training, use the updated consistency model to infer new solutions and screen the solution set through non-dominated sorting;

[0043] Remove the dominated solutions, retain and add new non-dominated solutions until the solution set converges or reaches the preset termination condition.

[0044] According to the multi-objective combinatorial optimization system based on generative artificial intelligence provided by the present invention, it includes:

[0045] Module M1: Construct a sparse graph neural network to model the graph structure of the combinatorial optimization problem, where the graph structure includes a node set and an edge set;

[0046] Module M2: Generate an initial solution set through a combinatorial optimization algorithm on multiple single objectives and embed the objective weights into the global features of the graph structure;

[0047] Module M3: Train a consistency model, directly map the objective weights and the corresponding solutions through conditional generation trajectories, and generate pre-trained model parameters;

[0048] Module M4: Initialize the reinforcement learning environment with the pre-trained consistency model parameters and generate training data through data sampling;

[0049] Module M5: Use the proximal policy optimization algorithm PPO to optimize the policy on multiple single objectives respectively and fit the multi-objective action distribution through the KL divergence;

[0050] Module M6: Iteratively update the Pareto optimal solution set and achieve multi-objective optimization through non-dominated solution screening.

[0051] Preferably, the training of the consistency model includes:

[0052] Gradually add Gaussian noise to the true solution x0 through the forward process of the diffusion model to generate the noisy solution x t , and its mathematical representation is:

[0053]

[0054] where β t is the noise scheduling parameter, N is the Gaussian distribution, I is the identity matrix, and x t-1 is the noisy solution at time step t - 1;

[0055] The consistency model directly learns from the noisy solution x through trainingt One-step mapping to the true solution x0, i.e.:

[0056]

[0057] where f θ is the model function, t is the time step, is the conditional input;

[0058] Concatenate the target weight vector with the global feature vector of the graph structure G to generate the conditional input:

[0059]

[0060] The global feature h G is extracted through the graph pooling layer of the sparse graph neural network, representing the overall information of the graph structure;

[0061] Using the conditional input and the noisy solution x t as inputs, directly output the true solution x0 through the consistency model f θ , and the loss function is defined as:

[0062]

[0063] where E is the expectation; by minimizing the loss function, the model learns the direct mapping relationship from the noise distribution and conditional weights to the Pareto optimal solution, achieving efficient conditional generation.

[0064] Preferably, the data sampling specifically includes:

[0065] Uniformly sample in the target weight space, embed the weights into the graph structure node features, and generate the distribution of solutions through the consistency model;

[0066] Use the greedy algorithm to extract candidate solutions that satisfy the constraints from the distribution, interact with the environment to obtain multi-objective reward values, and store them in the training buffer.

[0067] Preferably, the application of the KL divergence is specifically as follows:

[0068] For each optimization objective i, define the policy network π i (α|s; θ i ), representing the probability distribution of selecting action α in state s, where θ i are the policy parameters;

[0069] In each policy optimization iteration, record the old policy parameters θ i,old , calculate the new policy π i (α|s; θ i,newKL divergence from the old policy:

[0070]

[0071] where \(E\) represents expectation, and KL divergence is used to quantify the magnitude of policy change;

[0072] For each objective \(i\), design a loss function that includes the policy gradient and the KL divergence penalty term:

[0073]

[0074] where \(A\) i (s,a) is the advantage function of objective \(i\), which measures the quality of action \(a\) relative to the average performance; \(\alpha\) is the penalty coefficient, which controls the constraint strength of the KL divergence;

[0075] Simultaneously optimize the loss functions of all objectives by the gradient descent method:

[0076]

[0077] where \(\eta\) is the learning rate;

[0078] If the actual value of the KL divergence exceeds the preset threshold, adaptively adjust \(\alpha\) or clip the policy update magnitude to prevent policy mutation and maintain the balance of multi-objective optimization.

[0079] Preferably, the method for updating the Pareto optimal solution set includes:

[0080] In each generation of training, use the updated consistency model to infer new solutions and screen the solution set through non-dominated sorting;

[0081] Remove the dominated solutions, retain and add new non-dominated solutions until the solution set converges or reaches the preset termination condition.

[0082] Compared with the prior art, the present invention has the following beneficial effects:

[0083] The present invention solves the multi-objective combinatorial optimization problem based on the consistency model and reinforcement learning. The solution includes two stages: First, use a sparse graph neural network to model the combinatorial optimization problem and design a consistency model training method suitable for multi-objective optimization. The parameters of this consistency model are used as the pre-trained model of reinforcement learning; Second, use reinforcement learning to guide the pre-trained model to perform policy optimization in each single-objective direction and further train the fusion policy to achieve multi-objective optimization; This method can obtain a better Pareto front and its corresponding Pareto optimal solution set at a faster inference speed. Brief Description of the Drawings

[0084] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description of non - restrictive embodiments with reference to the following drawings:

[0085] Figure 1 is a consistency model based on a sparse graph neural network;

[0086] Figure 2 is multi - objective policy optimization based on reinforcement learning;

[0087] Figure 3 is the overall algorithm flow. Specific Embodiments

[0088] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.

[0089] Embodiment 1

[0090] For the general form of the multi - objective combinatorial optimization problem:

[0091]

[0092] where G is an instance of the combinatorial optimization problem, represented by the graph structure G(V, E), V and E represent the node set and the edge set respectively; x ∈ {0, 1} N is the optimization variable of this problem, N is the dimension of the optimization variable; the feasible region Ω represents the set of feasible solutions that satisfy the constraints; the optimization objective is to minimize m objective functions f1(x; G), f2(x; G),..., f m (x; G).

[0093] The present invention proposes a multi - objective combinatorial optimization method based on a consistency model and reinforcement learning. The consistency model CM has the distribution fitting ability of the diffusion model and significantly improves the inference speed by reducing the iterative sampling steps. Reinforcement learning RL is used to optimize the generation strategy of CM and fine - tune its parameters to explore the Pareto - optimal solution set; this method can effectively solve the multi - objective combinatorial optimization problem and obtain the corresponding Pareto front.

[0094] The technical solutions adopted by the present invention mainly include two stages: a consistency model based on a sparse graph neural network structure and multi - objective policy optimization based on reinforcement learning, where:

[0095] The first stage: a consistency model based on a sparse graph neural network structure;

[0096] Step 1: Construct a sparse graph neural network G(V,E) to describe the combinatorial optimization problem;

[0097] Step 2: Use existing combinatorial optimization problem-solving algorithms on m single-objectives to obtain solutions

[0098] Step 3: Set the objective weights w=(w1,w2,…,w m ), and the solution on the i-th single-objective corresponds to the objective weight Add w i to the global features of G to obtain G i ;

[0099] Step 4: Supervised learning of the consistency model θ, and directly map the conditional generation trajectory i obtained based on G in Step 3 to the corresponding solution That is for the above i∈{1,2,…,m};

[0100] Second stage: Multi-objective policy optimization based on reinforcement learning;

[0101] Step 5: Use the consistency model trained in Step 4 as the initial model parameters Initialize the Pareto optimal solution set

[0102] Step 6: Sample the training data D of reinforcement learning based on the consistency model parameters , which is specifically divided into the following three steps: Step 6-1, Step 6-2, and Step 6-3; j

[0103] Step 6-1: Select the objective weights w=(w1,w2,...,w m ), where w1,...,w m ~Uniform(0,1), and Select a combinatorial optimization problem instance G, and add w to the node features of G;

[0104] Step 6-2: The state s=[x T ,G], where x T is uniformly distributed random noise, and G is the graph structure with objective weights added in Step 6-1. Use the inference process of the consistency model to generate the distribution of solutions Obtain the solution x0 that maximizes the distribution and satisfies all constraints of the combinatorial optimization problem through the greedy algorithm;

[0105] Step 6-3: Interact with the environment to obtain the rewards on m single-objectives as r=(r1(x0,G),...,r​m (x0, G)), store the state s and the corresponding reward r in the buffer D j ;

[0106] The above steps 6-1 to 6-3 are repeatedly executed until the buffer D j is full;

[0107] Step 7: Adopt the Proximal Policy Optimization algorithm PPO in reinforcement learning to train the policies on m single objectives respectively in the buffer D j ;

[0108] Step 8: Based on the KL divergence, simultaneously fit the output action distributions on m objectives and train the parameters of the next-generation consistency model

[0109] Step 9: Use the consistency model trained in Step 8 to infer the solution of the multi-objective combinatorial optimization problem in the buffer D, add non-dominated solutions to the current Pareto optimal solution set P j and delete dominated solutions to obtain the next-generation Pareto optimal solution set P j+1 ;

[0110] The above steps 6 to 9 are repeatedly executed until the algorithm termination condition is met. The termination condition can adopt whether the non-dominated solution set changes, etc.

[0111] Preferably, in the above step 2, any single-objective solution selection method can be used, including but not limited to the branch and bound method, simulated annealing method, etc.

[0112] Example 2

[0113] For example, a multi-objective combinatorial optimization problem with two objectives:

[0114]

[0115] The first stage: A consistency model based on a sparse graph neural network structure;

[0116] Step 1: Construct a sparse graph neural network G(V, E) to describe the combinatorial optimization problem;

[0117] Step 2: Use the existing combinatorial optimization problem-solving algorithm on 2 single objectives to obtain solutions

[0118] Step 3: Set the objective weights w = (w1, w2), the solution corresponds to the objective weight w1 = (1, 0), the solution corresponds to the objective weight w2 = (0, 1), and add w i to the global feature of G to obtain Gi ;

[0119] Step 4: Train the consistency model θ, and map the G i conditional generation trajectory directly to the corresponding solution That is the above i ∈ {1, 2};

[0120] Second stage: Multi-objective policy optimization based on reinforcement learning;

[0121] Step 5: Use the consistency model obtained by training in Step 4 as the initial model parameters Initialize the Pareto optimal solution set

[0122] Step 6: Based on the consistency model parameters Sample the training data D for reinforcement learning j , which is specifically divided into the following three steps: Step 6-1, Step 6-2, and Step 6-3;

[0123] Step 6-1: Select the target weights w = (w1, w2), where w1, w2 ∼ Uniform(0, 1), and w1 + w2 = 1. Select the combinatorial optimization problem instance G and add w to the node features of G;

[0124] Step 6-2: The state s = [x T , G], where x T is uniformly distributed random noise, and G is the graph structure with target weights added in Step 6-1. Use the inference process of the consistency model to generate the distribution of solutions Obtain the solution x′0 that maximizes the distribution satisfying all the constraints of the combinatorial optimization problem through the greedy algorithm;

[0125] Step 6-3: Interact with the environment to obtain the rewards on two single objectives as r = (r1(x′0, G), r2(x′0, G)), and store the state s and the corresponding reward r in the buffer D j ;

[0126] The above steps 6-1 to 6-3 are repeatedly executed until the buffer D j is full;

[0127] Step 7: Use the proximal policy optimization algorithm PPO in reinforcement learning to train the policies on two single objectives respectively on the buffer D j ;

[0128] Step 8: Based on the KL divergence, simultaneously fit the output action distributions on two objectives and train the parameters of the next-generation consistency model

[0129] Step 9: Using the consistency model trained in Step 8 to infer the solution of the multi-objective combinatorial optimization problem in the inference buffer D, and add non-dominated solutions to the current Pareto optimal solution set P j and delete the dominated solutions to obtain the next-generation Pareto optimal solution set P j+1 ;

[0130] The above Steps 6 to 9 are repeatedly executed until the algorithm termination condition is satisfied. The termination condition can be whether the non-dominated solution set changes, etc. The overall step flow is as Figure 3 shown.

[0131] Those skilled in the art know that in addition to implementing the system, device and their respective modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system, device and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same program. Therefore, the system, device and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structure within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the method or the structure within the hardware component.

[0132] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.

Claims

1. A multi-objective combinatorial optimization method based on generative artificial intelligence, characterized in that Including: Step 1: Construct a sparse graph neural network to model the graph structure of the combinatorial optimization problem, where the graph structure includes a node set and an edge set; Step 2: Generate an initial solution set through a combinatorial optimization algorithm on multiple single objectives, and embed the objective weights into the global features of the graph structure; Step 3: Train a consistency model, and directly map the objective weights and the corresponding solutions through conditional generation trajectories to generate pre-trained model parameters; Step 4: Initialize the reinforcement learning environment with the pre-trained consistency model parameters, and generate training data through data sampling; Step 5: Adopt the proximal policy optimization algorithm PPO to optimize the policies on multiple single objectives respectively, and fit the multi-objective action distribution through the KL divergence; Step 6: Iteratively update the Pareto optimal solution set to achieve multi-objective optimization through non-dominated solution screening.

2. The multi-objective combinatorial optimization method based on generative artificial intelligence according to claim 1, wherein, The training of the consistency model includes: Gradually add Gaussian noise to the true solution x0 through the forward process of the diffusion model q to generate the noisy solution x t , and its mathematical representation is: Among them, β t is the noise scheduling parameter, N is the Gaussian distribution, I is the identity matrix, and x t-1 is the noise solution at time step t - 1; The consistency model directly learns, through training, the one-step mapping from the noisy solution x t to the ground truth solution x0, that is: where f θ is the model function, t is the time step, is the conditional input; Concatenate the target weight vector with the global feature vector of the graph structure G to generate a conditional input: The global feature h G is extracted by the graph pooling layer of the sparse graph neural network, representing the overall information of the graph structure; With conditional input and noise mitigation x t as input, directly output the true solution x0 through the consistency model f θ The loss function is defined as: where E is the expectation; by minimizing the loss function, the model learns the direct mapping relationship from the noise distribution and conditional weights to the Pareto optimal solution, realizing efficient conditional generation.

3. The multi-objective combinatorial optimization method based on generative artificial intelligence according to claim 1, wherein The data sampling specifically includes: Uniformly sample in the objective weight space, embed the weights into the node features of the graph structure, and generate the distribution of solutions through the consistency model; Adopt the greedy algorithm to extract candidate solutions that meet the constraints from the distribution, interact with the environment to obtain multi-objective reward values, and store them in the training buffer.

4. The multi-objective combinatorial optimization method based on generative artificial intelligence according to claim 1, characterized in that The application of the KL divergence is specifically as follows: For each optimization objective i, define the policy network π i (α|s; θ i ), which represents the probability distribution of selecting action α in state s, where θ i are the policy parameters; In each policy optimization iteration, record the old policy parameters θ i,old , calculate the new policy π i (α|s; θ i,new ) and the KL divergence D KL : where E represents the expectation, and the KL divergence is used to quantify the magnitude of policy changes; For each objective i, design a loss function that includes a policy gradient and a KL divergence penalty term: Among them, A i (s,a) is the advantage function of target i, which measures the quality of action a relative to the average performance; α is the penalty coefficient, which controls the constraint strength of the KL divergence; Simultaneously optimize the loss functions of all objectives through the gradient descent method: where η is the learning rate; If the actual value of the KL divergence exceeds the preset threshold, adaptively adjust α or clip the policy update amplitude to prevent policy mutations and maintain the balance of multi-objective optimization.

5. The multi-objective combinatorial optimization method based on generative artificial intelligence according to claim 1, wherein The update method of the Pareto optimal solution set includes: In each generation of training, use the updated consistency model to infer new solutions, and screen the solution set through non-dominated sorting; Remove the dominated solutions, retain and add new non-dominated solutions until the solution set converges or reaches the preset termination condition.

6. A multi-objective combinatorial optimization system based on generative artificial intelligence, characterized in that, Including: Module M1: Construct a sparse graph neural network to model the graph structure of the combinatorial optimization problem, where the graph structure includes a node set and an edge set; Module M2: Generate an initial solution set through a combinatorial optimization algorithm on multiple single objectives, and embed the objective weights into the global features of the graph structure; Module M3: Train a consistency model, and directly map the objective weights and the corresponding solutions through conditional generation trajectories to generate pre-trained model parameters; Module M4: Initialize the reinforcement learning environment with the pre-trained consistency model parameters, and generate training data through data sampling; Module M5: Adopt the proximal policy optimization algorithm PPO to optimize the policies on multiple single objectives respectively, and fit the multi-objective action distribution through the KL divergence; Module M6: Iteratively update the Pareto optimal solution set to achieve multi-objective optimization through non-dominated solution screening.

7. The multi-objective combinatorial optimization system based on generative artificial intelligence according to claim 6, wherein The training of the consistency model includes: During the forward process of the diffusion model, Gaussian noise is gradually added to the true solution x0 to generate the noisy solution x t , and its mathematical representation is: where β t is the noise scheduling parameter, N is the Gaussian distribution, and I is the identity matrix x t-1 is the noise solution at time step t - 1; The consistency model directly learns, through training, the one-step mapping from the noisy solution x t to the true solution x0, that is: where f θ is the model function, t is the time step, is the conditional input; Concatenate the target weight vector with the global feature vector of the graph structure G to generate the conditional input: The global feature h G is extracted by the graph pooling layer of the sparse graph neural network, representing the overall information of the graph structure; With conditional input and noise resolution x t as input, directly output the true solution x0 through the consistency model f θ The loss function is defined as: where E is the expectation; by minimizing the loss function, the model learns the direct mapping relationship from the noise distribution and conditional weights to the Pareto optimal solution, realizing efficient conditional generation.

8. The multi-objective combinatorial optimization system based on generative artificial intelligence according to claim 6, characterized in that The specific data sampling includes: Uniformly sample in the target weight space, embed the weights into the node features of the graph structure, and generate the distribution of solutions through the consistency model; Use the greedy algorithm to extract candidate solutions that meet the constraints from the distribution, interact with the environment to obtain multi-objective reward values, and store them in the training buffer.

9. The multi-objective combinatorial optimization system based on generative artificial intelligence according to claim 6, characterized in that, The application of the KL divergence is specifically as follows: For each optimization objective i, define the policy network π i (α|s; θ i ), representing the probability distribution of selecting action α in state s, where θ i are the policy parameters; In each policy optimization iteration, record the old policy parameters θ i,old , calculate the new policy π i (α|s; θ i,new ) and the KL divergence between the old policy: Among them, E represents the expectation, and the KL divergence is used to quantify the magnitude of the policy change; For each objective i, design a loss function that includes the policy gradient and the KL divergence penalty term: Among them, A i (s,a) is the advantage function of target i, which measures the quality of action a relative to the average performance; α is the penalty coefficient, which controls the constraint strength of the KL divergence; Simultaneously optimize the loss functions of all objectives by the gradient descent method: Among them, η is the learning rate; If the actual value of the KL divergence exceeds the preset threshold, adaptively adjust α or clip the policy update amplitude to prevent policy mutation and maintain the balance of multi-objective optimization.

10. The multi-objective combinatorial optimization system based on generative artificial intelligence according to claim 6, wherein The update method of the Pareto optimal solution set includes: In each generation of training, use the updated consistency model to infer new solutions, and screen the solution set through non-dominated sorting; Remove the dominated solutions, retain and add new non-dominated solutions until the solution set converges or reaches the preset termination condition.

Citation Information

Patent Citations

  • Super-multi-objective optimization method and system based on knowledge migration

    CN118036329A