Combination optimization solving system and method based on general attack and defense framework
By modeling combinatorial optimization problems as graph structures and utilizing a shared graph encoder and a dedicated decoder, the problem of needing to design attack and defense schemes separately for each problem in existing technologies is solved. This enables cross-task attack and defense, improves the robustness and efficiency of the solver, and reduces computational resources and time costs.
Patent Information
- Application Number
- CN202511491570.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-19
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-10-19
AI Technical Summary
Existing combinatorial optimization methods require separate attack and defense schemes for each problem when facing multiple combinatorial optimization problems, resulting in low efficiency and difficulty in adapting to different scales and constraints. Furthermore, existing technologies lack a unified attack paradigm and a systematic attack-defense collaborative optimization mechanism, leading to insufficient generalization of attack and defense effects.
A combined optimization solution system based on a general attack and defense framework is adopted. Different tasks are modeled as graph structures through a multi-task modeling module. A common graph encoder is used to extract general graph features, and perturbations that meet task constraints are generated through dedicated decoder and masking. By combining an attack model based on reinforcement learning architecture and a defense model based on graph neural network, cross-task attack and defense can be achieved.
It significantly improves the stability and security of the solver in adversarial environments, reduces the number of model training iterations and storage costs, enhances the model's adaptability, reduces computational resource consumption in adversarial environments, improves the efficiency of model training, simplifies model resource consumption, and reduces model training time and computational overhead.
Smart Images

Figure CN120975358A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of combinatorial optimization technology, specifically relating to a combinatorial optimization solution system and method based on a general attack and defense framework. Background Technology
[0002] Combinatorial optimization is a class of problems that seek the optimal solution in a discrete feasible solution set (such as the Traveling Salesman Problem (TSP), the knapsack problem, the scheduling problem, etc.). Its core challenge is that the solution space grows exponentially with the problem size, making it difficult to solve by exhaustive search.
[0003] Existing combinatorial optimization methods mainly fall into two categories: one is the traditional heuristic and exact methods, such as the Lin–Kernighan-Helsgaun (LKH) algorithm, branch and bound method, and mixed integer programming solvers (such as Gurobi and SCIP). These methods have high solution quality for single problems (such as ATSP, CVRP, and DAG scheduling), but cannot be directly applied to general adversarial attack and defense research. The other category is neural solvers based on deep learning. For example, PointerNetworks uses pointer attention mechanism to generate variable-length solution sequences, S2V-DQN combines graph embedding with deep reinforcement learning to achieve constructive solutions, POMO proposes multiple optimal policy optimization to improve the solution performance of TSP / CVRP, and MatNet efficiently models problems such as ATSP / CVRP through matrix encoding.
[0004] Current combinatorial optimization problem solving techniques face several key challenges, which severely restrict the reliability and generalization ability of solvers in practical applications.
[0005] First, existing solvers (including traditional heuristic algorithms and deep learning models) generally suffer from insufficient adversarial robustness. Research shows that even well-trained neural solvers (such as MatNet and POMO) experience a significant drop in solution quality when faced with specially designed adversarial perturbations. For example, in the Traveling Salesman Problem, even minor adjustments to the weights of a few key edges can increase the path length output by more than 50%. More seriously, this vulnerability could lead to significant economic losses or safety hazards in real-world applications such as vehicle routing and task scheduling.
[0006] Existing attack methods have significant limitations. Most adversarial attack techniques are designed specifically for particular problems, lacking a unified attack paradigm. These problem-specific attack schemes not only require the development of new attack models for each type of combinatorial optimization problem, resulting in enormous R&D costs, but more importantly, they fail to leverage common characteristics across different problems to improve attack efficiency. Furthermore, traditional attack methods often employ static optimization strategies, which are ill-suited to real-world problem instances of varying scales and constraints, leading to insufficient generalization of attack effectiveness.
[0007] In terms of defense, existing technologies also face severe challenges. Mainstream adversarial training methods typically use a fixed distribution of adversarial samples for training, and this static defense strategy is ill-suited to cope with dynamically changing attack methods. More importantly, most current defense solutions are developed independently of the attack process, lacking a systematic mechanism for collaborative optimization between offense and defense. This fragmented design paradigm limits the effectiveness of defenses and prevents the formation of a continuously evolving security protection system. Furthermore, existing defense methods often require designing separate defense strategies for each specific problem, which not only increases the complexity of engineering implementation but also makes it difficult to guarantee consistent defense effectiveness across different problems. Summary of the Invention
[0008] To address the aforementioned shortcomings of existing technologies, the combinatorial optimization solution system and method based on a general attack and defense framework provided by this invention solves the problems that existing combinatorial optimization solution methods are unable to solve multiple combinatorial optimization problems. They require the design of separate attack and defense schemes for each optimization problem, reducing the efficiency of solving combinatorial optimization problems and making it difficult to adapt to real-world problem instances of different scales and constraints. This results in insufficient generalization of attack and defense effects, reduced quality of solution results, and increased engineering implementation complexity. This invention not only effectively identifies the security weaknesses of existing solvers but also systematically improves the robustness of solvers in adversarial environments. By organically combining attack and defense models, it achieves full-process coverage from problem discovery to problem resolution, providing a solid security guarantee for the industrial application of combinatorial optimization technology.
[0009] To achieve the aforementioned objectives, the technical solution adopted by this invention is: a combined optimization solution system based on a general attack and defense framework, comprising: The multi-task modeling module is used to model the combined optimization problem of different tasks as a graph structure representation, and construct it as a graph structure Markov decision process according to the task type. The attack model is used to extract general graph feature representations from graph structures through a shared graph encoder, and generate perturbations that satisfy the constraints of the corresponding tasks through dedicated decoders and masking processes for different tasks. The defense model takes graph structure and node / edge features as input, and generates a decoding sequence that adapts to task perturbations and satisfies task constraints through a selected target defense solver, thereby achieving combinatorial optimization solution.
[0010] Furthermore, the task types include the Asymmetric Traveling Salesman Problem (ASTP), the Vehicle Routing Problem (CVRP) with capacity constraints, and the Directed Acyclic Graph (DAG) scheduling problem. In the constructed Markov decision process: Each state in the state space includes the current graph structure, the node feature matrix, and the global matrix; Each action in the action space is a perturbation operation on the node-to-graph structure; the perturbation operation is determined according to the task type. The reward function represents the degree of performance degradation of the target defense solver after adding perturbations.
[0011] Furthermore, the attack model is based on a reinforcement learning architecture and includes: The Actor network includes a shared encoder and a dedicated decoder. The shared encoder is used to extract general graph feature representations for optimization problems with different task combinations. The dedicated decoder is used to map the output of the shared encoder to the action space of the corresponding task and dynamically generate task legality masks corresponding to different task action spaces through masking, and then output the action probability distribution after masking. The Critic network is used to learn the state-value function for different tasks based on the action probability distribution output by the Actor network, and then feed it back to the Actor network. When the task is the Asymmetric Traveling Salesman Problem (ASTP), the learned state value function is the value function for estimating the path length of the ASTP; when the task is the Vehicle Routing Problem (CVRP) with capacity constraints, the learned state value function is the value function for estimating the total travel distance of the vehicle routing problem; when the task is the Directed Acyclic Graph (DAG) scheduling problem, the learned state value function is the value function for estimating the completion time of the scheduling problem.
[0012] Furthermore, the target defense solver is an end-to-end combinatorial optimization solver based on graph neural networks, MatNet.
[0013] A combinatorial optimization solution method includes the following steps: S100. The combinatorial optimization problem of the task to be solved is modeled as a Markov decision process with a graph structure. S200. In the attack model, a general graph feature representation of different task combination optimization problems is extracted by a shared graph encoder. The general graph feature representation is mapped to the corresponding action space by a dedicated decoder for the task to be solved. The task legality mask of the action space is dynamically generated by mask processing. S300. In the defense model, the node / edge features of the graph structure are used as input. The selected target defense solver gradually generates a decoding sequence that adapts to task perturbations and satisfies task constraints, thereby achieving combinatorial optimization solution.
[0014] Further, in step S200, extracting a general graph feature representation for optimization problems with different task combinations includes: S201. Extract features of the graph structure through a forward graph convolutional network, and extract features of the transposed graph structure through a backward graph convolutional network. S202. The extracted features are spliced together through a splicing layer; S203. The attention mechanism is used to calculate and extract the general feature representation of the spliced features through the global attention pooling layer.
[0015] Further, in step S200, mapping the general graph feature representation to the corresponding action space includes: S204. For any pair of nodes in the graph structure, concatenate the embedding vectors of the two nodes and the general graph feature representation. Among them, node pairs represent perturbation operations on the graph structure; S205. Based on the task type, the spliced features are used to generate the original score of the node pair through a multilayer perceptron; S206. The original scores of all node pairs are converted into corresponding probability distributions by using the Softmax activation function to obtain the corresponding action space. Specifically, when the task is the Asymmetric Traveling Salesman Problem (ASTP), an attention mechanism with query-key dot product is used to calculate the original scores of connectivity and path cost for edge replacement / perturbation. When the task is the Vehicle Routing Problem (CVRP) with capacity constraints, an attention mechanism is used to further calculate the node embedding vectors in the vehicle path, and then the demand node embedding vector is added to the coordinate node embedding vector to obtain the original score considering customer demand and vehicle remaining capacity. When the task is the Directed Acyclic Graph (DAG) scheduling problem, context information is enhanced by splicing, the embedding vectors of selected nodes are paired with the embedding vectors of other nodes, and the original score considering whether or not topological constraints are generated is obtained by using dynamic masking.
[0016] Furthermore, in step S200, the formula for generating the task validity mask of the action space is: In the formula, This indicates the new policy after weighted normalization in state. Select action The probability, Indicates the old strategy in state Select action The probability, This indicates the old strategy in the action state. Select action The probability, Indicates the action currently being evaluated. Indicates the current state. Indicates the legality of the action, when For illegal actions, when This was a legal action at the time.
[0017] Further, in step S200, the training loss function of the attack model... for: In the formula, Indicates the strategy loss. Represents the loss of the value function. Let represent the expectation for step t. Represents the entropy function. Indicates the state at step t The following strategy Represents the Actor network parameters. Indicates the parameters of the Critic network. Represents the state at step t. , Represents the policy distribution. and These represent the first weight hyperparameter and the second weight hyperparameter, respectively. This represents the function that takes the minimum value. This represents the ratio of the probability of the new strategy to the probability of the old strategy. This represents the estimation of the advantage function. This represents the clipping function. Indicates the clipping parameters. Represents the numerical stability constant; Indicates the parameters in the Actor network The current strategy is as follows: Indicates the old strategy, This represents the action at step t. Indicates the state at step t; This represents the output of the Critic network for the corresponding task; This represents the batch-normalized reward. Indicates a reward. and These represent the mean and standard deviation of the current batch of rewards, respectively.
[0018] Further, in step S300, the target defense solver is trained using an iterative adversarial retraining method, including: The original dataset is perturbed using a trained attack model to generate the first adversarial example; the perturbation satisfies the legality constraints of different tasks. The original dataset is mixed with the first adversarial example to obtain a hybrid dataset; The target defense solver is iteratively trained using a mixed dataset to obtain the defense model; In each iteration of training, the defense model obtained from the previous training is used as the attack target. New adversarial examples are generated using the attack model, and a new hybrid dataset is constructed with the stable test benchmark dataset to train the defense model in the current iteration round.
[0019] The beneficial effects of this invention are as follows: (1) High versatility, reducing redundant design and training This invention employs a unified multi-task attack model structure, sharing a multi-layer graph convolutional network (GCN) encoder, and setting task-specific decoding branches and a dynamic masking mechanism at the output end to achieve action generation for different combined optimization tasks (ATSP, CVRP, DAG). Since the parameters and feature extraction process of the encoder are shared by multiple tasks, and the specific constraints of different tasks are automatically guaranteed by the masking mechanism, a single model can adapt to multiple tasks without requiring separate attacker training for each task. Compared with existing methods, this design significantly reduces the number of model training iterations and storage costs, saves computational resources and network bandwidth (reducing model file transfer volume in distributed deployments), and improves overall deployment efficiency in multi-task environments.
[0020] (2) Improved adaptive robustness to ensure data security This invention introduces an iterative attack-defense closed-loop mechanism during training. The attack model generates new adversarial perturbation samples online, and the defense model uses these samples for retraining, updating in multiple rounds to form a dynamic game. This ensures that the defense model is continuously exposed to the latest and strongest attack instances during training. This mechanism enables the defender to continuously learn to cope with perturbation patterns of different distributions and strategies, overcoming the weakness of traditional static defenses that are easily bypassed by new attacks. Compared with existing defense methods, this invention significantly improves the defender's ability to cope with adaptive attacks, ensures the stability and security of the combinatorial optimization solver's results in long-term operation, and reduces security risks such as task scheduling errors or path planning failures.
[0021] (3) High resource utilization, reducing network and computing overhead. This invention reduces the total number of parameters by using a shared encoder, avoids full gradient backtracking by utilizing the PPO strategy for updates, and introduces a legality mask during the inference phase to directly block illegal actions, reducing invalid computations. During inference, only one forward propagation is needed to complete the action probability calculation, and the masking mechanism ensures that subsequent sampling is performed only on the set of legal actions, significantly reducing the evaluation of invalid node pairs. Compared with methods that run multiple models and tasks independently, this invention effectively reduces inference latency and memory usage, making it particularly suitable for distributed deployment scenarios with limited network bandwidth, such as edge computing nodes.
[0022] (4) Ensure the legitimacy of the disturbance and improve the effectiveness of the attack. This invention employs a task-specific legality masking mechanism to pre-filter perturbations that do not meet task constraints (such as CVRP capacity limits, DAG loop avoidance, and ATSP path connectivity) during the action generation phase, ensuring that the actions output by the policy are always legal. This mechanism performs legality judgment during the action selection phase, avoiding the risk of illegal perturbations entering the environment and causing abnormal states or contaminating training samples. Compared to schemes that require post-event removal of illegal samples, this invention directly improves the stability of attack training and the effectiveness of perturbation generation, allowing attack samples to be directly used for retraining the defense model, thereby shortening the attack-defense iteration cycle.
[0023] (5) Defensive generalization across solvers and distributions This invention employs a solver-independent reward function design, enabling the attacker to learn perturbation patterns applicable to different solvers (including heuristics, neural networks, and commercial optimizers) and different data distributions. Because the reward definition does not depend on the internal information of a specific algorithm, the attacker maintains stable attack performance in comprehensive evaluations across multiple distributions and multiple solvers.
[0024] When this attack method is used to train a defense model, the defense model maintains a high level of performance even when dealing with unseen solvers and datasets. Compared to existing defense methods that rely on training with a single solver, this invention significantly enhances cross-platform generalization capabilities and reduces security risks when deploying in new environments.
[0025] (6) Advantage of training time This invention offers significant advantages in training efficiency. Existing combinatorial optimization adversarial attack methods typically require designing separate network structures and feature processing flows for each task (such as ATSP, CVRP, and DAG), and training independent attack models for each. This means that when dealing with multi-task scenarios, the entire training process must be repeated three or more times. Each training iteration involves parameter initialization from scratch, training of the feature encoder, and iterative updates of policy optimization, which not only consumes a large amount of computational resources and GPU memory but also significantly extends the production cycle of usable model versions. For task combinations with large problem scales, this multi-model independent training mode leads to a linear or even superlinear increase in total training time, making it difficult to quickly update models within a limited computational budget.
[0026] The general multi-task attack model structure proposed in this invention utilizes a shared multi-layer graph convolutional network (GCN) encoder and global feature embedding, combined with task-specific decoding branches and a dynamic legitimacy mask mechanism, allowing it to adapt to multiple different tasks simultaneously with only one training iteration. Since the encoder parameters are fully shared across multiple tasks, the model can simultaneously learn the common structural features and task-specific constraints of multiple tasks in a single forward and backward propagation process, avoiding redundant computation and memory consumption caused by repeated training. Under the same task scale, the total training time of this invention is reduced by more than 60% compared to existing single-task multi-model methods. For example, the original method requires thousands of minutes of cumulative training on three tasks, while this invention can generate an attack model that can be applied to three tasks simultaneously in a single training iteration. This design not only significantly reduces computational energy consumption and hardware resource consumption but also greatly improves model iteration speed and deployment flexibility, making it particularly suitable for multi-task collaborative optimization and distributed computing environments with limited network bandwidth and computing resources. Attached Figure Description
[0027] Figure 1 The diagram shows the system structure for the combined optimization solution of the general attack and defense framework provided by this invention. Detailed Implementation
[0028] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0029] Example 1: This invention provides a combined optimization solution system based on a general attack and defense framework, comprising: The multi-task modeling module is used to model the combined optimization problem of different tasks as a graph structure representation, and construct it as a graph structure Markov decision process according to the task type. The attack model is used to extract general graph feature representations from graph structures through a shared graph encoder, and generate perturbations that satisfy the constraints of the corresponding tasks through dedicated decoders and masking processes for different tasks. The defense model takes graph structure and node / edge features as input, and generates a decoding sequence that adapts to task perturbations and satisfies task constraints through a selected target defense solver, thereby achieving combinatorial optimization solution.
[0030] In this embodiment of the invention, in order to achieve a unified attack and defense framework, different combinatorial optimization problems are first modeled as graph structures to ensure the consistency and processability of the input form of subsequent models.
[0031] The task types in this embodiment of the invention include the Asymmetric Traveling Salesman Problem (ASTP), the Vehicle Routing Problem with Capacity Constraints (CVRP), and the Directed Acyclic Graph (DAG). Specifically, in the directed graph structure representation obtained by modeling the Asymmetric Traveling Salesman Problem (ASTP), the node set represents cities, with each node corresponding to a city; the edge set represents the directed connection between cities; any edge represents the travel path between two cities; and the edge weight represents the travel cost between cities, with values derived from the distance matrix. In the undirected graph structure representation obtained by modeling the Vehicle Routing Problem with Capacity Constraints (CVRP), the node set includes warehouses and customers; the edge set represents the connection path between cities or customers; and the weight of any edge is obtained from the Euclidean distance. In the graph structure representation obtained by modeling the DAG, the node set represents computational tasks; the feature vector of each task node includes execution time and resource requirement vectors; the edge set represents the dependencies between tasks; and the edge features do not contain weights and are represented as empty vectors. Furthermore, the overall graph structure representation maintains acyclicity, and the correctness of dependency constraints is ensured through topological sorting.
[0032] By using the unified modeling approach described above, the combined optimization problem of different task types is transformed into a graph structure representation, which is then used for encoding and perturbation operations in the subsequent graph neural network.
[0033] In this embodiment of the invention, the attack process of solving the above combinatorial optimization problem is modeled as a Markov decision process; wherein, each state in the state space includes the current graph structure, node feature matrix and global matrix (such as current step number, remaining action budget, etc.); each action in the action space is a perturbation operation of the node to the graph structure; the perturbation operation is determined according to the task type; the reward function is the degree of performance degradation of the target defense solver after adding perturbation.
[0034] Furthermore, when determining the perturbation operation, for the Asymmetric Traveling Salesman Problem (ASTP), the main perturbation is the edge weight perturbation, such as halving or increasing the edge weight; for the Vehicle Routing Problem (CVRP) with capacity constraints, the main perturbation is the demand perturbation or edge weight modification; and for the Directed Acyclic Graph (DAG) scheduling problem, the main perturbation is the edge deletion or weight modification to maintain acyclicity.
[0035] The reward function is used to measure the attack effect of the attack model. It is defined as the difference between the objective function of the target defense solver before and after adding the perturbation operation, such as the path length in the asymmetric traveling salesman problem ASTP, the total travel distance in the vehicle routing problem CVRP with capacity constraints, and the maximum scheduling completion time in the directed acyclic graph scheduling problem DAG.
[0036] The attack model in this embodiment of the invention is based on a reinforcement learning architecture and includes: The Actor network includes a shared encoder and a dedicated decoder. The shared encoder is used to extract general graph feature representations for optimization problems with different task combinations. The dedicated decoder is used to map the output of the shared encoder to the action space of the corresponding task and dynamically generate task legality masks corresponding to different task action spaces through mask processing, and then output the processed action probability distribution to ensure that all generated actions satisfy the task constraints. The Critic network is used to learn the state-value function for different tasks based on the action probability distribution output by the Actor network, and then feed it back to the Actor network. When the task is the Asymmetric Traveling Salesman Problem (ASTP), the learned state value function is the value function for estimating the path length of the ASTP; when the task is the Vehicle Routing Problem (CVRP) with capacity constraints, the learned state value function is the value function for estimating the total travel distance of the vehicle routing problem; when the task is the Directed Acyclic Graph (DAG) scheduling problem, the learned state value function is the value function for estimating the completion time of the scheduling problem.
[0037] In this embodiment of the invention, in the network architecture of the above attack model, the Actor network adopts a unified encoder-decoder architecture, in which the encoder shares parameters among all tasks, and the decoder selects the corresponding branch according to the task type; the Critic network adopts a task-specific design, with each Critic network containing independent parameters, specifically handling the feature representation and value estimation of the corresponding task.
[0038] Based on the above Actor network structure design, the three tasks ATSP, CVRP and DAG can share the learned general graph feature representations through a shared encoder, while their respective decoders handle task-specific constraints and features. This design significantly reduces the number of model parameters, improves training efficiency, and achieves cross-task generalization capability.
[0039] In this embodiment of the invention, during the training of the aforementioned attack network, a corresponding Critic network is selected based on the current task type for value estimation and loss calculation. State value estimation is learned by minimizing the mean squared error, achieving knowledge sharing and specialized learning between tasks. Specifically, the Critic network outputs a state value, and the training objective is to minimize the squared error between this value and the discounted reward. The discounted reward is calculated by accumulating future rewards, using a discount factor to balance the importance of immediate and long-term rewards. This design allows each Critic network to specifically learn the reward structure and value distribution of its corresponding task, improving the accuracy of value estimation and training stability.
[0040] In this embodiment of the invention, the main purpose of the defense model is to improve the robustness of the target defense solver under adversarial perturbations. Without changing the solver structure, the invention enables the solver to gradually adapt to the multi-task perturbation instances generated by the unified attacker through iterative adversarial retraining, thereby significantly improving the performance stability of the solver under normal and abnormal input distributions.
[0041] In this embodiment of the invention, the target defense solver is an end-to-end combinatorial optimization solver MatNet based on graph neural networks, which can handle a variety of graph optimization tasks. Its input is graph structure and node / edge features, and its output is a decoded sequence that satisfies task constraints. The internal structure of MatNet consists of multiple layers of graph convolutional units and task-specific decoders.
[0042] In a specific embodiment of the present invention, the model parameters of the above-mentioned target defense solver are set as follows: embedding dimension is 384, encoder layer is 8 layers of graph convolution, attention head is 16, QKV dimension is 32, feedforward network hidden layer is 768, hybrid scoring hidden layer dimension is 32, and Logit pruning threshold is 12.
[0043] Example 2: Based on the combinatorial optimization solution system based on the general attack and defense framework in Example 1, this invention provides a corresponding combinatorial optimization solution method, such as... Figure 1 As shown, it includes the following steps: S100. The combinatorial optimization problem of the task to be solved is modeled as a Markov decision process with a graph structure. S200. In the attack model, a general graph feature representation of different task combination optimization problems is extracted by a shared graph encoder. The general graph feature representation is mapped to the corresponding action space by a dedicated decoder for the task to be solved. The task legality mask of the action space is dynamically generated by mask processing. S300. In the defense model, the node / edge features of the graph structure are used as input. The selected target defense solver gradually generates a decoding sequence that adapts to task perturbations and satisfies task constraints, thereby achieving combinatorial optimization solution.
[0044] In step S200 of this embodiment of the invention, based on the structure of the aforementioned attack model, a general graph feature representation of different task combination optimization problems is extracted from the Actor network, including: S201. Extract features of the graph structure through a forward graph convolutional network, and extract features of the transposed graph structure through a backward graph convolutional network. S202. The extracted features are spliced together through a splicing layer; S203. The attention mechanism is used to calculate and extract the general feature representation of the spliced features through the global attention pooling layer.
[0045] In the above process of this embodiment of the invention, three different tasks (ATSP, CVRP, and DAG scheduling) share the same encoder to achieve cross-task knowledge transfer and parameter sharing. Specifically, forward graph convolutional networks and backward graph convolutional networks are used to process the input graph structure to generate node-level representations. Regardless of whether the input is the distance matrix of ATSP, the demand graph of CVRP, or the dependency graph of DAG, features are extracted through the same encoder architecture, enabling the model to learn a general graph structure representation.
[0046] In step S200 of this embodiment of the invention, based on the structure of the aforementioned attack model, in the Actor network, the general graph feature representation is mapped to the corresponding action space, including: S204. For any pair of nodes in the graph structure, concatenate the embedding vectors of the two nodes and the general graph feature representation. Among them, node pairs represent perturbation operations on the graph structure; S205. Based on the task type, the spliced features are used to generate the original score of the node pair through a multilayer perceptron; S206. The original scores of all node pairs are converted into corresponding probability distributions by using the Softmax activation function to obtain the corresponding action space. Specifically, when the task is the Asymmetric Traveling Salesman Problem (ASTP), an attention mechanism with query-key dot product is used to calculate the original scores of connectivity and path cost for edge replacement / perturbation. When the task is the Vehicle Routing Problem (CVRP) with capacity constraints, an attention mechanism is used to further calculate the node embedding vectors in the vehicle path, and then the demand node embedding vector is added to the coordinate node embedding vector to obtain the original score considering customer demand and vehicle remaining capacity. When the task is the Directed Acyclic Graph (DAG) scheduling problem, context information is enhanced by splicing, the embedding vectors of selected nodes are paired with the embedding vectors of other nodes, and the original score considering whether or not topological constraints are generated is obtained by using dynamic masking.
[0047] In this embodiment, the decoding process follows a two-step mechanism: first, candidate perturbation actions (node pairs) are selected, and then the selected perturbation operation (edge deletion or weight halving) is executed. During this process, a masking mechanism forces the selection probability of invalid nodes to be 0, maintaining the legality of constraints for different tasks. To improve efficiency, edges currently included in the solver's predicted solution are excluded from the action space, and the complete attack trajectory consists of sequential actions, allowing the agent to explore the cumulative impact on the solver's output.
[0048] In step S200 of this embodiment of the invention, based on the structure of the aforementioned attack model, the formula for generating the task legality mask of the action space in the Actor network is as follows: In the formula, This indicates the new policy after weighted normalization in state. Select action The probability, Indicates the old strategy in state Select action The probability, This indicates the old strategy in the action state. Select action The probability, Indicates the action currently being evaluated. Indicates the current state. Indicates the legality of the action, when For illegal actions, when This was a legal action at the time.
[0049] In this embodiment of the invention, during the code implementation process of dynamically generating task legality masks corresponding to different task action spaces through masking, the score of illegal actions is directly added... For each valid action, add 0, and then perform a softmax calculation on (scores + mask). This implementation is mathematically equivalent to the mask formula mentioned above, and can be expressed as: When mask= When, exp( The probability of an illegal action is 0. When mask=0, exp(0)=1, preserving the original probability distribution. In this embodiment of the invention, a two-stage masking strategy is used to generate the mask corresponding to the Asymmetric Traveling Salesman Problem (ASTP). In the first stage, all nodes are initialized to... Then, the nodes with valid candidate edges are set to 0. Specifically, for each node... If the node has candidate edges (excluding edges on the current optimal path), then the task validity mask is used. ,otherwise In the second phase, given the nodes selected in the first phase, the mask is updated to only allow nodes connected to the selected nodes through valid candidate edges.
[0050] In this embodiment of the invention, the masking mechanism of the Vehicle Routing Problem with Capacity Constraints (CVRP) is the same as that of the Automatic Targeting Problem (ATSP), but it is specifically designed to respect path and demand constraints. Specifically, in the first stage, the mask excludes nodes that are already part of the current optimal path, or nodes whose selection would violate vehicle capacity. In the second stage, after selecting a node, the mask is updated to only allow the next feasiblely accessible node, while considering path continuity and remaining vehicle capacity.
[0051] In this embodiment of the invention, the Directed Acyclic Graph (DAG) scheduling problem is specifically designed to maintain the acyclic property of the graph. In the first stage, the mask excludes nodes without removable outgoing edges (i.e., nodes without valid dependencies that can be removed). In the second stage, after selecting the source node, the mask is updated to only allow target nodes connected by removable edges, and removal of these nodes will not introduce a cycle. This ensures that each perturbation corresponds to the deletion of valid edges that maintain the DAG structure and respect all scheduling constraints. Furthermore, the mask value is set as follows: illegal actions are... The probability of a legal action is 0, and the softmax function ensures that the probability of an illegal action is 0. Legal actions are assigned probabilities according to their original scores. This design ensures that the attacker only explores legal actions during training and inference, improving training efficiency and attack quality.
[0052] In this embodiment of the invention, based on the aforementioned attack model structure design, the training objective of the attack model is: to select the appropriate Critic network for value estimation and loss calculation according to the current task type, and to learn state value estimation by minimizing the mean square error, thereby achieving knowledge sharing and specialized learning among tasks.
[0053] Based on this, in this embodiment of the invention, the training loss function of the attack model for: In the formula, Indicates the strategy loss. Represents the loss of the value function. Let represent the expectation for step t. Represents the entropy function. Indicates the state at step t The following strategy Represents the Actor network parameters. Indicates the parameters of the Critic network. Represents the state at step t. , Represents the policy distribution. and These represent the first and second weight hyperparameters, respectively. This represents the function that takes the minimum value. This represents the ratio of the probability of the new strategy to the probability of the old strategy. This represents the estimation of the advantage function. This represents the clipping function. Indicates the clipping parameters. Represents the numerical stability constant. Indicates the parameters in the Actor network The current strategy is as follows: Indicates the old strategy, This represents the action at step t. This represents the state at step t. This represents the output of the Critic network for the corresponding task. This represents the batch-normalized reward. Indicates a reward. and These represent the mean and standard deviation of the current batch of rewards, respectively.
[0054] In a specific embodiment of the present invention, based on the above loss function design, the model training parameters are set as follows: 10 training epochs, 64 output dimensions of the forward / backward graph convolutional network, 1 batch size, 0.1 clip parameter, and 0.95 discount factor.
[0055] Furthermore, the training process of the attack model adopts the PPO strategy of "randomly selecting task type in each round of training". Through random sampling, it is ensured that the attacker can learn perturbation patterns of multiple task types at the same time, so as to achieve cross-task knowledge transfer.
[0056] In step S300 of this embodiment of the invention, the target defense solver is trained using an iterative adversarial retraining method, including: The original dataset is perturbed using the trained attack model to generate the first adversarial example; the perturbation satisfies the legality constraints of different tasks. The original dataset is mixed with the first adversarial example to obtain a hybrid dataset; The target defense solver is iteratively trained using a mixed dataset to obtain the defense model; In each iteration of training, the defense model obtained from the previous training is used as the attack target. New adversarial examples are generated using the attack model, and a new hybrid dataset is constructed with the stable test benchmark dataset to train the defense model in the current iteration round.
[0057] In a specific embodiment of the present invention, during the training process of the above-mentioned defense model, the training parameters of MatNet are set as follows: the number of training epochs is 30, the number of episodes is 2500, and the batch size is 64.
[0058] In one specific embodiment of the present invention, during the training process of the above-mentioned defense model, the optimizer parameters are set as follows: the learning rate is 1×10⁻⁶. -4 The weight decays to 1×10 -5 In the learning rate scheduler, the decay milestones are [10, 20, 30], and the decay rate is 0.85.
[0059] In a specific embodiment of the present invention, during the training process of the above-mentioned defense model, the iterative training strategy is further as follows: In the k-th iteration, the solver is fully trained using the defense training dataset of the k-th iteration. After training, the updated solver is passed to the attacker as the attack target, generating a new round of adversarial samples. These samples are then mixed with the original samples to form the training dataset for the next round of training, and then the next round of training begins. For example, the default number of iteration rounds is set to 3. Each round of training continues based on the solver weights of the previous round, rather than re-initializing them, to achieve continuous adaptation to new perturbation patterns. During training, the solver's loss function remains the same as the original task. For example, in path optimization tasks, the loss is minimized using path length; in scheduling tasks, the loss is minimized using maximum completion time. The optimizer, learning rate, and other training hyperparameters are consistent with the normal training phase to ensure comparability. In this embodiment of the invention, after each round of iterative training is completed, the defense model is tested on the generated first round of adversarial samples, and effective defense judgment conditions are set to evaluate the defense effect.
[0060] In this embodiment of the invention, the defense model is retrained using an iterative update adversarial approach, which enables the solver to maintain stable performance when faced with progressively increasing adversarial perturbations. By introducing multi-task perturbations, the defense capability is not only targeted at a single task, but also covers multiple combinatorial optimization problems. The defense process does not rely on task-specific prior knowledge and does not require modification of the solver structure, thus exhibiting high versatility and portability.
[0061] It should be noted that the combination optimization solution method of the general attack and defense framework provided by the present invention can be applied not only to the three tasks of ATSP, CVRP and DAG, but also to the combination optimization problem of other tasks according to the actual needs of users. It should not be assumed that the present invention can only achieve the solution of these three tasks just because these three types of tasks are mentioned in the present invention.
[0062] Example 3: This invention provides examples to verify the practicality and generalizability of the combinatorial optimization solution method in Example 2.
[0063] In this embodiment, the attack and defense models were evaluated on multiple test sets with different distributions. Unlike the training phase, which only used uniformly distributed random data, the testing phase introduced data distributions with significant differences to test the model's performance under unknown distributions.
[0064] For the ATSP and CVRP problems, during the training phase, the unified attacker and defense models are trained using only data generated by a random uniform distribution; during the testing phase, the following additional distribution is introduced: a) Tsplib and CVRPlib benchmark datasets: These classic datasets are derived from real-world transportation and logistics problems and have more complex geometric structures and non-uniform distribution characteristics.
[0065] b) Gaussian distribution example: Compared with the uniform distribution, the point set under the Gaussian distribution exhibits central clustering, which poses different challenges to path and capacity constraints.
[0066] Experimental results show that: a) The attack model can still significantly weaken the performance of the original solver on Tsplib, CVRPlib and Gaussian distribution instances, verifying its universal attack capability on data outside the training distribution.
[0067] b) The defended model also demonstrates robustness on these cross-distribution test sets. Compared with the undefended model, its performance under adversarial perturbations is significantly improved, proving that iterative adversarial retraining effectively improves the robustness of the model in the OOD environment.
[0068] For the DAG scheduling problem, during the training phase, the attack and defense models learn from uniformly randomly generated DAG data; during the testing phase, two distinctly different DAG distributions are introduced: a) TPC-H DAGs: Originating from database query optimization benchmarks, they have a typical task-dependent structure and uneven computational load.
[0069] b) Random graphs are obtained through a probabilistic random edge generation mechanism, resulting in denser and more irregular task dependencies.
[0070] Experimental results show that: a) The attack model is still able to generate efficient perturbations under both types of distributions, significantly increasing the scheduling completion time, indicating that the learned strategy has strong cross-distribution transferability. b) Defense model in the face of TPC-H and When subjected to perturbations, it can significantly reduce the performance gap before and after the attack, demonstrating robust adaptability to unknown DAG structures.
[0071] The above-described cross-distribution experimental results in the embodiments of the present invention show that: a) The unified attacker in the attack model of this invention is not only effective within the training distribution, but can also capture structural weaknesses in different combinatorial optimization problems; b) The defense model obtained through iterative adversarial retraining in this invention can not only resist perturbations under the training distribution, but also in Tsplib, CVRPlib, Gaussian distribution, TPC-H, and It maintains stable performance in diverse distributions.
[0072] This fully demonstrates that the attack and defense framework proposed in this invention has strong generalization ability and cross-segment robustness, and has important practical value in dealing with combinatorial optimization problems in real complex environments.
[0073] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
[0074] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A combined optimization solution system based on a general attack and defense framework, characterized in that, include: The multi-task modeling module is used to model the combined optimization problem of different tasks as a graph structure representation, and construct it as a graph structure Markov decision process according to the task type. The attack model is used to extract general graph feature representations from graph structures through a shared graph encoder, and generate perturbations that satisfy the constraints of the corresponding tasks through dedicated decoders and masking processes for different tasks. The defense model takes graph structure and node / edge features as input, and generates a decoding sequence that adapts to task perturbations and satisfies task constraints through a selected target defense solver, thereby achieving combinatorial optimization solution.
2. The combined optimization solution system based on a general attack and defense framework according to claim 1, characterized in that, The task types include the Asymmetric Traveling Salesman Problem (ASTP), the Vehicle Routing Problem (CVRP) with capacity constraints, and the Directed Acyclic Graph (DAG) scheduling problem. In the constructed Markov decision process: Each state in the state space includes the current graph structure, the node feature matrix, and the global matrix; Each action in the action space is a perturbation operation on the node-to-graph structure; the perturbation operation is determined according to the task type. The reward function represents the degree of performance degradation of the target defense solver after adding perturbations.
3. The combined optimization solution system based on a general attack and defense framework according to claim 2, characterized in that, The attack model is based on a reinforcement learning architecture and includes: The Actor network includes a shared encoder and a dedicated decoder. The shared encoder is used to extract general graph feature representations for optimization problems with different task combinations. The dedicated decoder is used to map the output of the shared encoder to the action space of the corresponding task and dynamically generate task legality masks corresponding to different task action spaces through masking, and then output the action probability distribution after masking. The Critic network is used to learn the state-value function for different tasks based on the action probability distribution output by the Actor network, and then feed it back to the Actor network. When the task is the Asymmetric Traveling Salesman Problem (ASTP), the learned state value function is the value function for estimating the path length of the ASTP; when the task is the Vehicle Routing Problem (CVRP) with capacity constraints, the learned state value function is the value function for estimating the total travel distance of the vehicle routing problem; when the task is the Directed Acyclic Graph (DAG) scheduling problem, the learned state value function is the value function for estimating the completion time of the scheduling problem.
4. The combined optimization solution system based on a general attack and defense framework according to claim 1, characterized in that, The target defense solver is an end-to-end combinatorial optimization solver based on graph neural networks, MatNet.
5. A combinatorial optimization solution method, implemented based on the combinatorial optimization solution system according to any one of claims 1 to 3, characterized in that, Includes the following steps: S100. The combinatorial optimization problem of the task to be solved is modeled as a Markov decision process with a graph structure. S200. In the attack model, a general graph feature representation of different task combination optimization problems is extracted by a shared graph encoder. The general graph feature representation is mapped to the corresponding action space by a dedicated decoder for the task to be solved. The task legality mask of the action space is dynamically generated by mask processing. S300. In the defense model, the node / edge features of the graph structure are used as input. The selected target defense solver gradually generates a decoding sequence that adapts to task perturbations and satisfies task constraints, thereby achieving combinatorial optimization solution.
6. The combinatorial optimization solution method according to claim 5, characterized in that, In step S200, the general graph feature representation of optimization problems with different task combinations is extracted, including: S201. Extract features of the graph structure through a forward graph convolutional network, and extract features of the transposed graph structure through a backward graph convolutional network. S202. The extracted features are spliced together through a splicing layer; S203. The attention mechanism is used to calculate and extract the general feature representation of the spliced features through the global attention pooling layer.
7. The combinatorial optimization solution method according to claim 5, characterized in that, In step S200, mapping the general graph feature representation to the corresponding action space includes: S204. For any pair of nodes in the graph structure, concatenate the embedding vectors of the two nodes and the general graph feature representation. Among them, node pairs represent perturbation operations on the graph structure; S205. Based on the task type, the spliced features are used to generate the original score of the node pair through a multilayer perceptron; S206. The original scores of all node pairs are converted into corresponding probability distributions by using the Softmax activation function to obtain the corresponding action space. Specifically, when the task is the Asymmetric Traveling Salesman Problem (ASTP), an attention mechanism with query-key dot product is used to calculate the original scores of connectivity and path cost for edge replacement / perturbation. When the task is the Vehicle Routing Problem (CVRP) with capacity constraints, an attention mechanism is used to further calculate the node embedding vectors in the vehicle path, and then the demand node embedding vector is added to the coordinate node embedding vector to obtain the original score considering customer demand and vehicle remaining capacity. When the task is the Directed Acyclic Graph (DAG) scheduling problem, context information is enhanced by splicing, the embedding vectors of selected nodes are paired with the embedding vectors of other nodes, and the original score considering whether or not topological constraints are generated is obtained by using dynamic masking.
8. The combinatorial optimization solution method according to claim 5, characterized in that, In step S200, the formula for generating the task validity mask of the action space is: In the formula, This indicates the new policy after weighted normalization in state. Select action The probability, Indicates the old strategy in state Select action The probability, This indicates the old strategy in the action state. Select action The probability, Indicates the action currently being evaluated. Indicates the current state. Indicates the legality of the action, when For illegal actions, when This was a legal action at the time.
9. The combinatorial optimization solution method according to claim 5, characterized in that, In step S200, the training loss function of the attack model for: In the formula, Indicates the strategy loss. Represents the loss of the value function. Let represent the expectation for step t. Represents the entropy function. Indicates the state at step t The following strategy Represents the Actor network parameters. Indicates the parameters of the Critic network. Represents the state at step t. , Represents the policy distribution. and These represent the first and second weight hyperparameters, respectively. This represents the function that takes the minimum value. This represents the ratio of the probability of the new strategy to the probability of the old strategy. This represents the estimation of the advantage function. This represents the clipping function. Indicates the clipping parameters. Represents the numerical stability constant. Indicates the parameters in the Actor network The current strategy is as follows: Indicates the old strategy, This represents the action at step t. This represents the state at step t. This represents the output of the Critic network for the corresponding task. This represents the batch-normalized reward. Indicates a reward. and These represent the mean and standard deviation of the current batch of rewards, respectively.
10. The combinatorial optimization solution method according to claim 5, characterized in that, In step S300, the target defense solver is trained using an iterative adversarial retraining method, including: The original dataset is perturbed using a trained attack model to generate the first adversarial example; the perturbation satisfies the legality constraints of different tasks. The original dataset is mixed with the first adversarial example to obtain a hybrid dataset; The target defense solver is iteratively trained using a mixed dataset to obtain the defense model; In each iteration of training, the defense model obtained from the previous training is used as the attack target. New adversarial examples are generated using the attack model, and a new hybrid dataset is constructed with the stable test benchmark dataset to train the defense model in the current iteration round.
Citation Information
Patent Citations
Knowledge graph embedding method based on relation path and double-layer attention
CN113806559A
Defense method for resisting bypass attack based on linear code mask and bit slicing technology
CN114048472A
Anti-patch attack-oriented defense method
CN119478567A
Black box directional countermeasure attack method based on general interaction mode
CN120281509A
Federal learning backdoor attack defense method based on multi-layer cooperative defense strategy
CN120434054A