Network adversarial decision-making method based on reinforcement learning
By modeling the network adversity process as partially observable Markov decision-making problems in network adversity, and combining graph neural networks and reinforcement learning algorithms, a network environment representation model and network adversity decision-making model are constructed. Using dynamic game optimization strategies, the problems of insufficient generalization ability for unknown environments and explosion in action space dimensions in the existing technology are solved, and stronger adaptability and generalization ability are achieved.
Patent Information
- Application Number
- CN202510077248.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing reinforcement learning has insufficient generalization ability to unknown environments in network confrontation, and the action space dimensions explode, making it difficult to adapt to dynamic changes in opponent strategies.
By modeling the network adversarial process as partially observable Markov decision-making problems, combining graph neural networks and reinforcement learning algorithms, a network environment representation model and a network adversarial decision-making model are constructed, and a dynamic game optimization strategy is used.
It improves the adaptability of the decision model to unknown network environments, reduces the dependence on reward functions, enhances the adaptability to high-dimensional action space and dynamic changes, and improves the generalization ability and convergence effect of the decision model.
Smart Images

Figure CN119966697A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network confrontation, and in particular to a network confrontation decision-making method based on reinforcement learning. Background Art
[0003] Applying reinforcement learning to network security is one of the current research hotspots, but there are some problems that need to be solved. First, the current reinforcement learning strategy optimization is mainly an end-to-end training paradigm. In this paradigm, the information used to improve decision-making ability is almost only rewards. The model tends to overfit to the training environment, resulting in insufficient generalization ability for unknown environments. However, unknown environments and incomplete information are important characteristics of network confrontation. Stronger generalization ability can enable decision models to make better decisions in unknown environments. Secondly, network confrontation involves the combination of multiple nodes and multiple technologies, and the node set may also change. Simply enumerating the combination of nodes and technologies in the strategy generation part will cause the action space dimension to explode, making it difficult for the training of the decision model to converge. Finally, in actual confrontation scenarios, the attacker and defender need to deal with opponents with different action strategies. If the attacker and defender interact and train independently in the network environment, it is difficult for the network confrontation decision model to adapt to the unknown opponent strategy. In summary, improving the adaptability of the decision model to the unknown network environment, solving the problem of high and dynamic changes in the action space dimension, and improving the adaptability of the decision model to different opponent strategies are the keys to building a high-quality network confrontation decision model. Summary of the invention
[0004] The purpose of the present invention is to provide a network confrontation decision-making method based on reinforcement learning, which mainly includes four parts: network confrontation process modeling, network environment characterization model, network confrontation decision-making model, dynamic game and strategy optimization. The network confrontation process is converted into a partially observable Markov decision problem, combined with a graph neural network and based on a reinforcement learning algorithm to realize the strategy generation of network confrontation, and the optimization of network confrontation strategy is realized through dynamic game between attack / defense agents.
[0005] To achieve the above object, the present invention provides a network confrontation decision-making method based on reinforcement learning, comprising the following steps:
[0006] S1: In the network environment, the attacker and defender are set as agents. Based on the assumption of Markov property and the characteristics of network confrontation, the six-tuple (S, A, T, R, Ω, O) is used to define the state space of the network environment, the agent action space, the agent observation space, the state transition probability, the conditional observation probability and the reward function.
[0007] S2: Establish a network environment representation model based on graph neural network;
[0008] S3: Combined with the network environment representation model, a network adversarial decision model is constructed as the attacker and defender;
[0009] S4: The attacker and defender optimize the corresponding network confrontation decision model through a dynamic game process;
[0010] S5: Output the optimized attacker and defender models.
[0011] Preferably, in step S1, the specific process of defining the sextuple is as follows:
[0012] The six-tuple (S, A, T, R, Ω, O) is defined as follows in the network confrontation scenario:
[0013] S is the state space of the network environment, and each state s t ∈S is expressed by network topology and node characteristics;
[0014] A is the action space of the agent, and each action a t ∈A contains the node identifier NID and the technology identifier TID;
[0015] Ω is the observation space of the agent, and each observation o t ∈Ω is the feedback received from the environment after the agent performs an action;
[0016] T(s t+1 |s t ,a t ) is the state transition probability, O(o t |s t+1 ,a t ) is the conditional observation probability; R(s t ,a t ) is the reward function, when s t+1 Compared to t Positive rewards are given when it is more beneficial to the agent, and negative rewards are given otherwise.
[0017] Preferably, in step S2, the specific process of constructing the network environment characterization model is as follows:
[0018] S21: Establish two auxiliary tasks, namely graph reconstruction and generation, to build a network environment representation model;
[0019] Define the state s of the network environment at time t t is a tuple consisting of the network topology and node features at that moment (M t ,X t ), where N is the number of nodes, K is the feature dimension, and M∈{0,1} N×N is the adjacency matrix of the network, is the feature matrix of all nodes, the observation o of the agent at time t t Represented as a binary group (M t ,X t ), the specific value depends on the predefined conditional observation probability O(o t |s t+1 ,a t );
[0020] S22: The network environment representation model mainly includes graph autoencoder GAE and variational graph autoencoder VGAE, which respectively implement parameter training of neural network through unsupervised tasks and self-supervised tasks of graphs;
[0021] S23: The unsupervised task is to reconstruct the state of the network environment, and the self-supervised task is to generate the future network state according to the Markov property;
[0022] During the training process, the encoder of GAE and the encoder responsible for generating the mean in VGAE share the graph attention layer; when there are two shared encoders, the parameters of the two GAT layers shared by GAE and VGAE are defined as and θ, the parameter of the GAT layer unique to VGAE is ρ;
[0023] The unsupervised task of GAE realizes the parameter training of the neural network by reconstructing the adjacency matrix and minimizing the reconstruction error, and generates the latent variable representation of the network environment through the encoder:
[0024]
[0025] In the above formula, The parameters are GAT network, GAT θ For a GAT network with parameter θ, use the dot product decoder to get the reconstructed adjacency matrix σ is the activation function:
[0026]
[0027] It's Z GAE The transposed matrix of the cross entropy loss function is used to calculate the true adjacency matrix M t and the reconstructed adjacency matrix The gap between the two is obtained, and the loss function L of GAE is obtained. GAE :
[0028]
[0029] In the above formula, m ij M t The i-th row and j-th column element of for The self-supervised task of VGAE is to generate the adjacency matrix of the next moment And maximize the lower bound of the evidence of variational inference to realize the parameter training of the neural network, and generate the distribution of hidden variables through the encoder:
[0030]
[0031] logσ=GAT ρ (GAT θ (M t ,X t ||D t ));
[0032]
[0033] In the above formula, D t is the Embedding vector of the action, μ i is the mean, is the variance, z i is the latent variable after sampling, GAT ρ is a GAT network with parameter ρ, is a normal distribution;
[0034] Sampling from the distribution to obtain the latent variable Z VGAE , the adjacency matrix is generated by the dot product decoder
[0035]
[0036] Finally, calculate the evidence lower bound of VGAE and obtain the loss function L of VGAE VGAE :
[0037]
[0038] In the above formula, β is the weight. Combining GAE and VGAE, the overall loss function of the network environment representation model is:
[0039]
[0040] In the above formula, α represents the weight;
[0041] S24: Through the above steps, the loss function threshold is set to obtain a trained network environment representation model, where the GAE encoder accepts zero-filled vectors and observations o t As input, the parameters are and θ, generating a network representation Z GAE .
[0042] Preferably, in step S3, the specific process is as follows:
[0043] S31: The network adversarial decision model is constructed using the proximal policy optimization model PPO, including the policy network Actor and the value network Critic;
[0044] S32: For the value network, input state s t and reward r t The front end of the value network is a graph neural network consisting of GAT and pooling operations. t and node feature X t The representation vector u of the graph is calculated. When L GAT layers are set, the formula is as follows:
[0045]
[0046] In the above formula, Pool l is the lth graph pooling layer, GAT l is the lth GAT layer;
[0047] Use the fully connected network to get the state value V(s) according to the graph representation vector u t ):
[0048] V(s t ) = MLP(u);
[0049] According to the Bellman formula and generalized advantage estimation, the TD error δ is obtained t and advantage value A(s t ,a t ):
[0050] δ t =r t +γ·V(s t+1 )-V(s t );
[0051]
[0052] In the above formula, γ is the discount factor, λ is an adjustable hyperparameter, and the value network outputs the advantage value A(s t ,a t );
[0053] For the policy network, it is necessary to consider the trajectory τ = {o1,o2,…,o t}, observe o t After the GAE encoder of the network environment representation model in step S2 GAE After that, it is input into a GRU model to obtain the node feature matrix H with time series informationt :
[0054] H t =GRU(Encoder GAE (M t ,X t ));
[0055] Due to the observation o t The number of nodes in the network is variable, so the attention mechanism is used to calculate the probability P(v i |τ), introduce a learnable parameter vector a and use it as the Query of the attention mechanism, then use the node feature vector h i As the key of the attention mechanism, the feature h of each node is calculated i The importance coefficient c of the parameter vector a i , and then use the Softmax function to get the normalized importance coefficient α i , and the coefficient α i The probability of being selected as a node is P(v i |τ), the formula is as follows:
[0056]
[0057]
[0058] In the above formula, W is the parameter matrix of the attention mechanism, and τ is the trajectory;
[0059] If the selected node is v i , then the next step is to consider the probability P(t j |v i ,τ), where the node v i Characteristics of h i Input to the fully connected network, and pass through the Softmax function at the end to get the probability P(t j |v i ,τ), and finally the attacker i Execution Technology j The probability of P(v i ,t j |τ), the formula is as follows:
[0060] P(v i ,t j |τ)=P(v i |τ)·P(t j |v i ,τ);
[0061]
[0062] In the above formula, is the jth bit of the vector obtained by MLP of the feature of the i-th node;
[0063] S33: Before use, the network decision model needs to be trained. The optimization process is as follows:
[0064] For the value network, its optimization goal is to minimize the error of the time difference method. The optimization goal is:
[0065]
[0066] In the above formula, φ and V φ are the parameters of the Critic network, Calculate for expectation;
[0067] For the policy network, its optimization goal is to maximize the return, so the optimization goal of the Actor is:
[0068]
[0069] In the above formula, θ is the parameter of the Actor network, π old For the old strategy;
[0070] In addition, in order to encourage reinforcement learning to explore new states, the entropy regularization term is introduced:
[0071]
[0072] In summary, the loss function of the network adversarial decision model is as follows:
[0073] L PPO (θ,φ)=c1·L Critic (φ)-c2·L Actor (θ)-c3·L Entripy (θ);
[0074] In the above formula, c1, c2, c3 are weight coefficients, L PPO (θ, φ) is the total loss function;
[0075] S34: Use the trained network decision model as the attacker and defender respectively.
[0076] Preferably, in step S4, the specific process is as follows:
[0077] S41: Establish a dynamic game environment, including the attacker and defender, and the corresponding checkpoint pools for the attacker and defender.
[0078] S42: Strategy optimization using dynamic game environment;
[0079] S421: Randomly initialize the attacker and defender, and set their respective checkpoint pools to empty;
[0080] S422: Let the attacker and the defender play multiple rounds of dynamic games in the network environment and collect battle data; randomly sample K checkpoints from the checkpoint pool each time, with the attacker sampling as RC and the defender sampling as BC, and let the attacker and the K BCs, and the defender and the K RCs, make decisions and execute them in sequence based on their own network confrontation decision-making model, the observation information and rewards returned by the network environment, so that the attacker and the defender make decisions and actions alternately until the goal of network attack or defense is achieved. The network environment state, rewards, and actions and observations of the intelligent agent generated during the dynamic game are the battle data required for model training;
[0081] S423: Use the battle data to update the network confrontation decision model of the agent, and store the updated agent as a checkpoint in the corresponding checkpoint pool;
[0082] S43: Repeat steps S421-S423. When the rewards obtained by the agent in the dynamic game tend to converge, the training is completed.
[0083] Preferably, in step S423, the process of updating the network confrontation decision model with battle data is as follows:
[0084] First initialize an empty experience replay pool Initialize K different network environments at the same time. In each training cycle, all trajectories obtained by interacting with all environments according to the current strategy form a trajectory set. And each of them (o j ,a j ,o j+1 ) tuple is stored in the experience replay pool Then use the experience replay pool respectively and trajectory collection The data is used to train the network environment representation model and the network adversarial decision-making model.
[0085] Therefore, the present invention adopts the above-mentioned network confrontation decision-making method based on reinforcement learning, which has the following advantages:
[0086] By analyzing the characteristics of network confrontation, this application models the network confrontation process as a Markov decision process POMDP, and the POMDP problem can be solved with the help of reinforcement learning methods. In view of the problem that traditional reinforcement learning methods are difficult to achieve good adaptability to unknown environments in network confrontation scenarios, this application proposes a network environment representation model based on graph neural networks, which realizes efficient representation of the network environment through unsupervised reconstruction tasks and self-supervised generation tasks of graphs, reduces dependence on reward functions, and improves the generalization ability of unknown environments. In view of the problems of variable node sets and high-dimensional action spaces, this application proposes a network confrontation decision model based on reinforcement learning. On the one hand, the network environment representation model is introduced in the Actor part of PPO to improve the generalization ability. On the other hand, the decision process is divided into two parts: node selection and technology selection, and the attention mechanism is used to process node sequences of indefinite length to improve the convergence effect of the model. Finally, in order to enable the decision model to cope with a variety of enemy action strategies, a strategy optimization framework based on dynamic game is proposed, which allows the intelligent agent to simulate confrontation with multiple defenders and attackers' Checkpoints, and trains the decision model based on the collected confrontation data to improve the decision-making ability of the intelligent agent.
[0087] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 This is an overall flow chart of a network confrontation decision-making method based on reinforcement learning in the present invention;
[0089] Figure 2 A schematic diagram of a decision-making process in a network confrontation decision-making method based on reinforcement learning of the present invention;
[0090] Figure 3 A schematic diagram of a network environment representation model in a network confrontation decision-making method based on reinforcement learning of the present invention;
[0091] Figure 4 Schematic diagram of a network confrontation decision model in a network confrontation decision method based on reinforcement learning of the present invention
[0092] Figure 5 It is a schematic diagram of a strategy optimization method based on dynamic game in a network confrontation decision-making method based on reinforcement learning of the present invention. DETAILED DESCRIPTION
[0093] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. The specific model specifications need to be selected and determined according to the actual specifications of the device, and the specific selection calculation method adopts the existing technology in the field, so it will not be described in detail.
[0094] Example
[0095] like Figure 1 As shown, the present invention provides a network confrontation decision-making method based on reinforcement learning, comprising the following steps:
[0096] S1: In the network environment, the attacker and defender are set as agents. Based on the assumption of Markov property and the characteristics of network confrontation, the six-tuple (S, A, T, R, Ω, O) is used to define the state space of the network environment, the agent action space, the agent observation space, the state transition probability, the conditional observation probability and the reward function.
[0097] The six-tuple (S, A, T, R, Ω, O) is defined as follows in the network confrontation scenario:
[0098] S is the state space of the network environment, and each state s t ∈S is expressed by network topology and node characteristics;
[0099] A is the action space of the agent, and each action a t ∈A contains the node identifier NID and the technology identifier TID;
[0100] Ω is the observation space of the agent, and each observation o t ∈Ω is the feedback received from the environment after the agent performs an action;
[0101] T(s t+1 |s t ,a t ) is the state transition probability, O(o t |s t+1 ,a t ) is the conditional observation probability; R(s t ,a t ) is the reward function, when s t+1 Compared to t Positive rewards are given when it is more beneficial to the agent, and negative rewards are given otherwise.
[0102] S2: Establish a network environment representation model based on graph neural network. The process is as follows:
[0103] S21: Establish two auxiliary tasks, namely, graph reconstruction and generation, and construct a network environment representation model;
[0104] Define the state s of the network environment at time t t is a tuple consisting of the network topology and node features at that moment (M t ,X t ), where N is the number of nodes, K is the feature dimension, and M∈{0,1} N×N is the adjacency matrix of the network, is the feature matrix of all nodes, the observation o of the agent at time t t Represented as a binary group (M t ,X t ), the specific value depends on the predefined conditional observation probability O(o t |s t+1 ,a t );
[0105] S22: The network environment representation model mainly includes graph autoencoder GAE and variational graph autoencoder VGAE, which respectively implement parameter training of neural network through unsupervised tasks and self-supervised tasks of graphs;
[0106] S23: The unsupervised task is to reconstruct the state of the network environment, and the self-supervised task is to generate the future network state according to the Markov property;
[0107] During the training process, the encoder of GAE and the encoder responsible for generating the mean in VGAE share the graph attention layer; when there are two shared encoders, the parameters of the two GAT layers shared by GAE and VGAE are defined as and θ, the parameter of the GAT layer unique to VGAE is ρ;
[0108] The unsupervised task of GAE realizes the parameter training of the neural network by reconstructing the adjacency matrix and minimizing the reconstruction error, and generates the latent variable representation of the network environment through the encoder:
[0109]
[0110] In the above formula, The parameters are GAT network, GAT θ For a GAT network with parameter θ, use the dot product decoder to get the reconstructed adjacency matrix σ is the activation function:
[0111]
[0112] It's Z GAEThe transposed matrix of , the true adjacency matrix M is calculated by the cross entropy loss function t and reconstruct the adjacency matrix The gap between the two is obtained, and the loss function L of GAE is obtained. GAE :
[0113]
[0114] In the above formula, m ij M t The i-th row and j-th column element of for The self-supervised task of VGAE is to generate the adjacency matrix of the next moment And maximize the lower bound of the evidence of variational inference to realize the parameter training of the neural network, and generate the distribution of hidden variables through the encoder:
[0115]
[0116] logσ=GAT ρ (GAT θ (M t ,X t ||D t ))
[0117]
[0118] In the above formula, D t is the Embedding vector of the action, μ i is the mean, is the variance, z i is the latent variable after sampling, GAT ρ is a GAT network with parameter ρ, is a normal distribution;
[0119] Sampling from the distribution to obtain the latent variable Z VGAE , the adjacency matrix is generated by the dot product decoder
[0120]
[0121] Finally, calculate the evidence lower bound of VGAE and obtain the loss function L of VGAE VGAE :
[0122]
[0123] In the above formula, β is the weight. Combining GAE and VGAE, the overall loss function of the network environment representation model is:
[0124]
[0125] In the above formula, α represents the weight;
[0126] S24: Through the above steps, the loss function threshold is set to obtain a trained network environment representation model, where the GAE encoder accepts zero-filled vectors and observations o t As input, the parameters are and θ, generating a network representation Z GAE .
[0127] S3: Combined with the network environment representation model, a network confrontation decision model is constructed as the attacker and defender. The specific process is as follows:
[0128] S31: The network adversarial decision model is constructed using the proximal policy optimization model PPO, including the policy network Actor and the value network Critic;
[0129] S32: For the value network, input state s t and reward r t The front end of the value network is a graph neural network consisting of GAT and pooling operations. t and node feature X t Calculate the representation vector u of the graph. Assume there are L GAT layers:
[0130]
[0131]
[0132] In the above formula, Pool l is the lth graph pooling layer, GAT l is the lth GAT layer; the state value V(s) is obtained based on the graph representation vector u using a fully connected network. t ):
[0133] V(s t ) = MLP(u);
[0134] According to the Bellman formula and generalized advantage estimation, the TD error δ is obtained t and advantage value A(s t ,a t ):
[0135] δ t =r t +γ·V(s t+1 )-V(s t );
[0136]
[0137] In the above formula, γ is the discount factor, λ is an adjustable hyperparameter, and the value network outputs the advantage value A(s t ,a t );
[0138] For the policy network, it is necessary to consider the trajectory τ = {o1,o2,…,o t}, observe o t After the GAE encoder of the network environment representation model in step S2 GAE After that, it is input into a GRU model to obtain the node feature matrix H with time series information t :
[0139] H t =GRU(Encoder GAE (M t ,X t ))
[0140] Due to the observation o t The number of nodes in the network is variable, so the attention mechanism is used to calculate the probability P(v i |τ), introduce a learnable parameter vector a and use it as the Query of the attention mechanism, then use the node feature vector h i As the key of the attention mechanism, the feature h of each node is calculated i The importance coefficient c of the parameter vector a i , and then use the Softmax function to get the normalized importance coefficient α i , and the coefficient α i The probability of being selected as a node is P(v i |τ), the formula is as follows:
[0141]
[0142]
[0143] In the above formula, W is the parameter matrix of the attention mechanism, τ is the trajectory; if the selected node is v i , then the next step is to consider the probability P(t j |v i ,τ), where the node v i Characteristics of h i Input to the fully connected network, and pass through the Softmax function at the end to get the probability P(t j |v i ,τ), and finally the attacker i Execution Technology jThe probability of P(v i ,t j |τ), the formula is as follows:
[0144] P(v i ,t j |τ)=P(v i |τ)·P(t j |v i ,τ);
[0145]
[0146] In the above formula, is the jth bit of the vector obtained by MLP of the feature of the i-th node;
[0147] S33: Before use, the network decision model needs to be trained. The optimization process is as follows:
[0148] For the value network, its optimization goal is to minimize the error of the time difference method. The optimization goal is:
[0149]
[0150] In the above formula, φ and V φ are the parameters of the Critic network, Calculate for expectation;
[0151] For the policy network, its optimization goal is to maximize the return, so the optimization goal of the Actor is:
[0152]
[0153] In the above formula, θ is the parameter of the Actor network, π old For the old strategy;
[0154] In addition, in order to encourage reinforcement learning to explore new states, the entropy regularization term is introduced:
[0155]
[0156] In summary, the loss function of the network adversarial decision model is as follows:
[0157] L PPO (θ,φ)=c1·L Critic (φ)-c2·L Actor (θ)-c3·L Entropy (θ)
[0158] In the above formula, c1, c2, c3 are weight coefficients, L PPO (θ, φ) is the total loss function;
[0159] S34: Use the trained network decision model as the attacker and defender respectively.
[0160] S4: The attacker and defender optimize the corresponding network confrontation decision model through a dynamic game process. The specific process is as follows:
[0161] S41: Establish a dynamic game environment, including the attacker and defender, and the corresponding checkpoint pools for the attacker and defender.
[0162] S42: Strategy optimization using dynamic game environment;
[0163] S421: Randomly initialize the attacker and defender, and set their respective checkpoint pools to empty;
[0164] S422: Let the attacker and the defender play multiple rounds of dynamic games in the network environment and collect battle data; randomly sample K checkpoints from the checkpoint pool each time, with the attacker sampling as RC and the defender sampling as BC, and let the attacker and the K BCs, and the defender and the K RCs, make decisions and execute them in sequence based on their own network confrontation decision-making model, the observation information and rewards returned by the network environment, so that the attacker and the defender make decisions and actions alternately until the goal of network attack or defense is achieved. The network environment state, rewards, and actions and observations of the intelligent agent generated during the dynamic game are the battle data required for model training;
[0165] S423: Using the battle data to update the network confrontation decision model of the intelligent agent, the process of updating the network confrontation decision model with the battle data is as follows:
[0166] First initialize an empty experience replay pool Initialize K different network environments at the same time. In each training cycle, all trajectories obtained by interacting with all environments according to the current strategy form a trajectory set. And each of them (o j ,a j ,o j+1 ) tuple is stored in the experience replay pool Then use the experience replay pool respectively and trajectory collection The data is used to train the network environment representation model and the network adversarial decision model. The corresponding algorithms are executed as follows:
[0167] Input: K different network environments The learning rates α1, α2 and the loss function weights c1, c2, c3 are implemented as follows:
[0168] Initialize the experience replay pool Do the following until convergence:
[0169] Initialize the trajectory collection For each environment Strategy π θ Interactively get the trajectory And the trajectory All (o j ,a j ,o j+1 ) tuple is stored in the experience replay pool
[0170] In each training cycle of the network environment representation model, the experience replay pool Sampling to get small batches of experience Calculate L based on small batch experience b feat :
[0171]
[0172] During the reinforcement learning training cycle, from the trajectory cache Sampling to get small batches of trajectories Calculate L based on the mini-batch trajectory d Actor , L Critic and L Entropy :
[0173]
[0174] Store the updated agent as a Checkpoint in the corresponding Checkpoint pool;
[0175] S43: Repeat steps S421-S423. When the rewards obtained by the agent in the dynamic game tend to converge, the training is completed.
[0176] Therefore, the present invention adopts a network confrontation decision-making method based on reinforcement learning. By analyzing the characteristics of network confrontation, this application models the network confrontation process as a Markov decision process POMDP, and the POMDP problem can be solved with the help of reinforcement learning methods. In view of the problem that traditional reinforcement learning methods are difficult to achieve good adaptability to unknown environments in network confrontation scenarios, this application proposes a network environment representation model based on graph neural networks, which realizes efficient representation of the network environment through unsupervised reconstruction tasks and self-supervised generation tasks of graphs, reduces dependence on reward functions, and improves the generalization ability of unknown environments. In view of the problems of variable node sets and high-dimensional action spaces, this application proposes a network confrontation decision-making model based on reinforcement learning. On the one hand, the network environment representation model is introduced in the Actor part of PPO to improve the generalization ability. On the other hand, the decision-making process is split into two parts: node selection and technology selection, and the attention mechanism is used to process node sequences of indefinite length to improve the convergence effect of the model. Finally, in order to enable the decision-making model to cope with a variety of enemy action strategies, a strategy optimization framework based on dynamic game is proposed, which allows the intelligent agent to simulate confrontation with multiple defenders and attackers' checkpoints. The decision-making model is trained based on the collected confrontation data to improve the decision-making ability of the intelligent agent.
[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A network confrontation decision-making method based on reinforcement learning, characterized by: The following steps are involved: S1: In the network environment, the attacker and defender are set as agents. Based on the assumption of Markov property and the characteristics of network confrontation, the six-tuple (S, A, T, R, Ω, O) is used to define the state space of the network environment, the agent action space, the agent observation space, the state transition probability, the conditional observation probability and the reward function. S2: Establish a network environment representation model based on graph neural network; S3: Combined with the network environment representation model, a network adversarial decision model is constructed as the attacker and defender; S4: The attacker and defender optimize the corresponding network confrontation decision model through a dynamic game process; S5: Output the optimized attacker and defender models.
2. The network confrontation decision-making method based on reinforcement learning according to claim 1 is characterized in that: In step S1, the specific process of defining the six-tuple is as follows: The six-tuple (S, A, T, R, Ω, O) is defined as follows in the network confrontation scenario: S is the state space of the network environment, and each state s t ∈S is expressed by network topology and node characteristics; A is the action space of the agent, and each action a t ∈A contains the node identifier NID and the technology identifier TID; Ω is the observation space of the agent, and each observation o t ∈Ω is the feedback received from the environment after the agent performs an action; T(s t+1 |s t ,a t ) is the state transition probability, O(o t |s t+1 ,a t ) is the conditional observation probability; R(s t ,a t ) is the reward function, when s t+1 Compared to t Positive rewards are given when it is more beneficial to the agent, and negative rewards are given otherwise.
3. The network confrontation decision-making method based on reinforcement learning according to claim 2 is characterized in that: In step S2, the specific process of constructing the network environment representation model is as follows: S21: Establish two auxiliary tasks, namely graph reconstruction and generation, to build a network environment representation model; Define the state s of the network environment at time t t is a tuple consisting of the network topology and node features at that moment (M t ,X t ), where N is the number of nodes, K is the feature dimension, and M∈{0,1} N×N is the adjacency matrix of the network, is the feature matrix of all nodes, the observation o of the agent at time t t Represented as a binary group (M t ,X t ), the specific value depends on the predefined conditional observation probability O(o t |s t+1 ,a t ); S22: The network environment representation model mainly includes graph autoencoder GAE and variational graph autoencoder VGAE, which respectively implement parameter training of neural network through unsupervised tasks and self-supervised tasks of graphs; S23: The unsupervised task is to reconstruct the state of the network environment, and the self-supervised task is to generate the future network state according to the Markov property; During the training process, the encoder of GAE and the encoder responsible for generating the mean in VGAE share the graph attention layer; when there are two shared encoders, the parameters of the two GAT layers shared by GAE and VGAE are defined as and θ, the parameter of the GAT layer unique to VGAE is ρ; The unsupervised task of GAE realizes the parameter training of the neural network by reconstructing the adjacency matrix and minimizing the reconstruction error, and generates the latent variable representation of the network environment through the encoder: In the above formula, The parameters are GAT network, GAT θ For a GAT network with parameter θ, use the dot product decoder to get the reconstructed adjacency matrix σ is the activation function: It's Z GAE The transposed matrix of the cross entropy loss function is used to calculate the true adjacency matrix M t and the reconstructed adjacency matrix The gap between the two is obtained, and the loss function L of GAE is obtained. GAE : In the above formula, m ij M t The i-th row and j-th column element of for The self-supervised task of VGAE is to generate the adjacency matrix of the next moment And maximize the lower bound of the evidence of variational inference to realize the parameter training of the neural network, and generate the distribution of hidden variables through the encoder: logσ=GAT ρ (GAT θ (M t ,X t ||D t )); In the above formula, D t is the Embedding vector of the action, μ i is the mean, is the variance, z i is the latent variable after sampling, GAT ρ is a GAT network with parameter ρ, is a normal distribution; Sampling from the distribution to obtain the latent variable Z VGAE , the adjacency matrix is generated by the dot product decoder Finally, calculate the evidence lower bound of VGAE and obtain the loss function L of VGAE VGAE : In the above formula, β is the weight. Combining GAE and VGAE, the overall loss function of the network environment representation model is: In the above formula, α represents the weight; S24: Through the above steps, the loss function threshold is set to obtain a trained network environment representation model, in which the GAE encoder accepts zero-filled vectors and observations o t As input, the parameters are and θ, generating a network representation Z GAE .
4. The network confrontation decision-making method based on reinforcement learning according to claim 2 is characterized in that: In step S3, the specific process is as follows: S31: The network adversarial decision model is constructed using the proximal policy optimization model PPO, including the policy network Actor and the value network Critic; S32: For the value network, input state s t and reward r t The front end of the value network is a graph neural network consisting of GAT and pooling operations. t and node feature X t The representation vector u of the graph is calculated. When L GAT layers are set, the formula is as follows: In the above formula, Pool l is the lth graph pooling layer, GAT l is the lth GAT layer; Use the fully connected network to get the state value V(s) according to the graph representation vector u t ): V(s t )=MLP(u); According to the Bellman formula and generalized advantage estimation, the TD error δ is obtained t and advantage value A(s t ,a t ): δ t =r t +γ·V(s t+1 )-V(s t ); In the above formula, γ is the discount factor, λ is an adjustable hyperparameter, and the value network outputs the advantage value A(s t ,a t ); For the policy network, it is necessary to consider the trajectory τ = {o1,o2,…,o t }, observe o t After the GAE encoder of the network environment representation model in step S2 GAE After that, it is input into a GRU model to obtain the node feature matrix H with time series information t : H t =GRU(Encoder GAE (M t ,X t )); Due to the observation o t The number of nodes in the network is variable, so the attention mechanism is used to calculate the probability P(v i |τ), introduce a learnable parameter vector a and use it as the Query of the attention mechanism, then use the node feature vector h i As the key of the attention mechanism, the feature h of each node is calculated i The importance coefficient c of the parameter vector a i , and then use the Softmax function to get the normalized importance coefficient α i , and the coefficient α i The probability of being selected as a node is P(v i |τ), the formula is as follows: In the above formula, W is the parameter matrix of the attention mechanism, and τ is the trajectory; If the selected node is v i , then the next step is to consider the probability P(t j |v i ,τ), where the node v i Characteristics of h i Input to the fully connected network, and pass through the Softmax function at the end to get the probability P(t j |v i ,τ), and finally the attacker i Execution Technology j The probability of P(v i ,t j |τ), the formula is as follows: P(v i ,t j |τ)=P(v i |τ)·P(t j |v i ,τ); In the above formula, is the jth bit of the vector obtained by MLP of the feature of the i-th node; S33: Before use, the network decision model needs to be trained. The optimization process is as follows: For the value network, its optimization goal is to minimize the error of the time difference method. The optimization goal is: In the above formula, φ and V φ are the parameters of the Critic network, Calculate for expectation; For the policy network, its optimization goal is to maximize the return, so the optimization goal of the Actor is: In the above formula, θ is the parameter of the Actor network, π old For the old strategy; In addition, in order to encourage reinforcement learning to explore new states, the entropy regularization term is introduced: In summary, the loss function of the network adversarial decision model is as follows: L PPO (θ,φ)=c1·L Critic (φ)-c2·L Actor (θ)-c3·L Entripy (i); In the above formula, c1, c2, c3 are weight coefficients, L PPO (θ, φ) is the total loss function; S34: Use the trained network decision model as the attacker and defender respectively.
5. The network confrontation decision-making method based on reinforcement learning according to claim 2 is characterized in that: In step S4, the specific process is as follows: S41: Establish a dynamic game environment, including the attacker and defender, and the corresponding checkpoint pools for the attacker and defender. S42: Strategy optimization using dynamic game environment; S421: Randomly initialize the attacker and defender, and set their respective checkpoint pools to empty; S422: Let the attacker and the defender play multiple rounds of dynamic games in the network environment and collect battle data; randomly sample K checkpoints from the checkpoint pool each time, with the attacker sampling as RC and the defender sampling as BC, and let the attacker and the K BCs, and the defender and the K RCs, make decisions and execute them in sequence based on their own network confrontation decision-making model, the observation information and rewards returned by the network environment, so that the attacker and the defender make decisions and actions alternately until the goal of network attack or defense is achieved. The network environment state, rewards, and actions and observations of the intelligent agent generated during the dynamic game are the battle data required for model training; S423: Using the battle data to update the network confrontation decision model of the agent, and storing the updated agent as a checkpoint in the corresponding checkpoint pool; S43: Repeat steps S421-S423. When the rewards obtained by the agent in the dynamic game tend to converge, the training is completed.
6. The network confrontation decision-making method based on reinforcement learning according to claim 5 is characterized in that: In step S423, the process of updating the network confrontation decision model with battle data is as follows: First initialize an empty experience replay pool Initialize K different network environments at the same time. In each training cycle, all trajectories obtained by interacting with all environments according to the current strategy form a trajectory set. And each of them (o j ,a j ,o j+1 ) tuple is stored in the experience replay pool Then use the experience replay pool separately and trajectory collection The data is used to train the network environment representation model and the network adversarial decision-making model.
Citation Information
Patent Citations
Semi-supervised graph representation learning method based on fusion of transfer learning and deep learning and device thereof
CN112990295A
Multi-agent attack and defense decision-making method based on deep reinforcement learning
CN115544898A
Terminal side alarm traceability graph filtering method based on graph neural network clustering
CN117349744A
Network processing analysis method and device, electronic equipment and storage medium
CN117376164A
Network risk analysis method and system based on multilevel game model
CN119254483A
Cited By
Defense strategy self-generation method and system for intelligent device cluster
CN120768612A
Defensive strategy self-generation method and system for intelligent device cluster
CN120768612B
Power distribution network dispatching method, computer equipment and storage medium
CN120855523A
Power distribution network dispatching method, computer device and storage medium
CN120855523B
Multi-task acoustic decoy strategy migration method and system based on meta-reinforcement learning
CN121457561A