APT node isolation defense method based on population training

By establishing a game model of APT intelligent lateral movement and a defense strategy based on population training, the problem of lateral movement in the existing technology that is difficult to effectively defend against APT attackers is solved, and a more efficient and robust defense effect is achieved.

CN119995950AActive Publication Date: 2025-05-13GUANGZHOU UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510061472.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and defend against the lateral movement of advanced persistent threat (APT) attackers in institutional networks, making it difficult for defenders to detect and isolate attackers in time.

Method used

The APT node isolation defense method based on population training is adopted. By establishing a game model of APT intelligent lateral movement, the attacker's initial position and lateral movement method are determined, the defender and attacker's strategic population is obtained, and the strategy is updated through the metagame solver until the Nash equilibrium is reached.

Benefits of technology

It improves the robustness and efficiency of the defender's defense strategy, can detect and isolate APT attackers more effectively, and reduces the attacker's probability of success and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995950A_ABST
    Figure CN119995950A_ABST
Patent Text Reader

Abstract

The invention provides an APT node isolation defense method based on population training, and relates to the technical field of network security, and the method comprises the steps: constructing an APT intelligent transverse movement game model, firstly determining an initial position and a target node of an attacker, then setting a transverse movement mode, constructing a strategy and a strategy space of the attacker and a defender, and constructing an APT intelligent transverse movement game model; and respectively calculating a first value function aiming at the target node and a second value function based on the strategy of the two parties. Then, defining strategy populations of both parties and a meta-game form, and generating a meta-game income matrix; after the strategy population and the meta-strategy are initialized, attack and defense benefits are obtained through gaming, and a meta-gaming benefit matrix is updated. And solving by utilizing a meta game solver to obtain a double-party meta strategy, calculating a new Q value of each target node, and updating an attacker strategy population. Then, training a defender response strategy, adding the defender response strategy into a strategy population of the defender response strategy for iterative updating, and obtaining Nash equilibrium; according to the invention, the dynamic game and intelligent optimization of the APT attack and defense strategy are realized, and the network security defense capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to an APT node isolation defense method based on population training. Background Art

[0002] Advanced persistent threats (APTs) have long been considered one of the most notorious types of cyberattacks. Compared with conventional cyberattacks, APTs have a high penetration success rate because well-funded APT attackers can exploit the inherent weaknesses of human nature and conduct meticulous social engineering attacks to successfully hijack host nodes in an organization's network. Once an APT attacker has established a foothold in an organization's intranet, he will try to approach the target nodes, the most valuable nodes in the organization, such as data centers and servers, through lateral movement, with the goal of stealing critical data and information from these nodes. Due to its high concealment, defenders often find it difficult to detect APTs, making APTs one of the most serious threats to modern organizations.

[0003] Existing research has shown that it is necessary to consider lateral movement when optimizing the defender's APT defense strategy, because through lateral movement, APT attackers can effectively expand the scope of attack. However, most existing studies assume that the attacker's lateral movement follows a fixed pattern, such as an uncompromised node will be compromised at a certain rate due to communication with a compromised node. In actual scenarios, this assumption may still underestimate the capabilities of APT attackers to a certain extent. Although some studies consider the controllable rate of lateral movement, in reality APT attackers may adopt smarter strategies.

[0004] Specifically, consider the worst case scenario for the defender - the APT attacker chooses the shortest path, starting from the initially compromised node, and approaches the organization's target node along this path. From the perspective of the APT attacker, the attacker only needs to control fewer nodes, thus reducing the probability of being detected and cleared, and also reducing the cost of the attack. On the other hand, a longer lateral movement path is beneficial to the defender because the defender can monitor the organization's network through global information, thereby increasing the success rate of APT detection.

[0005] Therefore, it is urgent to provide a solution to improve the above problems. Summary of the invention

[0006] The purpose of the present invention is to provide an APT node isolation defense method based on population training, so as to improve the problem of low defense effectiveness of existing network security technology.

[0007] The present invention provides an APT node isolation defense method based on population training, which adopts the following technical solutions:

[0008] Establish a game model for APT intelligent lateral movement, and determine the attacker's initial position and lateral movement method based on the target node;

[0009] Based on the game model, establish the attacker's strategy and strategy space, the defender's strategy and strategy space, obtain a first value function based on the defender's strategy and the target node, and obtain a second value function based on the defender's strategy and the attacker's strategy;

[0010] Obtain the form of the defender and attacker strategy populations, and obtain the form of the defender and attacker meta-games, and generate a meta-game payoff matrix;

[0011] Initialize the strategy population and meta-strategy of the attacker and defender, play games with the defender and attacker based on the meta-strategy to obtain attack and defense benefits, and fill the attack and defense benefits into the meta-game benefit matrix for updating to obtain an updated meta-game benefit matrix;

[0012] Solving the updated meta-game payoff matrix based on the meta-game solver to obtain the meta-strategies of the defender and the attacker, and calculating the new Q value of each target node to obtain the updated strategy population of the attacker;

[0013] The response strategy of the defender is trained based on the strategy population to obtain a trained new strategy, which is added to the strategy population of the defender for iterative updating until the current number of iterations is greater than the preset maximum number of iterations, and the updating is stopped to obtain the Nash equilibrium.

[0014] The present invention provides an APT node isolation defense method based on population training, which has the beneficial effects that: the present invention applies game theory to APT lateral movement behavior modeling, considers the worst case from the defender's perspective: APT attackers may adopt the most efficient attack strategy, establishes a game theory mathematical model, and characterizes the strategic interaction process between APT attackers and defenders. When constructing a defense strategy, it may be inefficient to always select the most important target nodes (i.e., nodes in V′) for security review, because careful review of these nodes may require higher costs. In the present invention, the defender determines the nodes for security review based on the current state, and the detection is more efficient. In addition, since the worst case is taken into account, the defender's defense strategy can be made more robust.

[0015] Optionally, the process of establishing the game model of APT intelligent lateral movement includes: constructing an undirected graph of an organization, the undirected graph consisting of multiple nodes and multiple edges, the nodes communicating with each other through the organization's intranet, the nodes including important nodes and ordinary nodes in the organization, wherein the important nodes are the target nodes of the APT attackers, consisting of data centers and servers, and are used to store the organization's core data, perform data analysis tasks or control critical infrastructure, and the remaining nodes are ordinary nodes.

[0016] Optionally, the mathematical expression of the first value function is:

[0017]

[0018] The mathematical expression of the second value function is:

[0019]

[0020] in, and The distribution is the rewards of the defender and the attacker at the time 0≤t≤T, Ε represents the expected calculation, T represents the maximum time step, represents the first value function of the defender, represents the attacker's first value function, V D (π D ,π A ) represents the second value function of the defender, V A (π D ,π A ) represents the attacker’s second value function.

[0021] Optionally, the process of obtaining the form of the defender and attacker strategy populations includes:

[0022] Use the defender strategy population Indicates that represents the defender's first strategy, represents the defender’s K+1th strategy, where K ≥ 0;

[0023] The attacker strategy population is used Indicates that represents the attacker's first strategy, represents the attacker’s K+1th strategy, where K≥0.

[0024] Optionally, when obtaining the defender's meta-game form, the meta-game is a matrix game where the defender's strategy set is represented as Contains K+1 strategies, among which, represents a meta-strategy of the defender, which is expressed as a probability distribution of the defender's strategy. The conditions that are satisfied include:

[0025]

[0026] in, Indicates the selection of the i-th strategy probability.

[0027] Optionally, when obtaining the attacker's meta-game form, the attacker's strategy set is expressed as Contains K+1 strategies, using represents a meta-strategy of the attacker, which represents a probability distribution about the attacker's strategy. The conditions that are satisfied include:

[0028]

[0029] in, Indicates the selection of the jth strategy probability.

[0030] Optionally, the process of initializing the attacker's strategy includes: obtaining a Q value of the target node, and obtaining the attacker's strategy based on the Q value, wherein:

[0031] The mathematical expression of the Q value is:

[0032]

[0033] The attacker's strategy The mathematical expression is:

[0034]

[0035] When the Q values ​​of all target nodes are initialized to 0, the attacker's strategy becomes a uniform strategy, and the mathematical expression is:

[0036]

[0037] Where |V′| represents the number of target nodes, represents the attacker’s uniform strategy, represents the first value function of the attacker.

[0038] Optionally, the process of obtaining the updated attacker's strategy population includes:

[0039]

[0040] in, represents the attacker’s strategy population, represents the attacker's new response strategy, Represents a union operation.

[0041] Optionally, the process of training the defender's response strategy based on the strategy population to obtain a trained new strategy and adding it to the defender's strategy population includes:

[0042]

[0043] in, represents the defender’s strategy population, represents the trained new strategy, Represents a union operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flow chart of an APT node isolation defense method based on population training provided by the present invention;

[0045] Figure 2 It is a result graph of an experiment conducted on a synthetic network including 100 nodes provided by the present invention;

[0046] Figure 3 This is a result diagram of an experiment conducted on a real communication network including 100 nodes provided by the present invention;

[0047] Figure 4 This is a result diagram of an experiment conducted on a real researcher collaboration network containing 100 nodes provided by the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be understood by people with general skills in the field to which the present invention belongs. "Including" and similar words used in this article mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.

[0049] The embodiment of the present invention provides an APT node isolation defense method based on population training, including:

[0050] S1. Establish a game model for APT intelligent lateral movement, and determine the attacker's initial position and lateral movement method based on the target node;

[0051] S2. Establishing the attacker's strategy and strategy space, the defender's strategy and strategy space based on the game model, obtaining a first value function based on the defender's strategy and the target node, and obtaining a second value function based on the defender's strategy and the attacker's strategy;

[0052] S3, obtaining the form of the defender and attacker strategy populations, and obtaining the form of the defender and attacker meta-games, and generating a meta-game payoff matrix;

[0053] S4, initializing the strategy population and meta-strategy of the attacker and defender, playing games with the defender and attacker based on the same strategy to obtain attack and defense benefits, and filling the attack and defense benefits into the meta-game benefit matrix for updating to obtain an updated meta-game benefit matrix;

[0054] S5. Solve the updated meta-game payoff matrix based on the meta-game solver to obtain the meta-strategies of the defender and the attacker, and calculate the new Q value of each target node to obtain the updated strategy population of the attacker;

[0055] S6. Based on the strategy population, the defender's response strategy is trained to obtain a trained new strategy, which is added to the defender's strategy population for iterative updating until the current number of iterations is greater than the preset maximum number of iterations, and the updating is stopped to obtain a Nash equilibrium.

[0056] In some embodiments, the process of executing step S1 to establish a game model of APT intelligent lateral movement includes: constructing an undirected graph of an organization, the undirected graph consists of multiple nodes and multiple edges, the nodes communicate with each other through the intranet of the organization, the nodes include important nodes and ordinary nodes in the organization, wherein the important nodes are the target nodes of the APT attacker, composed of data centers and servers, and are used to store the core data of the organization, perform data analysis tasks or control critical infrastructure, and the remaining nodes are ordinary nodes.

[0057] Specifically, the intranet of an organization is represented by a graph G = (V, E), where V = {v 1 ,…,v N} is the node set, and E is the edge set. i ,v j}∈E indicates that node i and node j communicate through the organization's intranet. V′∈V indicates important nodes in the organization, and other nodes in the organization are ordinary nodes. express.

[0058] Furthermore, when executing step S1, the process of determining the initial position and lateral movement mode of the attacker based on the target node includes:

[0059] use Represents the attacker's starting attack node. At t = 0, the attacker selects a target node Then by arrive The shortest path for lateral movement is expressed as

[0060] In some embodiments, the process of establishing the attacker's strategy and strategy space in step S2 includes:

[0061] At t = 0, the attacker follows the strategy π A : Select a target node and then move horizontally according to the shortest path from the initial node to the selected target node. A is the attacker’s strategy space, then π A ∈Π A .

[0062] Further, the process of establishing the defender's strategy and strategy space includes:

[0063] At each time step, the defender selects a node for isolation review based on the current state to determine whether the node is under APT attack. The system state observed by the defender consists of two parts: the current time step and the historical action sequence, namely Use S D represents the space of states observed by the defender, then the defender’s strategy is defined as π D :S D →Δ(V). Let Π D is the defender’s strategy space, then π D ∈Π D In the present invention, the defender's strategy is represented using a neural network.

[0064] In some embodiments, when executing step S2, given the defender strategy π D and the target node When , the value function of the attacker and the defender is defined as the first value function, and the mathematical expression of the first value function is:

[0065]

[0066] in, and The distribution is the rewards of the defender and the attacker at the time 0≤t≤T, where Ε represents the expected calculation and T represents the maximum time step. When the attacker reaches the target node within T time steps If the attacker is not captured by the defender, or the defender still fails to capture the attacker when the game reaches the maximum time step T, the attacker wins and the reward is The defender gets punished On the contrary, if the defender successfully captures the attacker within T time steps, the defender wins and the reward is The attacker gets punished For a given strategy combination (π D ,π A ), then the value function of the attacker and the defender is defined as the second value function, and the mathematical expression of the second value function is:

[0067]

[0068] in, and The distribution is the rewards of the defender and the attacker at the time 0≤t≤T, Ε represents the expected calculation, T represents the maximum time step, represents the first value function of the defender, represents the attacker's first value function, V D (π D ,π A ) represents the second value function of the defender, V A (π D ,π A ) represents the attacker’s second value function.

[0069] In some embodiments, when executing step S3, the process of obtaining the form of the defender and attacker strategy populations includes:

[0070] S3-1. Use the defender strategy population Indicates that represents the defender's first strategy, represents the K+1th strategy of the defender, where K ≥ 0.

[0071] S3-2, use the attacker strategy population Indicates that represents the attacker's first strategy, represents the attacker’s K+1th strategy, where K≥0.

[0072] In some embodiments, when executing step S3, when obtaining the form of the defender meta-game, the meta-game is a matrix game, in which the defender's strategy set is represented as Contains K+1 strategies, among which, represents a meta-strategy of the defender, which is expressed as a probability distribution of the defender's strategy. The conditions that are satisfied include:

[0073]

[0074] in, Indicates the selection of the i-th strategy probability.

[0075] Furthermore, when obtaining the attacker's meta-game form, the attacker's strategy set is expressed as Contains K+1 strategies. represents a meta-strategy of the attacker, which represents a probability distribution about the attacker's strategy. The conditions that are satisfied include:

[0076]

[0077] in, Indicates the selection of the jth strategy probability.

[0078] In some embodiments, when executing step S4, the process of initializing the attacker's strategy includes: obtaining the Q value of the target node, and obtaining the attacker's strategy based on the Q value, wherein:

[0079] The mathematical expression of the Q value is:

[0080]

[0081] The mathematical expression of the attacker's strategy is:

[0082]

[0083] When the Q values ​​of all target nodes are initialized to 0, the attacker's strategy becomes a uniform strategy, and the mathematical expression is:

[0084]

[0085] Where |V′| represents the number of target nodes, represents the attacker’s uniform strategy. At this time, the attacker’s initial strategy population is The meta-strategy is σ A =(1).

[0086] In some embodiments, when executing step S4, the attack and defense benefits are filled into the meta-game benefit matrix for updating. When obtaining the updated meta-game benefit matrix, the meta-game benefit matrix is ​​represented as M. Assuming that the defender and the attacker use and The defender's payoff is The attacker's profit is Let M(i,j)=(u D ,u A ), the updated meta-game payoff matrix is ​​shown in Table 1.

[0087] Table 1 Updated meta-game payoff matrix

[0088]

[0089] In some embodiments, when executing step S4, the process of initializing the strategy population and meta-strategy of the attacker and defender includes: obtaining the Q value of the target node, and obtaining the attacker's strategy based on the Q value, wherein the mathematical expression of the Q value is:

[0090]

[0091] Specifically, for the target node The value represents the expected profit of the attacker of this node.

[0092] Furthermore, when the defender's strategy is based on σ D When sampling, if the attacker chooses As the target node, the expected benefit that the attacker can obtain. The attacker's strategy is defined according to the Q value, and the mathematical expression is:

[0093]

[0094] Furthermore, initialize the defender's strategy population and meta-strategy σ D . Randomly initialize a neural network policy Join the defender's strategy population The corresponding meta-strategy is σ D =(1).

[0095] Furthermore, we complete the meta-game payoff matrix. For each combination of defense strategy and attack strategy and Through simulation, the defender and the attacker adopt corresponding strategies to play the game, calculate the benefits of both, and fill them into the meta-game benefit matrix.

[0096] In some embodiments, when executing step S5, the process of obtaining the updated attacker's strategy population includes: defining the new defender meta-strategy σ according to the Q value D Calculate the Q value of each target node and get the attacker's new response strategy Then add the new strategy to the attacker's strategy population:

[0097]

[0098] in, represents the attacker’s strategy population, represents the attacker's new response strategy, Represents a union operation.

[0099] Furthermore, the process of training the defender's response strategy based on the strategy population to obtain a trained new strategy and adding it to the defender's strategy population includes:

[0100] Given the attacker's strategy population and meta-strategy σ A , the defender uses a deep reinforcement learning oracle to train the defender’s optimal response strategy Then the trained new strategy Adding it to the defender's strategy population, we get:

[0101]

[0102] in, represents the defender’s strategy population, represents the trained new strategy, Represents a union operation.

[0103] In fact, the resulting strategy population and meta-strategy together constitute an approximate Nash equilibrium for the APT intelligent lateral movement defense problem. Intuitively, the defender follows the meta-strategy σ D From the strategy population The sampling node isolation strategy is used to deal with APT attacks.

[0104] See also Figure 1 ,At the beginning, an APT intelligent lateral movement defense game model is established, the form of the defender and attacker strategy population is determined, the form of the defender and attacker meta-game is determined, and the form of the defender and attacker meta-game payoff matrix is ​​determined. Then, the defender and attacker strategy population and meta-strategy are initialized, and it is determined whether the current number of iterations is less than the maximum number of iterations K. If so, the meta-game payoff matrix is ​​completed through simulation, and the defender and attacker meta-strategies are sought through the meta-solver. Then, the response strategies of the defender and attacker are calculated / trained and added to the strategy population. If not, the loop ends directly.

[0105] See also Figure 2 , represents the result of the experiment on a synthetic network graph containing 100 nodes. The problem of APT intelligent lateral movement defense is studied on the synthetic network graph. The target node set is V′={v 28 ,v 43 ,v 66 ,v 82}, the attacker's initial position is The time range is T = 10, and the number of iterations of the population-based training algorithm is The experimental results are shown in the figure above. The worst-case benefit of the population-based training algorithm defender is significantly better than all baseline methods. This shows that the population-based training algorithm can effectively learn a robust node isolation review strategy and has a significant performance advantage when fighting against intelligent advanced attackers.

[0106] See also Figure 3 , represents the result of an experiment on a real communication network containing 100 nodes. The network is extracted from a real communication network with a sparse but complex topology. The target node set is T′={v 35 ,v46 ,v 85 ,v 90}, the attacker's initial position is The time range is T = 10, and the number of iterations of the population-based training algorithm is The experimental results are shown in the figure above. The population-based training algorithm significantly outperforms all baseline methods in terms of the defender's worst-case utility. This further proves that the population-based training algorithm can effectively learn robust defense strategies when dealing with advanced persistent threat attacks and maintain stable performance advantages under different network structures.

[0107] See also Figure 4 , represents the result graph of an experiment on a real researcher collaboration network containing 100 nodes. The network is extracted from the researcher collaboration network in the field of general relativity and quantum cosmology, which is highly complex and irregular. The target node set is V′={v 25 ,v 36 ,v 83 ,v 92}, the attacker's initial position is The time range is T = 10, and the number of iterations of the population-based training algorithm is The experimental results are shown in the figure above, indicating that the population-based training algorithm significantly outperforms all baseline methods in terms of the defender's worst-case utility. This further verifies the ability of the population-based training algorithm to effectively learn robust defense strategies in different network topologies, while demonstrating its wide applicability and excellent performance in dealing with complex network threats.

[0108] Although the embodiments of the present invention are described in detail above, it is obvious to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are within the scope and spirit of the present invention as described in the claims. Moreover, the present invention described herein may have other embodiments and may be implemented or realized in a variety of ways.

Claims

1. A method for isolating and defending APT nodes based on population training, characterized in that: include: Establish a game model for APT intelligent lateral movement, and determine the attacker's initial position and lateral movement method based on the target node; Based on the game model, establish the attacker's strategy and strategy space, the defender's strategy and strategy space, obtain a first value function based on the defender's strategy and the target node, and obtain a second value function based on the defender's strategy and the attacker's strategy; Obtain the form of the defender and attacker strategy populations, and obtain the form of the defender and attacker meta-games, and generate a meta-game payoff matrix; Initialize the strategy population and meta-strategy of the attacker and defender, play games with the defender and attacker based on the meta-strategy to obtain attack and defense benefits, and fill the attack and defense benefits into the meta-game benefit matrix for updating to obtain an updated meta-game benefit matrix; Solving the updated meta-game payoff matrix based on the meta-game solver to obtain the meta-strategies of the defender and the attacker, and calculating the new Q value of each target node to obtain the updated strategy population of the attacker; The response strategy of the defender is trained based on the strategy population to obtain a trained new strategy, which is added to the strategy population of the defender for iterative updating until the current number of iterations is greater than the preset maximum number of iterations, and the updating is stopped to obtain the Nash equilibrium.

2. According to claim 1, a population training-based APT node isolation defense method is characterized in that: The process of establishing the game model of APT intelligent lateral movement includes: constructing an undirected graph of an organization, the undirected graph consists of multiple nodes and multiple edges, the nodes communicate with each other through the organization's intranet, the nodes include important nodes and ordinary nodes in the organization, wherein the important nodes are the target nodes of the APT attacker, composed of data centers and servers, and are used to store the organization's core data, perform data analysis tasks or control key infrastructure, and the remaining nodes are ordinary nodes.

3. The APT node isolation defense method based on population training according to claim 1 is characterized in that: The mathematical expression of the first value function is: The mathematical expression of the second value function is: in, and The distribution is the rewards of the defender and the attacker at the time 0≤t≤T, Ε represents the expected calculation, T represents the maximum time step, represents the first value function of the defender, represents the attacker's first value function, V D (π D ,π A ) represents the second value function of the defender, V A (π D ,π A ) represents the attacker’s second value function.

4. The APT node isolation defense method based on population training according to claim 1 is characterized in that: The process of obtaining the form of the defender and attacker strategy populations includes: Use the defender strategy population Indicates that represents the defender's first strategy, represents the K+1th strategy of the defender, where K ≥ 0; The attacker strategy population is used Indicates that represents the attacker's first strategy, represents the attacker’s K+1th strategy, where K≥0.

5. The APT node isolation defense method based on population training according to claim 1 is characterized in that: When taking the form of the defender meta-game, the meta-game is a matrix game where the defender's strategy set is represented as Contains K+1 strategies, among which, represents a meta-strategy of the defender, which is expressed as a probability distribution of the defender's strategy. The conditions that are satisfied include: in, Indicates the selection of the i-th strategy probability.

6. The APT node isolation defense method based on population training according to claim 1 is characterized in that: When obtaining the form of the attacker's meta-game, the attacker's strategy set is expressed as Contains K+1 strategies, using represents a meta-strategy of the attacker, which represents a probability distribution about the attacker's strategy. The conditions that are satisfied include: in, Indicates the selection of the jth strategy probability.

7. The APT node isolation defense method based on population training according to claim 1 is characterized in that: The process of initializing the attacker's strategy includes: obtaining the Q value of the target node, and obtaining the attacker's strategy based on the Q value, wherein: The mathematical expression of the Q value is: The attacker's strategy The mathematical expression is: When the Q values ​​of all target nodes are initialized to 0, the attacker's strategy becomes a uniform strategy, and the mathematical expression is: Where |V′| represents the number of target nodes, represents the attacker’s uniform strategy, represents the first value function of the attacker.

8. The APT node isolation defense method based on population training according to claim 1 is characterized in that: The process of obtaining the updated attacker's strategy population includes: in, represents the attacker’s strategy population, represents the attacker's new response strategy, Represents a union operation.

9. The APT node isolation defense method based on population training according to claim 1 is characterized in that: The process of training the defender's response strategy based on the strategy population to obtain a trained new strategy and adding it to the defender's strategy population includes: in, represents the defender’s strategy population, represents the trained new strategy, Represents a union operation.

Citation Information

Patent Citations

  • Confrontation decision evaluation method for unmanned aerial vehicle

    CN105427032A

  • Self-adaptive platform migration defense method based on field programmable logic gate array

    CN116418548A

  • CPPS optimal defense strategy game method for uncertain attacks

    CN117439794A

  • Computing power network attack and defense confrontation game model based on supergame

    CN118821937A

  • Game-based optimal resource allocation method for cyber-physical power system (cpps) to defend against false data injection (fdi) attack

    GB202218851D0