An APT node isolation defense method based on population training
By establishing a game theory model and population training, constructing the policy spaces of defenders and attackers, generating a meta-game payoff matrix, and training the defender's response strategy, the problem of low defense efficiency caused by the intelligent lateral movement of APT attackers is solved, and more efficient APT detection and elimination is achieved.
Patent Information
- Application Number
- CN202510061472.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-01-15
AI Technical Summary
Existing technologies are ill-equipped to effectively counter the intelligent lateral movement strategies employed by advanced persistent threat (APT) attackers, resulting in inefficiencies for defenders in detecting and eliminating attackers.
A population-based APT node isolation defense method is adopted to establish a game model. By determining the attacker's initial position and lateral movement, the policy space of the defender and the attacker is constructed, a meta-game payoff matrix is generated, and the policies of the defender and the attacker are solved by the meta-game solver. The defender's response policy is trained to reach Nash equilibrium.
It improves the efficiency of defenders in detecting and eliminating APT attacks, effectively responds to worst-case attack strategies, and enhances the robustness of defense strategies.
Smart Images

Figure CN119995950B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of network security, and particularly relates to an APT node isolation defense method based on population training. BACKGROUND
[0002] Advanced persistent threat (APT) has long been considered as one of the notorious types of cyber attacks. Compared with conventional cyber attacks, APT has the characteristic of high penetration success rate, because the well-funded APT attackers can take advantage of the inherent weakness of human nature, conduct meticulous social engineering attacks, and thus successfully hijack the host nodes in the network of an institution. Once the APT attacker establishes a foothold in the intranet of an institution, he will try to approach the target nodes, i.e. the most valuable nodes in the institution such as data centers and servers, by lateral movement, with the purpose of stealing critical data and information from these nodes. Due to its high concealment, it is often difficult for the defender to detect the APT, making it one of the most serious threats to modern institutions.
[0003] Existing research has shown that it is necessary to consider lateral movement when optimizing the APT defense strategy of the defender, because through lateral movement, the APT attacker can effectively expand the attack range. However, most existing research assumes that the lateral movement of the attacker follows a certain fixed way, such as an unbroken node will be broken at a certain rate because of communication with a broken node. In actual scenarios, this assumption may still underestimate the ability of the APT attacker to some extent. Although some research considers the controllable rate of lateral movement, in practice the APT attacker may adopt a more intelligent strategy.
[0004] Specifically, the worst case for the defender is considered, i.e. the APT attacker chooses a shortest path to approach the target nodes of the institution from the initially broken node along this path. From the perspective of the APT attacker, the attacker only needs to control fewer nodes, thus can reduce the probability of being detected and eliminated, and also can reduce the attack cost. On the other hand, the longer lateral movement path is advantageous to the defender, because the defender can monitor the institution network through global information, thus improving the success rate of APT detection.
[0005] Therefore, there is an urgent need to provide a solution to improve the above problems. SUMMARY
[0006] The present application aims to provide an APT node isolation defense method based on population training, to improve the problem of low defense effectiveness of existing network security technology.
[0007] The APT node isolation defense method based on population training provided by the present application adopts the following technical solution:
[0008] A game model of APT intelligent horizontal movement is established, and the initial position of the attacker and the horizontal movement mode are determined based on the target node;
[0009] The strategy and strategy space of the attacker, the strategy and strategy space of the defender are established based on the game model, the first value function is obtained based on the strategy of the defender and the target node, and the second value function is obtained based on the strategy of the defender and the strategy of the attacker;
[0010] The form of the strategy population of the defender and the attacker is obtained, and the form of the meta-game of the defender and the attacker is obtained, and a meta-game payoff matrix is generated;
[0011] The strategy population and meta-strategy of the attacker and the defender are initialized, the attack and defense payoff is obtained by game between the defender and the attacker based on the meta-strategy, and the attack and defense payoff is filled into the meta-game payoff matrix for updating to obtain an updated meta-game payoff matrix;
[0012] The meta-strategy of the defender and the attacker is obtained by solving the updated meta-game payoff matrix based on a meta-game solver, and the new Q value of each target node is calculated to obtain an updated strategy population of the attacker;
[0013] The response strategy of the defender is trained based on the strategy population to obtain a trained new strategy, which is added to the strategy population of the defender for iterative updating, and the updating is stopped when the current iteration number is greater than the preset maximum iteration number to obtain a Nash equilibrium.
[0014] The APT node isolation defense method based on population training has the beneficial effects that the game theory is applied to APT horizontal movement behavior modeling, the worst case is considered from the perspective of the defender, that is, the APT attacker may adopt the most efficient attack strategy, a game theory mathematical model is established to describe the strategic interaction process between the APT attacker and the defender. When constructing the defense strategy, it is possible that it is inefficient to always select the most important target node (i.e. the node in V') for security review, because it may require a higher cost to carefully review these nodes. In the present application, the defender determines the node for security review according to the current state, which is more efficient. In addition, since the worst case is considered, the defense strategy of the defender is more robust.
[0015] Optionally, the process of establishing the game model of APT intelligent horizontal movement comprises: constructing an undirected graph of the organization, the undirected graph is composed of multiple nodes and multiple edges, the nodes communicate through the intranet of the organization, the nodes include important nodes and ordinary nodes in the organization, wherein the important nodes are target nodes of the APT attacker, composed of data centers and servers, used to store core data of the organization, perform data analysis tasks or control critical infrastructure, and the remaining nodes are ordinary nodes.
[0016] Optionally, the mathematical expression of the first value function is:
[0017]
[0018] The mathematical expression of the second value function is:
[0019]
[0020] wherein, and is the distribution of the rewards of the defender and the attacker at time 0≤t≤T, E represents the expected calculation, and T represents the maximum time step, represents the first value function of the defender, represents the first value function of the attacker, V D (π D ,π A ) represents the second value function of the defender, V A (π D ,π A ) represents the second value function of the attacker.
[0021] Optionally, the process of obtaining the form of the strategy population of the defender and the attacker includes:
[0022] The strategy population of the defender is represented by , wherein represents the first strategy of the defender, represents the K+1th strategy of the defender, wherein K≥0.
[0023] The strategy population of the attacker is represented by , wherein represents the first strategy of the attacker, represents the K+1th strategy of the attacker, wherein K≥0.
[0024] Optionally, when the form of the meta-game of the defender is obtained, the meta-game is a matrix game, wherein the strategy set of the defender is represented by contains K+1 strategies, wherein represents a meta-strategy of the defender, and the meta-strategy is represented as a probability distribution of the strategies of the defender, and the conditions satisfied include:
[0025]
[0026] wherein, represents the probability of selecting the ith strategy .
[0027] Optionally, when the form of the attacker's metagame is obtained, the strategy set of the attacker is represented as where K+1 represents the number of strategies of the attacker, and represents a metagame of the attacker, which represents a probability distribution about the strategy of the attacker, and the conditions met include:
[0028]
[0029] wherein, represents the probability of selecting the jth strategy.
[0030] Optionally, the process of initializing the strategy of the attacker includes: obtaining the Q value of the target node, and obtaining the strategy of the attacker based on the Q value, wherein,
[0031] the mathematical expression of the Q value is:
[0032]
[0033] the mathematical expression of the strategy of the attacker is:
[0034]
[0035] when the Q values of all target nodes are initialized as 0, the strategy of the attacker becomes a uniform strategy, and the mathematical expression is:
[0036]
[0037] wherein, |V'| represents the number of target nodes, represents the uniform strategy of the attacker, represents the first value function of the attacker.
[0038] Optionally, the process of obtaining the updated strategy population of the attacker includes:
[0039]
[0040] wherein, represents the strategy population of the attacker, represents the new response strategy of the attacker, represents the set operation.
[0041] Optionally, the process of adding the trained new strategy to the strategy population of the defender based on the training of the response strategy of the defender includes:
[0042]
[0043] wherein, representing the strategy population of the defender, representing the trained new strategy, representing the union operation. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 is a flowchart of an APT node isolation defense method based on population training provided by the present application;
[0045] Figure 2 is a result graph of an experiment performed on a synthetic network containing 100 nodes provided by the present application;
[0046] Figure 3 is a result graph of an experiment performed on a real communication network containing 100 nodes provided by the present application;
[0047] Figure 4 is a result graph of an experiment performed on a real researcher cooperation network containing 100 nodes provided by the present application. DETAILED DESCRIPTION
[0048] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application. Unless otherwise defined, the technical terms or scientific terms used herein should be understood as the general meaning understood by those of ordinary skill in the art to which the present application belongs. The words such as “comprise” and similar words used herein mean that the elements or objects before the words cover the elements or objects listed after the words and their equivalents, and do not exclude other elements or objects.
[0049] The embodiments of the present application provide an APT node isolation defense method based on population training, comprising:
[0050] S1, establishing a game model of APT intelligent horizontal movement, determining the initial position and horizontal movement mode of the attacker based on the target node;
[0051] S2, establishing the strategy and strategy space of the attacker and the strategy and strategy space of the defender based on the game model, obtaining a first value function based on the strategy of the defender and the target node, and obtaining a second value function based on the strategy of the defender and the strategy of the attacker;
[0052] S3, obtaining the form of the strategy population of the defender and the attacker, and obtaining the form of the meta-game of the defender and the attacker, and generating a meta-game payoff matrix;
[0053] S4, initialize the strategy population and meta-strategy of the attacker and the defender, obtain attack and defense benefits based on the same strategy game between the defender and the attacker, fill the attack and defense benefits into the meta-game benefit matrix to update the meta-game benefit matrix and obtain an updated meta-game benefit matrix;
[0054] S5, solve the updated meta-game benefit matrix based on a meta-game solver to obtain the meta-strategy of the defender and the attacker, and calculate the new Q value of each target node to obtain an updated strategy population of the attacker;
[0055] S6, based on the strategy population, train the response strategy of the defender to obtain a trained new strategy, and add it to the strategy population of the defender for iterative updating until the current iteration number is greater than the preset maximum iteration number to stop updating and obtain Nash equilibrium.
[0056] In some embodiments, in the process of performing step S1 to establish the game model of APT intelligent horizontal movement, the process includes: constructing an undirected graph of the organization, the undirected graph being composed of a plurality of nodes and a plurality of edges, the nodes being communicated through the intranet of the organization, the nodes including important nodes and ordinary nodes in the organization, wherein the important nodes are target nodes of the APT attacker, composed of data centers and servers, used to store core data of the organization, perform data analysis tasks or control critical infrastructure, and the remaining nodes are ordinary nodes.
[0057] Specifically, the intranet of the organization is represented by a graph G=(V,E), wherein V={v1,…,v N} is a node set, and E is an edge set. {v i ,v j}∈E represents that node i and node j communicate through the intranet of the organization. V'∈V represents important nodes in the organization, and other nodes in the organization are ordinary nodes, represented by a set .
[0058] Further, in performing step S1, the process of determining the initial position and horizontal movement mode of the attacker based on the target node includes:
[0059] Let represent the starting attack node of the attacker, at t=0, the attacker selects a target node Then move horizontally through the shortest path from to , the path is represented as
[0060] In some embodiments, in the process of performing step S2 to establish the strategy and strategy space of the attacker, the process includes:
[0061] The attacker selects a target node at time t = 0 according to the strategy π A : Select a target node, and then move horizontally according to the shortest path from the initial node to the selected target node. Let Π A be the strategy space of the attacker, then π A ∈ Π A .
[0062] Further, the process of establishing the strategy and strategy space of the defender includes:
[0063] The defender selects a node for isolation review according to the current state at each time step to determine whether the node is subject to APT attack. The system state observed by the defender includes two parts: the current time step and the historical action sequence, i.e. Let S D represent the space of the state observed by the defender, then the strategy of the defender is defined as π D : S D → Δ(V). Let Π D be the strategy space of the defender, then π D ∈ Π D . In the present application, the strategy of the defender is represented by a neural network.
[0064] In some embodiments, when performing step S2, given the defender strategy π D and the target node , the value functions of the attacker and the defender are defined as a first value function, and the mathematical expression of the first value function is:
[0065]
[0066] wherein, and are the rewards of the defender and the attacker at 0 ≤ t ≤ T, E represents an expected calculation, and T represents the maximum time step. When the attacker reaches the target node at the T time step and is not captured by the defender, or the game is played to the maximum time step T and the defender still fails to capture the attacker, the attacker wins, and the reward is while the defender is punished On the contrary, if the defender successfully captures the attacker within the T time step, the defender wins, and the reward is while the attacker is punished For a given strategy combination (π D , π A ), the value functions of the attacker and the defender are defined as a second value function, and the mathematical expression of the second value function is:
[0067]
[0068] wherein, and are the rewards of the defender and the attacker at time step t, E denotes the expectation calculation, and T denotes the maximum time step, denotes the first value function of the defender, denotes the first value function of the attacker, V D (π D ,π A ) denotes the second value function of the defender, V A (π D ,π A ) denotes the second value function of the attacker.
[0069] In some embodiments, when performing step S3, the process of obtaining the form of the defender and the attacker strategy population, comprises:
[0070] S3-1, the defender strategy population is represented by , wherein denotes the first strategy of the defender, denotes the K+1th strategy of the defender, wherein K≥0.
[0071] S3-2, the attacker strategy population is represented by , wherein denotes the first strategy of the attacker, denotes the K+1th strategy of the attacker, wherein K≥0.
[0072] In some embodiments, when performing step S3, the form of the defender metagame is obtained, the metagame is a matrix game, wherein the strategy set of the defender is represented by contains K+1 strategies, wherein denotes a meta-strategy of the defender, which is represented as a probability distribution of the defender strategies, and the conditions satisfied include:
[0073]
[0074] wherein, denotes the probability of selecting the i-th strategy .
[0075] Further, when the form of the attacker metagame is obtained, the strategy set of the attacker is represented by contains K+1 strategies. Denote a meta-strategy of the attacker by , which represents a probability distribution of the attacker strategies, and the conditions satisfied include:
[0076]
[0077] wherein, denotes the probability of selecting the jth strategy.
[0078] In some embodiments, when performing step S4, the process of initializing the attacker strategy includes: obtaining the Q value of the target node, and obtaining the strategy of the attacker based on the Q value, wherein,
[0079] The mathematical expression of the Q value is:
[0080]
[0081] The mathematical expression of the strategy of the attacker is:
[0082]
[0083] When the Q values of all target nodes are initialized to 0, the strategy of the attacker becomes a uniform strategy, and the mathematical expression is:
[0084]
[0085] wherein, |V'| denotes the number of target nodes, denotes the uniform strategy of the attacker, and at this time, the initial strategy population of the attacker is The meta-strategy is σ A =(1).
[0086] In some embodiments, when performing step S4, the attack-defense payoff is filled into the meta-game payoff matrix to update and obtain an updated meta-game payoff matrix, and the meta-game payoff matrix is denoted as M. Assuming that the defender and the attacker use and strategies to play the game, the payoff of the defender is The payoff of the attacker is Let M(i,j)=(u D ,u A ), and the updated meta-game payoff matrix is shown in Table 1.
[0087] Table 1: Updated meta-game payoff matrix
[0088]
[0089] In some embodiments, when performing step S4, the process of initializing the strategy population of the attacker and the defender and the meta-strategy includes: obtaining the Q value of the target node, and obtaining the strategy of the attacker based on the Q value, wherein the mathematical expression of the Q value is:
[0090]
[0091] Specifically, for a target node The value represents the expected payoff of the node attacker.
[0092] Further, when the defender's strategy is sampled according to σ D , if the attacker chooses as the target node, the expected payoff that the attacker can obtain. According to the Q value definition, the strategy of the attacker is defined as:
[0093]
[0094] Further, the defender's strategy population and the meta-strategy σ D are initialized. A neural network strategy is randomly initialized and added to the defender's strategy population The corresponding meta-strategy is σ D =(1).
[0095] Further, the meta-game payoff matrix is completed. For each combination of defender strategy and attacker strategy and Through simulation, that is, the defender and the attacker adopt the corresponding strategies to play the game, the payoffs of the two are calculated, and are filled into the meta-game payoff matrix.
[0096] In some embodiments, when performing step S5, the process of obtaining the updated strategy population of the attacker includes: calculating the Q value of each target node according to the Q value definition and the new defender meta-strategy σ D , obtaining the new response strategy of the attacker Then the new strategy is added to the strategy population of the attacker:
[0097]
[0098] Wherein, represents the strategy population of the attacker, represents the new response strategy of the attacker, represents the union operation.
[0099] Further, the process of training the response strategy of the defender based on the strategy population to obtain a trained new strategy and adding the trained new strategy to the strategy population of the defender includes:
[0100] Given the strategy population of the attacker and the meta-strategy σ A , the defender uses a deep reinforcement learning oracle to train the optimal response strategy of the defender The trained new strategy is then added to the defender's strategy population, resulting in:
[0101]
[0102] wherein, denotes the defender's strategy population, denotes the trained new strategy, denotes the union operation.
[0103] In fact, the resulting strategy population together with the metastategy constitutes an approximate Nash equilibrium for the APT intelligent lateral movement defense problem. Intuitively, the defender samples a node quarantine strategy from the strategy population D according to the metastategy to counter the attack of the APT.
[0104] Referring to Figure 1 , at the beginning, an APT intelligent lateral movement defense game model is established, the forms of the defender's and attacker's strategy populations are determined, the form of the defender's and attacker's metastategy is determined, and the form of the payoff matrix of the defender's and attacker's metastategy is determined, then the defender's and attacker's strategy populations and metastategy are initialized, it is judged whether the current iteration number is less than the maximum iteration number K, if yes, the payoff matrix of the metastategy is completed through simulation, the defender's and attacker's metastategy is solved through a metastategy solver, then the response strategies of the defender and the attacker are calculated / trained and the response strategies are added to the strategy population, if no, the loop is directly ended.
[0105] Referring to Figure 2 , the results of experiments conducted on a synthetic network graph containing 100 nodes are shown. The APT intelligent lateral movement defense problem is studied on a synthetic network graph, the target node set is V' = {v 28 ,v 43 ,v 66 ,v 82}, the initial position of the attacker is the time range is T = 10, and the iteration number of the population-based training algorithm is The experimental results are shown in the above graph, the worst-case payoff of the defender based on the population-based training algorithm is significantly better than all baseline methods. This shows that the population-based training algorithm can effectively learn a robust node quarantine strategy and has a significant performance advantage in countering intelligent advanced attackers.
[0106] Referring to Figure 3 , the result graph of experiments conducted on a real communication network containing 100 nodes is shown. The network is extracted from a real communication network and has a sparse but complex topology structure, the target node set is T' = {v 35 ,v46 ,v 85 ,v 90},attackers' initial positions are The time horizon is T = 10, and the number of iterations of the population-based training algorithm is The experimental results are shown in the above figure, and the population-based training algorithm is significantly better than all baseline methods in terms of the defender's worst-case utility. It is further proved that the population-based training algorithm can effectively learn robust defense strategies when facing advanced persistent threat attacks, and maintains stable performance advantages under different network structures.
[0107] Referring to Figure 4 , the figure shows the results of experiments conducted on a real researcher collaboration network containing 100 nodes. This network is extracted from the researcher collaboration network in the field of general relativity and quantum cosmology, and has high complexity and irregularity. The target node set is V' = {v 25 ,v 36 ,v 83 ,v 92},attackers' initial positions are The time horizon is T = 10, and the number of iterations of the population-based training algorithm is The experimental results are shown in the above figure, and the population-based training algorithm is significantly better than all baseline methods in terms of the defender's worst-case utility. It is further verified that the population-based training algorithm can effectively learn robust defense strategies in different network topologies, while demonstrating its wide applicability and excellent performance in dealing with complex network threats.
[0108] Although the embodiments of the present application are described in detail above, it is obvious to those skilled in the art that various modifications and changes can be made to these embodiments. However, it should be understood that such modifications and changes are within the scope and spirit of the present application as described in the claims. Moreover, the present application described herein can have other embodiments and can be implemented or realized in various ways.
Claims
1. A population-based APT node isolation and defense method, characterized in that, include: Establish a game theory model for APT intelligent lateral movement, determining the attacker's initial position and lateral movement method based on the target node, including: using Indicates the attacker's starting attack node, in At any given moment, the attacker selects a target node. Then through from arrive Move laterally along the shortest path; Based on the game model, establish the attacker's strategy and strategy space, the defender's strategy and strategy space, obtain a first value function based on the defender's strategy and the target node, and obtain a second value function based on the defender's strategy and the attacker's strategy. The mathematical expression for the first value function is: ; ; The mathematical expression for the second-valued function is: ; ; in, and For the defender and the attacker respectively Momentary rewards This indicates the expected calculation. Indicates the maximum time step. Describes the first-valued function of the defender. This represents the attacker's first-value function. This represents the second-valued function of the defender. The second-valued function representing the attacker. Indicates the target node; Obtain the forms of the defender's and attacker's policy populations, and obtain the forms of the defender's and attacker's meta-game. Generate the meta-game payoff matrix. The defender and attacker use their corresponding strategies to play the game, calculate their payoffs, and fill them into the meta-game payoff matrix, including: Use the defender strategy population It means that, among them This indicates the defender's first strategy. The first one represents the defender One strategy, among which ; Use attacker strategy population It means that, among them This represents the attacker's first strategy. The attacker's first One strategy, among which ; Metagame is a matrix game, where the set of strategies for the defenders is represented as follows: ,Include One strategy, among which, using Let represent a meta-policy of the defender, which is expressed as a probability distribution of the defender's policies, satisfying the following conditions: ; in, Indicates selecting the first One strategy The probability of; The attacker's set of strategies is represented as ,Include One strategy, using Let represent an attacker's meta-policy, which is a probability distribution of the attacker's policy that satisfies the following conditions: ; in, Indicates selecting the first One strategy The probability of; The process involves initializing the attacker's and defender's policy populations and meta-policies, engaging in game theory between the attacker and defender based on the meta-policies to obtain attack and defense payoffs, and then updating the meta-game payoff matrix by filling the payoffs into the meta-game payoff matrix. The initialization of the attacker's policy includes: obtaining the Q-value of the target node and obtaining the attacker's policy based on the Q-value. The mathematical expression for the Q value is: ; The attacker's strategy The mathematical expression is: ; When the Q-values of all target nodes are initialized to 0, the attacker's strategy becomes a uniform strategy, mathematically expressed as: ; in, Indicates the number of target nodes. This represents the attacker's uniform strategy. Represents the attacker's first-value function; The updated meta-game payoff matrix is solved using a meta-game solver to obtain the meta-policies of the defender and attacker, and the new Q-value of each target node is calculated to obtain the updated attacker's policy population. The defender's response strategy is trained based on the strategy population to obtain a new trained strategy, which is then added to the defender's strategy population for iterative updates until the current iteration count exceeds the preset maximum iteration count, at which point the updates stop and a Nash equilibrium is obtained.
2. The APT node isolation and defense method based on population training according to claim 1, characterized in that, The process of establishing the game model for APT intelligent lateral movement includes: constructing an undirected graph of the organization, which consists of multiple nodes and multiple edges. The nodes communicate with each other through the organization's intranet. The nodes include important nodes and ordinary nodes in the organization. The important nodes are the target nodes of the APT attackers, which consist of data centers and servers, and are used to store the organization's core data, perform data analysis tasks, or control critical infrastructure. The remaining nodes are ordinary nodes.
3. The APT node isolation and defense method based on population training according to claim 1, characterized in that, The process of obtaining the updated attacker strategy population includes: ; in, This represents the attacker's strategy population. This indicates a new response strategy from the attacker. This indicates the union operation.
4. The APT node isolation and defense method based on population training according to claim 1, characterized in that, The process of training a new, well-trained strategy for the defender based on the aforementioned strategy population and adding it to the defender's strategy population includes: ; in, This represents the strategy population of the defenders. This indicates a well-trained new strategy. This indicates the union operation.
Citation Information
Patent Citations
Confrontation decision evaluation method for unmanned aerial vehicle
CN105427032A
Self-adaptive platform migration defense method based on field programmable logic gate array
CN116418548A