A Dynamic Defense Method and System for APT Network Kill Chain Based on Dual Deep Reinforcement Learning

The dynamic defense method for APT network kill chains, which utilizes dual deep reinforcement learning, combines network topology and node information to optimize resource allocation and defense strategies. This addresses the problem of inappropriate resource allocation in advanced persistent threat attacks, thereby improving network defense effectiveness and system security.

CN119788348BActive Publication Date: 2025-12-02XIAMEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411868838.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-12-02
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Existing technologies struggle to allocate limited defense resources rationally and dynamically across the multi-stage chain of advanced persistent threat attacks, leading to resource waste and impacting system performance. Network deception defense technologies also consume significant computational resources.

Method used

A dynamic defense method for APT network kill chains based on dual deep reinforcement learning is adopted. By combining network topology changes, node importance and vulnerability information, the method dynamically optimizes resource allocation and defense strategies by acquiring network status and attacker behavior in real time and using specialized defense devices.

Benefits of technology

It achieves optimal deception asset deployment in complex attack scenarios, enhances network defense capabilities, reduces attackers' profitability, and improves system security and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119788348B_ABST
    Figure CN119788348B_ABST
Patent Text Reader

Abstract

A dynamic defense method and system for APT network kill chains based on dual deep reinforcement learning is proposed. The system pre-configures a network system including host nodes, security defenders, malicious attackers, and security defense devices. Malicious attackers execute attacks on host nodes according to the network kill chain. Security defenders monitor these attacks and acquire network status. Security defense devices, based on dual deep reinforcement learning algorithms and combined with the network status, dynamically generate optimal defense strategies and feed them back to the security defenders. The security defenders then dynamically respond to the attacks using these optimal strategies to protect against the network kill chain. This approach achieves efficient identification and accurate response to APT attacks, effectively improving system security and resource utilization efficiency. It provides an innovative solution for dynamic defense in complex network environments and significantly enhances the resistance to APT attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network attack and defense security, and in particular to a dynamic defense method and system for APT network kill chains based on dual deep reinforcement learning. Background Technology

[0002] In recent years, with the increasing complexity and stealth of cyberattacks, especially the frequent occurrence of Advanced Persistent Threat (APTS) attacks, effectively detecting and defending against such attacks has become a key challenge in the field of cybersecurity. APTS attacks achieve their objectives through a multi-stage network kill chain, including reconnaissance, weaponization, delivery, exploitation, installation, command and control, and ultimately, achieving the attack goal. This presents significant challenges for defenders in identifying and intervening in APTS attacks. Chinese Patent CN117395072A discloses a method for generating network kill chains. This method analyzes network traffic and terminal behavior to generate multi-source logs and security alert logs, aggregates the alert logs, and then correlates the multi-source logs to generate user behavior sequences. The method uses a pre-trained detection model to identify attack traffic within the behavior sequences. If attack traffic is detected, the corresponding network kill chain is generated using the generative model. This method significantly improves the efficiency and accuracy of attack traffic identification and accurately reconstructs the network kill chain of attack traffic, helping cybersecurity personnel comprehensively assess and respond to potential threats in a timely manner.

[0003] While existing technologies can improve the accuracy and comprehensiveness of generating kill chains, how to rationally and dynamically allocate limited defense resources across the multi-stage chain of advanced persistent threat (APT) attacks to ensure optimal defense effectiveness at each stage remains a key challenge. Network deception technology has been introduced into APT defense systems as an innovative defense method. Chinese patent CN116614296A discloses a honeypot deception defense method that helps security personnel efficiently locate attack methods and optimize defense responses by identifying attacker behavior and displaying it based on an ATT&CK matrix view. AHAnwar et al. [AHAnwar, CAKamhoua, NOLeslie and C. Kiekintveld, "Honeypot Allocation for Cyber ​​Deception Under Uncertainty," in IEEE Transactions on Network and Service Management, vol.19, no.3, pp.3438-3452, Sept.2022, doi:10.1109 / TNSM.2022.3179965.] proposed a deception method combining game theory and reinforcement learning models. They modeled the reactive deception problem as a partially observable Markov decision process based on a game-theoretic dynamic model to address the incomplete monitoring of attacker behavior.

[0004] However, implementing network deception defense technologies requires significant computing resources and network bandwidth. Improper resource allocation can lead to high costs and even affect normal system performance. Therefore, rationally allocating network deception defense resources and optimizing defense strategies have become important issues in the defense against advanced persistent threat (APT) attacks. Zhang, L et al. [L.Zhang,T.Zhu,FKHussain,D.Ye and W.Zhou, "A Game-Theoretic Method for Defending Against Advanced Persistent Threats in Cyber ​​Systems,"in IEEE Transactions on Information Forensics and Security, vol.18,pp.1349-1364,2023,doi:10.1109 / TIFS.2022.3229595.] helped defenders achieve optimal defense results when dealing with APT attacks by optimizing defense strategies, adjusting timing, and allocating resources, thereby maximizing defense efficiency and reducing resource waste. B. Peng et al. [B. Peng, J. Liu and J. Zeng, "Dynamic Analysis of Multiplex Networks With Hybrid Maintenance Strategies," in IEEE Transactions on Information Forensics and Security, vol. 19, pp. 555-570, 2024, doi:10.1109 / TIFS.2023.3324386.] proposed a propagation suppression strategy based on the RCMO model. Combining static and dynamic control methods, this strategy effectively suppresses malware propagation and optimizes resource allocation to improve network security and reduce security risks.

[0005] Dual deep reinforcement learning, as an intelligent optimization method, can adaptively learn in complex and unknown environments, gradually optimizing strategies through continuous feedback. It is highly suitable for complex and dynamic defense tasks such as advanced persistent threat attacks. Y.Yu et al. [Y.Yu, W.Yang, W.Ding and J.Zhou, "Reinforcement Learning Solution for Cyber-Physical Systems Security Against Replay Attacks," in IEEE Transactions on Information Forensics and Security, vol.18, pp.2583-2595, 2023, doi:10.1109 / TIFS.2023.3268532.] proposed an attack detection method based on reinforcement learning. Through a model-free learning framework, it can automatically identify and respond to the attacker's dynamic strategies, while simultaneously designing an optimization learning-based defense strategy, effectively improving network security protection capabilities. W. He et al. [W. He, J. Tan, Y. Guo, K. Shang and H. Zhang, "A Deep Reinforcement Learning-Based Deception Asset Selection Algorithm in Differential Games," in IEEE Transactions on Information Forensics and Security, vol. 19, pp. 8353-8368, 2024, doi:10.1109 / TIFS.2024.3451189.] proposed a differential game-based deception asset selection algorithm based on multi-agent deep reinforcement learning. By constructing a network security state evolution analysis and a differential game model, they optimized the deployment of deception assets in the defense against advanced persistent threat attacks, achieving more efficient attack defense strategy selection.

[0006] In APT attacks, malicious APT attackers, using tools such as Burp Suite and OpenVAS, can only obtain information about compromised nodes and their neighboring nodes. Based on the importance of the nodes and the characteristics of the vulnerabilities, they choose the optimal path, balancing minimum attack cost with maximum attack benefit. Subsequently, the attackers launch various threat attacks against the target node, including scanning attacks, phishing attacks, botnet attacks, DDoS attacks, zero-day attacks, key corruption attacks, fake key attacks, and data breach attacks. If the attack is successful, the attacker will advance to the next stage of the network kill chain, continuously interacting with defenders employing a binding strategy. Summary of the Invention

[0007] The main objective of this invention is to overcome the network deception defense problem in the existing technology and propose a dynamic defense method and system for APT network kill chains based on dual deep reinforcement learning. It comprehensively considers network topology changes, node importance and vulnerability information, and aims to dynamically optimize resource allocation and defense strategy selection by acquiring network status and dangerous attacker behavior information in real time.

[0008] The present invention adopts the following technical solution:

[0009] A dynamic defense method for APT network kill chain based on dual deep reinforcement learning is proposed, which pre-configures a network system, including host nodes, security defenders, malicious attackers, and security defense devices.

[0010] The attacker executes attacks on the host node according to the network kill chain. The security defender monitors the attack and obtains the network status. The security defense device, based on a dual deep reinforcement learning algorithm, dynamically generates the optimal defense strategy in combination with the network status and feeds it back to the security defender. The security defender uses the optimal defense strategy to dynamically respond to the attack to achieve protection against the network kill chain.

[0011] The attack on the host node by the dangerous attacker according to the network kill chain specifically refers to:

[0012] The attacker gradually utilizes the set of zombie nodes according to each stage of the network kill chain, selects targets for attack based on the characteristics of host nodes adjacent to the zombie nodes, allocates dangerous attack resources according to the attack budget, and launches dangerous attacks. The set of host nodes that are successfully damaged will be transformed into a new set of zombie nodes.

[0013] The security defender models the dynamic risk level of the host node being attacked by zombie nodes as an exponential function; it introduces a response indication vector to describe the resources that the security defender allocates to the damaged host node using a limited defense budget.

[0014] In time slot t, the security defender observes the network state of the previous time slot t-1 and constructs the observation state function S[t] for time slot t, where S[t] = {I[t-1], V[t-1], c A [t-1],r A [t-1],d D [t-1],c D [t-1],r D[t-1], TD[t-1], p[t-1]}, I[t-1] is the set of importance of the host nodes of the network system at time slot t-1, V[t-1] is the set of vulnerability of the host nodes of the network system at time slot t-1, c A [t-1] represents the cost set of the dangerous attacker employing the dangerous attack strategy at time slot t-1, r A [t-1] represents the set of rewards obtained by the dangerous attacker at time slot t-1, d D [t-1] represents the set of resources allocated by the security defender to the compromised host node at time slot t-1, c D [t-1] represents the cost set of the security defense strategy adopted by the security defender at time slot t-1, r D [t-1] represents the reward obtained by the security defender at time slot t-1, TD[t-1] represents the set of host nodes that were damaged at time slot t-1, and p[t-1] represents the set of risk levels of the host nodes that were damaged at time slot t-1 under dangerous attacks.

[0015] The security defense device, based on the observed state function S[t] of the network state in time slot t, uses a dual deep reinforcement learning algorithm to obtain the spatial set A[t] = {d} of the network deception defense deployment strategy adopted in time slot t. D [t],π[t]},d D [t] represents the resource allocation of the security defender at time slot t, π[t] represents the network deception defense strategy selected by the security defender at time slot t, and the utility of the security bundling defense strategy is calculated as follows: Among them DR i [t] represents the response indication vector.

[0016] Where TD[t] is the set of host nodes damaged at time slot t, and DR i [t] indicates whether the security defender takes action to protect the damaged host node i at time slot t. i [t] represents the risk level of the host node i being compromised at time slot t, indicating a dangerous attack. This refers to the amount of resources allocated by the security defender to the compromised host node i during time slot t. It is the cost of the security defender employing a security defense strategy against the compromised host node i during time slot t.

[0017] The dual deep reinforcement learning algorithm includes a main network and a target network. The main network selects the optimal action to perform based on the current state, while the target network evaluates the value of the action selected by the main network and calculates the target Q-value matrix accordingly. The main network is updated and trained at each time step, and the target network synchronously updates its weights at fixed time intervals to track the changes of the main network.

[0018] The Q-value matrix is ​​Q a (S[t], A[t]), which represents the Q value of the security defender when selecting the network deception defense deployment strategy space set A[t] under the observed state function S[t] of the network state. The expression for updating the Q value matrix is ​​as follows:

[0019] Q a (S[t],A[t])=(1-α)Q a (S[t],A[t])+α(R D [t+1]+γ·Q b (S[t+1],max a∈A Q a (S[t+1],A[t])));

[0020] Where α is the learning factor, γ is the discount factor, and R0 D [t+1] and S[t+1] represent the security defender's gains and state in time slot t+1, respectively. A[t] represents the security defender's action in time slot t. The max(.) function is the function to take the maximum value.

[0021] The security defense device also receives defense results from the security defender to adjust the optimal defense strategy, forming a closed-loop feedback mechanism.

[0022] It also includes determining whether the network system has failed based on failure condition thresholds. Network system failure refers to the network system's operational logic ceasing, making further attack and defense interactions impossible. The failure threshold conditions are as follows:

[0023]

[0024] Where Th1 and Th2 are the information importance leakage threshold and the available node threshold, respectively, I i |VN0| is the importance value of host node i in the network system, and |VN0| is the total number of host nodes in the initial network system before being attacked by a malicious attacker. t | is the number of all available host nodes in the network system at time slot t, N i N represents the security, damage, or eviction status of a host node. i =1 indicates that the host node has been damaged or evicted; otherwise, N i =0.

[0025] If the network system fails or reaches the maximum number of iterations, the process ends; otherwise, it is determined whether the attacker succeeded. If the attack failed, the current stage of the network kill chain continues; if the attack succeeded, the process proceeds to the next stage of the network kill chain.

[0026] A dynamic defense system for APT network kill chains based on dual deep reinforcement learning, comprising:

[0027] host node,

[0028] A dangerous attacker executes attack actions against the host node according to the network kill chain;

[0029] The security defender monitors the attack behavior and obtains the network status, and uses the optimal defense strategy to dynamically respond to the attack behavior in order to protect the network kill chain.

[0030] The security defense device, based on a dual deep reinforcement learning algorithm, dynamically generates the optimal defense strategy by combining the network state and feeds it back to the security defender.

[0031] As can be seen from the above description of the present invention, compared with the prior art, the present invention has the following beneficial effects:

[0032] This invention analyzes and obtains network status by importing real-time network traffic from a security defender, and uses a dual deep reinforcement learning algorithm to iteratively update the bundled defense strategy. This enables the selection of the optimal deception asset deployment strategy in multi-stage attack scenarios, effectively responding to complex attacks. While protecting system host nodes and improving system security, it also reduces the potential profit of dangerous attackers.

[0033] This invention systematically discusses how dangerous attackers shift their operations at different stages of the network kill chain and interact with security defenders employing a binding strategy. Based on this, a security defense device is designed. Through intelligent analysis and algorithm modules, it helps defenders accurately predict attacker behavior, assess risk intensity, and select appropriate optimized resource allocation and defense strategies. This significantly improves network defense capabilities and enhances the resistance to APT attacks, demonstrating significant advantages and promising application prospects.

[0034] This invention innovatively discusses the vulnerability of nodes, defining it as a dynamic and comprehensive vulnerability related to the allocation of offensive and defensive resources in time and interaction. It also considers that in the case of network topology changes, a dangerous attacker can only know the information of adjacent nodes and uses the importance and vulnerability of nodes to select the optimal path.

[0035] This invention achieves efficient identification and accurate response to APT attacks, effectively improving system security and resource utilization efficiency. It provides an innovative solution for dynamic defense in complex network environments, significantly enhancing the resistance to APT attacks, and possesses significant technological innovation and practical value. Attached Figure Description

[0036] Figure 1 This is a diagram of the network system architecture of the present invention.

[0037] Figure 2 This is a block diagram of the security defense device of the present invention.

[0038] Figure 3 This is a flowchart of the method of the present invention.

[0039] Figure 4 This is the state transition diagram of the network kill chain in this invention.

[0040] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Detailed Implementation

[0041] The present invention will be further described below through specific embodiments.

[0042] See Figure 3 This paper presents a dynamic defense method for APT network kill chains based on dual deep reinforcement learning. The method pre-configures a network system, which includes host nodes, security defenders, malicious attackers, and security defense devices. The network system model is built on an SDN architecture, containing multiple host nodes with varying importance and known vulnerabilities, and a core component where the SDN controller acts as the security defender. The security defender deploys a network intrusion detection system to monitor network traffic in real time and responds to changes in complex network topology through dynamic security defense devices. To counter the diverse attacks launched by malicious APT attackers at different stages of the kill chain, the security defense device can dynamically adjust its output defense strategies. For example, it can output honeypot defense strategies, allowing the security defender to utilize honeypots to capture malicious attackers, analyze attack behavior, and generate decoy data.

[0043] Initialize the network system and related parameters, including the Q-value matrix, learning factor α, and discount factor γ. Configure the number of host nodes (e.g., 100), the number of security defenders (e.g., 1), and the number of malicious attackers (e.g., 2). Determine the network system failure condition threshold (e.g., set to...). and Determine the total amount of resources available to the security defender (e.g., 50), including time and computing resources, and initialize the importance and vulnerability of each host node.

[0044] The method of the present invention further includes the following steps:

[0045] S1 attackers execute attacks on host nodes according to the network kill chain.

[0046] Specifically, attackers gradually utilize the set of zombie nodes (i.e., the set of compromised nodes) according to each stage of the network kill chain (e.g., from reconnaissance to action), select targets for attack based on the characteristics of host nodes adjacent to the zombie nodes, allocate dangerous attack resources according to the attack budget, and launch dangerous attacks. The set of host nodes that are successfully compromised will be transformed into a new set of zombie nodes, and the next dangerous attack will be carried out.

[0047] The neighbor node path selection weights of the zombie node set controlled by a dangerous attacker are modeled as follows:

[0048]

[0049] Wherein, β1, β2, and β3 are used to balance the impact of node importance, node vulnerability, and attack cost considered by malicious attackers when selecting a path, respectively. i V is the importance value of host node i. i It is the vulnerability value of host node i. It is the attack cost required for a dangerous attacker to attack host node i.

[0050] The state transition diagram of the network kill chain is as follows: Figure 4 As shown, its transition matrix is ​​modeled as follows:

[0051]

[0052] Where, λ i and μ i Let P be the eigenvalue of the transition matrix of the kill chain in the i-th layer of the network. i [t], i∈[0,7] represents the transition probability of the kill chain in the i-th layer of the network at time slot t.

[0053] S2 Security Defender monitors attack behavior and obtains network status.

[0054] Security defenders utilize the network intrusion detection system deployed on them to obtain information about abnormal and dangerous attacks, and use this information to model the characteristics of the attacked host node, i.e., the compromised node i, into a binary tuple θ. i (t i risk i ), where t i Indicates the time of the attack, risk i This indicates the initial risk level of the attack.

[0055] The security defender determines the dynamic risk level p of the damaged node i in the current time slot t being attacked by zombie nodes. i [t] is modeled as an exponential function Introducing Response Indication Vector (DR) i[t] describes how security defenders utilize a limited defense budget. D Assign to damaged node i Resources.

[0056] In this step, at time slot t, the security defender observes the network state of the previous time slot t-1 and constructs the observation state function S[t]={I[t-1],V[t-1],c A [t-1],r A [t-1],d D [t-1],c D [t-1],r D [t-1], TD[t-1], p[t-1]}, where I[t-1] is the set of importance of host nodes in the network system at time slot t-1, V[t-1] is the set of vulnerability of host nodes in the network system at time slot t-1, and c A [t-1] represents the cost set of the dangerous attacker employing the dangerous attack strategy at time slot t-1, r A [t-1] represents the set of rewards obtained by the dangerous attacker at time slot t-1, d D [t-1] represents the set of resources allocated by the security defender to the compromised host node at time slot t-1, c D [t-1] represents the cost set of the security defense strategy adopted by the security defender at time slot t-1, r D [t-1] represents the reward obtained by the security defender at time slot t-1, TD[t-1] represents the set of host nodes that were damaged at time slot t-1, and p[t-1] represents the set of risk levels of the host nodes that were damaged at time slot t-1 under dangerous attacks.

[0057] In time slot t, the dynamic change formula for the vulnerability k of host node i is defined as follows:

[0058]

[0059] in, δ is the initial value of the vulnerability k of host node i, and δ is the discount factor of the time-varying vulnerability function. This represents the interaction between a dangerous attacker and a security defender regarding resource allocation on host node i. This refers to the amount of resources allocated by the security defender to the compromised host node i during time slot t. This refers to the amount of resources allocated by the attacker to the host node i targeted by the attack at time slot t. `m` controls the non-linear effect of the ratio of attack resources to defense resources. When `m` is large, the ratio of attack to defense in the formula has a more significant impact on the result; conversely, when `m` is small, the impact of the attack-to-defense ratio is relatively gradual, and its change has a smaller impact on the result. It is a constant value. This represents the maximum vulnerability growth time for the system's host nodes.

[0060] The S3 security defense device is based on a dual deep reinforcement learning algorithm, which dynamically generates the optimal defense strategy by combining network conditions and feeds it back to the security defender.

[0061] The security defense device uses the observed state function S[t] of the network state in time slot t to obtain the space set A[t]={d} of the network deception defense deployment strategy adopted in time slot t through a dual deep reinforcement learning algorithm. D [t],π[t]},d D [t] represents the resource allocation of the security defender at time slot t, and π[t] represents the network deception defense strategy selected by the security defender at time slot t, where π = Pπ i Let i ∈ [0,7]}, where i = 0 refers to the firewall, i = 1 refers to patching, i = 2 refers to redefining the encryption key, i = 3 refers to expelling the compromised node, i = 4 refers to deploying low-interaction and high-interaction honeypots, i = 5 refers to setting honey information, i = 6 refers to setting a fake key, and i = 7 refers to hiding the edges of the network topology. The utility of the security bundling defense strategy is calculated as follows: Among them DR i [t] represents the response indication vector.

[0062] Where TD[t] is the set of host nodes damaged at time slot t, and DR i [t] indicates whether the security defender takes action to protect the damaged host node i at time slot t. i [t] represents the risk level of the host node i being compromised at time slot t, indicating a dangerous attack. This refers to the amount of resources allocated by the security defender to the compromised host node i during time slot t. It is the cost of the security defender employing a security defense strategy against the compromised host node i during time slot t.

[0063] The dual-deep reinforcement learning algorithm of this invention employs two neural network architectures, including a main network and a target network. The main network selects the optimal action to execute based on the current state, while the target network evaluates the value of the action selected by the main network and calculates the target Q-value matrix accordingly. The main network is updated and trained at each time step, while the target network synchronously updates its weights at fixed time intervals to track the changes in the main network. By reducing the frequency of updates to the target network, dual-deep reinforcement learning effectively reduces estimation bias, thereby significantly improving the stability and convergence speed of the policy.

[0064] The Q-value matrix is ​​Q a(S[t], A[t]), which represents the Q value of the security defender when choosing the network deception defense deployment strategy space set A[t] under the function S[t] of the network state. The specific Q value matrix expression is as follows:

[0065] Q a (S[t],A[t])=(1-α)Q a (S[t],A[t])+α(R D [t+1]+γ·Q b (S[t+1],max a∈A Q a (S[t+1],A[t])));

[0066] Where α is the learning factor, γ is the discount factor, 0≤α≤1, 0≤γ≤1, R D [t+1] and S[t+1] represent the security defender's gains and state in time slot t+1, respectively. A[t] represents the security defender's action in time slot t. The max(.) function is the function to take the maximum value.

[0067] S4 Security Defenders utilize optimal defense strategies to dynamically respond to attacks in order to protect the network kill chain.

[0068] In this step, the security defender employs network deception defense strategies to respond to a surge in dangerous attacks. In time slot t, the space set of network deception defense deployment strategies adopted by the security defender is defined as A[t]={d D [t],dπ[t]},d D [t] represents the resource allocation of the security defender at time slot t, and π[t] represents the network deception defense strategy selected by the security defender at time slot t.

[0069] The security defense device also receives feedback from the security defenders on the defense results to adjust the optimal defense strategy, forming a closed-loop feedback mechanism.

[0070] This invention repeats the above steps until the system fails or time slot t equals the iteration termination condition T (e.g., T = 1000), at which point the iteration stops. The failure condition threshold is used to determine whether the network system has failed. Here, network system failure refers to the network system's inability to function properly or provide the expected service due to some reason (e.g., node attack, data leakage, or insufficient number of nodes to meet operational requirements). In other words, the system's operational logic is halted, preventing further attack and defense interactions, and potentially causing network service paralysis or interruption, affecting actual business or functional requirements. The failure threshold conditions are as follows:

[0071]

[0072] Where Th1 and Th2 are the information importance leakage threshold and the available node threshold, respectively, I i |VN0| is the importance value of host node i in the network system, and |VN0| is the total number of host nodes in the initial network system before being attacked by a malicious attacker. t | is the number of all available host nodes in the network system at time slot t, N i N represents the security, damage, or eviction status of a host node. i =1 indicates that the host node has been damaged or evicted; otherwise, N i =0.

[0073] If the network system fails or the maximum number of iterations is reached, the process ends. Otherwise, it is determined whether the attack by the dangerous attacker was successful. If the attack was unsuccessful, the current stage of the network kill chain continues. If the attack was successful, the process proceeds to the next stage of the network kill chain.

[0074] Based on this, see Figure 1 , Figure 2 This invention also proposes a dynamic defense system for APT network kill chains based on dual deep reinforcement learning, built on an SDN architecture, and employing the aforementioned dynamic defense method for APT network kill chains based on dual deep reinforcement learning, including:

[0075] A host node consists of multiple host nodes with different levels of importance and known vulnerabilities.

[0076] A dangerous attacker executes attacks on host nodes according to the network kill chain.

[0077] The security defender monitors attack behavior and acquires network status, dynamically responding to attacks using optimal defense strategies to protect the network kill chain. The SDN controller serves as the core component of the security defender. The defender deploys a network intrusion detection system to monitor network traffic in real time and responds to changes in complex network topology through dynamic security defense devices.

[0078] This security defense device targets diverse attacks launched by dangerous APT attackers at different stages of the network kill chain. Based on a dual deep reinforcement learning algorithm, it dynamically generates optimal defense strategies by combining network conditions and feeds these strategies back to security defenders. For example, it outputs honeypot defense strategies, allowing security defenders to use honeypots to capture dangerous attackers, analyze their attack behavior, and generate decoy data.

[0079] The security defense device of the present invention can acquire network status in real time and generate security defense strategies. A schematic diagram of the specific device is shown below. Figure 2As shown, the system mainly includes an analysis module and a defense strategy generation module. The analysis module is responsible for preprocessing the real-time network traffic input to the system by the security defender, extracting and generating real-time network status information, including the importance and vulnerability values ​​of each node. The defense strategy generation module, based on a dual deep reinforcement learning algorithm, dynamically generates the optimal defense strategy by combining the threat level of dangerous attack behaviors and the network status, and sends it to the security defender for real-time protection. Simultaneously, the defense strategy generation module receives feedback on the defense results, forming a closed-loop feedback mechanism, continuously improving the system's defense performance through continuous learning and optimization. In the dynamic defense device of this invention, the dual deep reinforcement learning algorithm is its core component, which can optimize the allocation of computing and time resources and select and deploy network deception defense strategies such as honeypots. Specific deployment schemes are common to the specific steps of the above methods.

[0080] Ultimately, security defenders utilize the generated optimal defense strategies to dynamically respond to dangerous attacks, thereby achieving precise protection against APT attacks and improving the intelligence and adaptability of the defense. This device integrates real-time traffic detection, automatic generation and distribution of defense strategies, and dynamic optimization of defense effectiveness, effectively enhancing the intelligence level and response efficiency of network security systems and significantly strengthening their ability to cope with complex network threats.

[0081] The above are merely specific embodiments of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantial modifications made to the present invention using this concept shall be considered as infringing upon the protection scope of the present invention.

Claims

1. A dynamic defense method for APT network kill chains based on dual deep reinforcement learning, characterized in that, A pre-configured network system is included, comprising host nodes, security defenders, malicious attackers, and security defense devices. The attacker performs attacks on the host node according to the network kill chain. The security defender monitors the attack and obtains the network status. The security defense device dynamically generates the optimal defense strategy based on the dual deep reinforcement learning algorithm and combines the network status, and feeds it back to the security defender. The security defender uses the optimal defense strategy to dynamically respond to the attack to achieve protection against the network kill chain. In time slot t, the security defender observes the network state of the previous time slot t-1 and constructs the observation state function S[t] for time slot t, where S[t] = {I[t-1], V[t-1], c A [t-1], r A [t-1], d D [t-1], c D [t-1], r D [t-1], TD[t-1], p[t-1]}, I[t-1] is the set of importance of the host nodes of the network system at time slot t-1, V[t-1] is the set of vulnerability of the host nodes of the network system at time slot t-1, c A [t-1] represents the cost set of the dangerous attacker employing the dangerous attack strategy at time slot t-1, r A [t-1] represents the set of rewards obtained by the dangerous attacker at time slot t-1, d D [t-1] represents the set of resources allocated by the security defender to the compromised host node at time slot t-1, c D [t-1] represents the cost set of the security defense strategy adopted by the security defender at time slot t-1, r D [t-1] represents the reward obtained by the security defender at time slot t-1, TD[t-1] represents the set of host nodes that were damaged at time slot t-1, and p[t-1] represents the set of risk levels of the host nodes that were damaged at time slot t-1 under dangerous attacks; The security defense device, based on the observed state function S[t] of the network state in time slot t, uses a dual deep reinforcement learning algorithm to obtain the spatial set A[t] = {d} of the network deception defense deployment strategy adopted in time slot t. D [t],π[t]},d D [t] represents the resource allocation of the security defender at time slot t, π[t] represents the network deception defense strategy selected by the security defender at time slot t, and the utility of the security bundling defense strategy is calculated as follows: Among them DR i [t] represents the response indication vector; Where TD[t] is the set of host nodes damaged at time slot t, and DR i [t] indicates whether the security defender takes action to protect the damaged host node i at time slot t. i [t] represents the risk level of the host node i being compromised at time slot t, indicating a dangerous attack. This refers to the amount of resources allocated by the security defender to the compromised host node i during time slot t. It is the cost of the security defender employing a security defense strategy against the compromised host node i at time slot t; The dual deep reinforcement learning algorithm includes a main network and a target network. The main network selects the optimal action to perform based on the current state, while the target network evaluates the value of the action selected by the main network and calculates the target Q-value matrix accordingly. The main network is updated and trained at each time step, and the target network synchronously updates its weights at fixed time intervals to track the changes of the main network. The Q-value matrix is ​​Q a (S[t], A[t]), which represents the Q value of the security defender when selecting the network deception defense deployment strategy space set A[t] under the observed state function S[t] of the network state. The expression for updating the Q value matrix is ​​as follows: Where α is the learning factor, γ is the discount factor, and R0 D [t+1] and S[t+1] represent the utility and state of the security defender in time slot t+1, respectively. A[t] represents the action of the security defender in time slot t. The max(.) function is the function to take the maximum value.

2. The dynamic defense method for APT network kill chains based on dual deep reinforcement learning as described in claim 1, characterized in that, The attack on the host node by the dangerous attacker according to the network kill chain specifically refers to: The attacker gradually utilizes the set of zombie nodes according to each stage of the network kill chain, selects targets for attack based on the characteristics of host nodes adjacent to the zombie nodes, allocates dangerous attack resources according to the attack budget, and launches dangerous attacks. The set of host nodes that are successfully damaged will be transformed into a new set of zombie nodes.

3. The dynamic defense method for APT network kill chains based on dual deep reinforcement learning as described in claim 1, characterized in that, The security defender models the dynamic risk level of the host node being attacked by zombie nodes as an exponential function; it introduces a response indication vector to describe the resources that the security defender allocates to the damaged host node using a limited defense budget.

4. The dynamic defense method for APT network kill chains based on dual deep reinforcement learning as described in claim 1, characterized in that, The security defense device also receives defense results from the security defender to adjust the optimal defense strategy, forming a closed-loop feedback mechanism.

5. The dynamic defense method for APT network kill chains based on dual deep reinforcement learning as described in claim 1, characterized in that, It also includes determining whether the network system has failed based on failure condition thresholds. Network system failure refers to the network system's operational logic ceasing, making further attack and defense interactions impossible. The failure condition thresholds are as follows: Where Th1 and Th2 are the information importance leakage threshold and the available node threshold, respectively, I i |VN0| is the importance value of host node i in the network system, and |VN0| is the total number of host nodes in the initial network system before being attacked by a malicious attacker. t | is the number of all available host nodes in the network system at time slot t, N i N represents the security, damage, or eviction status of a host node. i =1 indicates that the host node has been damaged or evicted; otherwise, N i =0; If the network system fails or reaches the maximum number of iterations, the process ends; otherwise, it is determined whether the attacker succeeded. If the attack failed, the current stage of the network kill chain continues; if the attack succeeded, the process proceeds to the next stage of the network kill chain.

6. A dynamic defense system for APT network kill chains based on dual deep reinforcement learning, characterized in that, A method for dynamic defense of APT network kill chains based on dual deep reinforcement learning, as described in any one of claims 1 to 5, includes: host node, A dangerous attacker executes attack actions against the host node according to the network kill chain; The security defender monitors the attack behavior and obtains the network status, and uses the optimal defense strategy to dynamically respond to the attack behavior in order to protect the network kill chain. The security defense device, based on a dual deep reinforcement learning algorithm, dynamically generates the optimal defense strategy by combining the network state and feeds it back to the security defender.

Citation Information

Patent Citations

  • Honeypot deception defense method, device and equipment and storage medium

    CN116614296A

  • Method and device for generating network killing chain and electronic equipment

    CN117395072A

  • Attack and defense evolution game based network safety situation assessment method and system

    CN108512837A

  • Dynamic industrial control honeypot deployment method based on deep reinforcement learning

    CN117792749A