Multi-step attack decoy method and system with dynamic deception defense enablement

By generating relevant honey baits in a cloud-native environment and utilizing security game modeling and reinforcement learning algorithms to optimize the honey bait deployment strategy, the problems of insufficient generalization capability and dynamic decision-making evaluation limitations of cloud-native deception defense technology are solved, achieving effective defense against APT attacks.

CN119966674BActive Publication Date: 2025-10-10Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510010095.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-10-10
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

Existing cloud-native deception defense technologies lack generalization capabilities when facing different cloud service providers and diverse attack scenarios. They lack comprehensive consideration of multiple attack modes and defense strategies, have limitations in dynamic decision-making and evaluation, and cannot effectively defend against the hidden characteristics of APT attacks.

Method used

By generating different types of honey baits (operating system honey baits, process honey baits, and service honey baits), combining security game modeling and reinforcement learning algorithms, optimizing honey bait deployment strategies, and utilizing dynamic adjustments of honeychain defense resources, timely detection and induction of attack behaviors can be achieved.

Benefits of technology

It improves the security protection capabilities of cloud-native environments, can effectively lure attackers into preset traps, and achieve timely detection and defense of attack behaviors, and is suitable for distributed system security protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119966674B_ABST
    Figure CN119966674B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of network security, and particularly relates to a multi-step attack trapping method and system enabled by dynamic deception defense, different types of honeypots are generated according to the resources required by microservices and the relevance between microservices to perceive network attack behaviors by using the honeypots; a security game model for describing the attack and defense confrontation process of the honeynet is constructed according to the network scale and defense requirements, the security game model is represented by using attack and defense participants, the current state of the network environment, the attack and defense action set, the discount factor, the network state transition probability and the attack and defense payoff function; the security game model is optimized and solved by using a reinforcement learning algorithm to obtain the optimal number and position of the honeypots in the spatial dimension, the honeypots are deployed into the network according to the optimal number and position, and the state and position of the honeypots in the network are dynamically adjusted according to a hopping strategy. The present application can effectively induce the attacker to enter a preset line diameter, realize the timely discovery of attack behaviors and improve the network security protection performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to a multi-step attack trapping method and system enabled by dynamic deception defense. Background Art

[0002] As global digitalization continues to increase, cloud computing has become a critical infrastructure for information technology development. Since Pivotal first proposed the cloud-native architecture in 2013, cloud-native technologies such as containers and microservices have enabled rapid deployment and elastic scaling, offering advantages such as high scalability, high concurrency, and high availability. However, the widespread adoption of cloud-native technologies has brought with it complex security threats, including cyberattacks, program vulnerabilities, identity authentication and access control, and data leaks. These threats can expose enterprises to supply chain and operational risks. Due to significant information asymmetry, time asymmetry, and cost asymmetry between attackers and defenders, attackers can exploit vulnerabilities in containers and microservices to covertly move laterally, escalate privileges, and ultimately steal sensitive data. Furthermore, the constant evolution of cloud-native applications increases the uncertainty of the attack surface, while the interactions and dependencies between microservices further exacerbate the challenges of designing effective defense strategies.

[0003] Network deception defense increases attack complexity and cost through covert obfuscation and dynamic migration. Moving target defense, a game-changing technology, enhances the uncertainty and randomness of the network, enhancing the inherent system defense capabilities. Both approaches offer a new approach to proactive security protection for information and communication systems without changing the existing network framework. However, cloud-native deception defense technologies lack generalizability when applied to deception defense scenarios across different cloud service providers and diverse attack scenarios. For example, the existing Duplicity game framework is a game-theoretic deception defense decision-making method. By expanding the application scenarios of deception technology and exploring how to design deception mechanisms when goals are inconsistent, the game-theory-based zero-trust security framework GAZETA is proposed. This framework uses a dynamic game model to design interdependent trust assessment and authentication strategies. However, it has limitations in dynamic decision-making and evaluation during deployment and application. Research on this framework often focuses on specific attack scenarios or defense mechanisms, lacking comprehensive consideration of multiple attack modes and defense strategies. Existing game models also lack dynamicity and real-time performance. Another example is the dynamic game-based APT attack detection solution, which uses a hill-climbing policy algorithm to increase the uncertainty of defense strategies and deceive attackers. Compared to the Q-learning algorithm solution, this approach effectively improves APT attack detection performance, but it cannot effectively defend against the stealthy nature of APT attacks. Therefore, a defense method that can perceive penetration threats is urgently needed to enhance the security of cloud-native environments. Summary of the Invention

[0004] To this end, the present invention provides a multi-step attack trapping method and system enabled by dynamic deception defense to solve the problems of insufficient generalization capabilities of existing microservice deception defense attack and defense scenarios and lack of comprehensive consideration of multiple attack modes and defense strategies. By modeling the honey bait attack and defense process through security game modeling and combining reinforcement learning algorithms to calculate the optimal strategy for deploying honey baits on the target network, attackers are effectively induced to follow the preset path, thereby achieving timely and effective discovery of attack behaviors.

[0005] According to the design scheme provided by the present invention, on the one hand, a multi-step attack trapping method with dynamic deception defense enablement is provided, comprising:

[0006] Different types of honey baits are generated based on the resources required by microservices and the correlation between microservices. The honey baits are used to perceive network attack behaviors. The honey baits include operating system honey baits, process honey baits, and service honey baits. The honey baits are correlated with each other in terms of network resource configuration.

[0007] A security game model is constructed based on the network scale and defense requirements to describe the honeylink attack and defense process. The security game model is represented by the attack and defense participants, the current state of the network environment, the attack and defense action set, the discount factor, the network state transition probability, and the attack and defense payoff function. The current state of the network environment includes the number of honey baits in the current network, the location of the defender's honey bait nodes, the location of the attacker's penetration nodes, and the location of the vulnerability.

[0008] A reinforcement learning algorithm is used to optimize and solve the security game model to obtain the optimal number and location of honey bait deployment in the spatial dimension. The honey bait is deployed in the network based on the optimal number and location, and the status and location of the honey bait in the network are dynamically adjusted according to the hopping strategy. The hopping strategy is used to adjust the honey bait in the time dimension for network defense.

[0009] As a multi-step attack trapping method enabled by dynamic deception defense of the present invention, different types of honey baits are further generated based on the resources required by microservices and the correlation between microservices, including:

[0010] Creating a honey bait network resource based on microservice resource elements, wherein the microservice resource elements include an operating system, a service, and a process;

[0011] The network resource type and quantity in the honey bait are configured according to the microservice correlation parameters, and the honey bait configuration elements are integrated into the host to generate the corresponding honey bait. The correlation parameters include a probability parameter for indicating that the microservice selects a new configuration, an association parameter for indicating whether there is an association between microservice network resources, and a Poisson distribution parameter for indicating the configuration quantity system. The host is a virtual or physical network environment entity consistent with an ordinary network host.

[0012] As a multi-step attack trapping method enabled by dynamic deception defense in the present invention, a security game model for describing the honeylink attack and defense confrontation process is further constructed based on the network scale and defense requirements, including:

[0013] The network environment is represented as a graph data consisting of a set of nodes and a set of edges, and the adjacency matrix of the network topology is used to represent the connectivity of the network.

[0014] The path that the attacker takes to move laterally in the network and reach the target node is considered the attack path. The attacker's state is represented by the attacker's current node, the set of nodes the attacker can access, and the degree of proximity to the target node. The attacker's actions are represented by the current state of the network node and the attack target.

[0015] The defender's state is obtained based on the attacker's known attack actions and the honey bait information deployed on the network. The defender's actions are represented by the location of the honey bait deployed in the network. The defender's state and actions are used to represent the attacker's attack strategy. The defender's current network state is represented by the attacker's attack strategy to select the defense strategy for deploying the honey bait node.

[0016] Construct a profit function for both attackers and defenders based on the rewards they receive from taking corresponding offensive and defensive actions under the current network state.

[0017] Based on the attack path, attack and defense status, attack and defense actions, attack and defense strategies and attack and defense benefit functions, a security game model is established to describe the honeychain attack and defense confrontation process.

[0018] As a multi-step attack trapping method enabled by dynamic deception defense of the present invention, a security game model for describing the honeylink attack and defense confrontation process is further constructed according to the network scale and defense requirements, which also includes:

[0019] Divide the roles of participants in the honeylink attack and defense process into leaders and followers;

[0020] According to the game order of the attack and defense confrontation process, the leader makes the decision first, and the followers respond accordingly based on the leader's decision. When the leader makes a decision, he maximizes his own interests by predicting the followers' action responses. The followers choose the best strategy to perform the corresponding actions based on the leader's decision.

[0021] As a multi-step attack trapping method enabled by dynamic deception defense of the present invention, a reinforcement learning algorithm is further used to optimize and solve the security game model, including:

[0022] Randomly select a network state and input it into the Actor network. The Actor network selects an action based on the input network state and executes the action in the network environment, so that the network environment outputs the next network state and reward. The selected network state, selected action, and output next network state and reward are used as training data and put into the experience replay pool until the number of training data in the experience replay pool reaches the threshold.

[0023] The critic network is trained using the training data in the experience replay pool, the Q-value function is used to carefully evaluate the current state-action strategy, and the critic network parameters are updated by minimizing the mean square error loss.

[0024] Until the model converges to Nash equilibrium, and based on the Nash equilibrium state, the optimal number and position of honey bait deployment in the spatial dimension are obtained.

[0025] As a multi-step attack trapping method enabled by dynamic deception defense of the present invention, the state and position of the honey bait in the network are dynamically adjusted according to the hopping strategy, including:

[0026] IP hopping is adopted as the hopping strategy. Based on the hopping strategy, the host addresses in the network subnet are randomly remapped and the honey bait network configuration in the network is dynamically changed.

[0027] As a multi-step attack trapping method enabled by dynamic deception defense of the present invention, the jumping strategy further includes: setting a time interval hyperparameter for triggering IP jumping, when the time interval is reached, triggering the target host to randomly generate a new IP address, and checking whether the new IP address has been used by other hosts. If it has been used, a new IP address is regenerated; if it has not been used, the randomly generated new IP address is assigned to the target host, and the host address mapping relationship in the network is updated.

[0028] On the other hand, the present invention also provides a multi-step attack trapping system with dynamic deception defense capability, comprising: a honey bait configuration module, a model setting module and an attack defense module, wherein:

[0029] A honey bait configuration module is used to generate different types of honey baits based on the resources required by microservices and the correlation between microservices, so as to use the honey baits to perceive network attack behaviors. The honey baits include operating system honey baits, process honey baits, and service honey baits, and the honey baits are correlated with each other in terms of network resource configuration;

[0030] The model setting module is used to construct a security game model for describing the honeychain attack and defense confrontation process based on the network scale and defense requirements. The security game model is represented by the attack and defense participants, the current state of the network environment, the attack and defense action set, the discount factor, the network state transition probability, and the attack and defense benefit function. The current state of the network environment includes the number of honey baits in the current network, the location of the defender's honey bait nodes, the location of the attacker's penetration nodes, and the location of the vulnerability.

[0031] The attack and defense module is used to optimize and solve the security game model using a reinforcement learning algorithm, obtain the optimal number and location of honey bait deployment in the spatial dimension, deploy the honey bait into the network based on the optimal number and location, and dynamically adjust the status and location of the honey bait in the network according to the hopping strategy. The hopping strategy is used to adjust the honey bait in the time dimension for network defense.

[0032] In another aspect, the present invention further provides a multi-level defense system, comprising: a defense platform and a protected server, wherein:

[0033] The defense platform is used to generate honey baits using the above method and configure the honey bait hosts as deceptive resources to disguise themselves as nodes in the target network. The honey bait hosts are connected in series to form a spatiotemporal honeychain and distributed at different locations along the attacker's intrusion depth. The network addresses of the honey bait hosts are set to a random hopping pattern to present a dynamic false network view to the attacker.

[0034] Protected servers include: legitimate users and / or network servers in network services.

[0035] Beneficial effects of the present invention:

[0036] The present invention deploys a series of deceptive honey baits on potential attack paths, uses security game modeling to model the attack and defense process of honey baits, combines reinforcement learning algorithms to schedule honey baits, combines spatiotemporal honey chain defense resource deployment, integrates and promotes the concepts of dynamic defense and deception defense, and uses simulation methods to evaluate the deployment of temporal and spatial deception resources in the network, optimizes the configuration required for detecting honey baits and implementing mobile target defense in the network, and solves the problems of insufficient generalization ability of deception defense attack and defense scenarios in cloud native environments, inability to effectively reduce APT covert characteristic attacks, limitations of dynamic decision-making evaluation, and lack of multiple attack modes and defense strategies. It can accurately calculate the optimal number and location of honey baits to deploy on the target network, effectively induce attackers to step into preset traps, and realize timely detection of attack behaviors. It can not only improve security protection in cloud native environments, but also provide reference for security protection of other distributed systems. It has great application prospects in the field of network security. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1This is a schematic diagram of the multi-step attack trapping process that enables dynamic deception defense in an embodiment;

[0038] Figure 2 This is a schematic diagram of the division of covert infiltration stages in the embodiment;

[0039] Figure 3 Schematic diagram of the honey bait generation process in the embodiment;

[0040] Figure 4 This is an illustration of the cloud-native spatiotemporal honeylink principle in the embodiment;

[0041] Figure 5 This is a schematic diagram of the DDPG algorithm process in the embodiment;

[0042] Figure 6 This is a schematic diagram of the overall structure of the spatiotemporal honeylink in the embodiment;

[0043] Figure 7 This is a schematic diagram of the spatiotemporal honeychain defense architecture in the embodiment. DETAILED DESCRIPTION

[0044] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention is further described in detail below with reference to the accompanying drawings and technical solutions.

[0045] In view of the fact that the current deception defense technology cannot solve the problem of insufficient generalization capability of deception defense attack and defense scenarios and lack of comprehensive consideration of multiple attack modes and defense strategies, the embodiment of the present invention, see Figure 1 As shown, a multi-step attack trapping method with dynamic deception defense capability is provided, including:

[0046] S101. Generate different types of honey baits based on the resources required by microservices and the correlation between microservices, so as to use the honey baits to perceive network attack behaviors. The honey baits include operating system honey baits, process honey baits, and service honey baits, and the honey baits are correlated in network resource configuration.

[0047] Specifically, different types of honey bait are generated based on the resources required by microservices and the correlation between microservices, including:

[0048] Creating a honey bait network resource based on microservice resource elements, wherein the microservice resource elements include an operating system, a service, and a process;

[0049] The network resource type and quantity in the honey bait are configured according to the microservice correlation parameters, and the honey bait configuration elements are integrated into the host to generate the corresponding honey bait. The correlation parameters include a probability parameter for indicating that the microservice selects a new configuration, an association parameter for indicating whether there is an association between microservice network resources, and a Poisson distribution parameter for indicating the configuration quantity system. The host is a virtual or physical network environment entity consistent with an ordinary network host.

[0050] Due to the complexity of microservice architectures, attackers can exploit the call relationships between microservices to move laterally between them, gaining access to the host, stealing data, and attacking other containers. At the network level, attackers can conduct eavesdropping scans during communications to eavesdrop on data or discover vulnerabilities in microservices and containers. Attacks that exploit unknown vulnerabilities are the most difficult to prevent.

[0051] ATT&CK (Adversarial Tactics, Techniques, and Common Knowledge) is an attack model framework developed by the MITRE organization. It consists of two main parts: attack tactics (Tactics) and attack techniques (Techniques). According to the 14 tactics and their functions in ATT&CK, such as Figure 3 As shown in the figure, cloud native environment penetration threats are divided into three specific attack methods:

[0052] 1. Attack scanning, which can proactively detect vulnerability fingerprints of network resources;

[0053] 2. Vulnerability exploitation: using known vulnerabilities to penetrate and elevate privileges to achieve attack control;

[0054] 3. Covert infiltration, stealing unencrypted sensitive data through lateral movement.

[0055] Based on the information asymmetry between attackers and defenders regarding vulnerability information, we can categorize attackers' attacks into the following two types:

[0056] Type 1 known unknown attack refers to the attacker having access to a common vulnerability library, but the defender cannot grasp the number and specific types of vulnerabilities in the vulnerability library exploited by the attacker. Therefore, there is a certain probability that the attacker will successfully penetrate the library by exploiting these vulnerabilities.

[0057] Type 2, unknown unknown attacks, refers to attacks in which the attacker possesses a larger library of vulnerabilities (i.e., 0-day vulnerabilities) than the defender. When the attacker exploits a 0-day vulnerability, the penetration success rate is assumed to be 100%. When the attacker exploits a vulnerability that the defender also possesses, the attacker also has a certain probability of successful penetration.

[0058] Based on the penetration capability of a certain type of attacker, the penetration strategy is divided into three stages:

[0059] Phase 1, a global penetration strategy, begins by performing all possible scans on all existing containers, including the operating system, microservices, vulnerabilities, and processes. A container is then randomly selected and the optimal vulnerability is exploited to gain user or root privileges. If the attacker gains user privileges, they escalate privileges to exploit the vulnerability. If the exploitation is successful, a process scan is performed. If root privileges are obtained, data theft is carried out.

[0060] The second phase uses a specific penetration strategy. First, a container scan is performed to obtain the IP address of the active container. Then, a container is randomly selected and only the operating system, microservices, and processes of the container are scanned. Finally, vulnerability exploitation and data theft are performed based on the permission level of the scanned container.

[0061] Phase three involves a random penetration strategy, which first scans containers in the cloud-native environment, exploits known vulnerabilities to elevate privileges and implement attack control, and then penetrates all containers in random order. If penetration is successful, data is stolen.

[0062] There is an attack-defense mapping relationship between ATT&CK and Engage. Adversary Engagement is the comprehensive use of network blocking (Denial) and network deception (Deception) in the context of strategic planning and analysis. By utilizing vulnerabilities and lures (Introduced Vulnerabilities and Lures), the methods of detection (Detect), diversion (Direct), destruction (Disrupt), and motivation (Motivate) can be implemented simultaneously to achieve the goals of exposure (Expose), influence (Affect), and elicit (Elicit), exposing deception resources, influencing attacker decisions, and attracting attacks. Therefore, in the embodiment of this case, a honey bait deception unit can be designed according to the guidance of the Engage adversarial framework to achieve the purpose of denying attacks and confusing opponents.

[0063] To illustrate the deceptive nature of honey bait, we break it down into three main properties: simulation, relevance, and diversity. They are as follows:

[0064] Property 1: Simulation: Simulation is a requirement for the underlying design of the honey bait. The construction of the honey bait is based on normal microservice applications, and different network resources are configured to generate the honey bait, so as to ensure that the honey bait is consistent with the normal host in appearance and can be disguised as a real microservice.

[0065] Property 2: Connectivity: The degree of connectivity between microservices simulates the varying scales of microservice applications in real networks. A low degree of connectivity represents the independent configuration of individual microservices in real networks, while a high degree of connectivity represents a relatively large scale of microservices within a network.

[0066] Diversity: Diversity refers to designing different honey baits based on the proportion of vulnerabilities in different aspects. For related honey baits, the number of vulnerabilities in their network resources can be configured to increase the proportion of vulnerabilities targeting a specific network resource. This creates greater diversity in certain aspects, making it more likely to attract corresponding attack measures, thereby enabling targeted defense.

[0067] Based on the above characteristics, the honey bait construction process is as follows Figure 3 As shown, the first step is to generate network resources. The essential resources for configuring a functioning microservice are an operating system (OS), services, and processes. Available operating systems include Windows, Linux, and Unix; services include SSH, FTP, Http, DNS, and LDAP; and processes include Tomcat, Daclsvc, Schtask, and DNS.

[0068] By configuring the necessary network resources for the honey bait, we can make it appear functionally consistent with normal microservices, thus meeting the simulation requirements. Next, we consider the correlation between microservices. The correlation parameters can be summarized as follows:

[0069] (alpha_H). Represents the probability of a microservice selecting a new configuration. When configuring a microservice, if it is the first microservice, a new configuration will be selected for the microservice.

[0070] (alpha_V). Controls whether a network resource in a microservice is reselected or selected from the previous microservice configuration options. This determines whether the network resource of the microservice is associated with the previous microservice application.

[0071] (lambda_V). Represents the Poisson distribution parameter, which controls the number of configuration options. It is at least 1. A larger value means that the configuration contains more options.

[0072] By configuring the above parameters, a set of honey baits with correlation in network resource configuration can be obtained. The pseudo code of the honey bait generation algorithm can be shown as Algorithm 1:

[0073]

[0074]

[0075] The construction of the honeypot first needs to generate the basic network resources required by the honeypot, including the operating system, services and processes. These resources are generated using predefined functions (such as lines 1-3 of the code). Subsequently, line 4 configures the specific types and quantities of network resources in the honeypot according to the correlation parameters (defined by variables alpha_h, alpha_v and lambda_v). This is achieved through the GenerateCorrelatedHosts() function, which ensures that the honeypot exhibits a real network environment with interconnected resources. Finally, lines 5-14 of the code integrate the generated configuration elements into a host that is indistinguishable from a normal network host. This is achieved by creating a new host object with parameters including network address, operating system, services, processes, value, discovery value, credentials and vulnerabilities. In this way, a complete honeypot configuration is obtained, which can be deployed in a network environment.

[0076] The selection of network resource configurations with correlation lays the foundation for the diversity of honeypots. The diversity of honeypot deception is mainly reflected in the deployment of vulnerability diversity. The higher the proportion of vulnerabilities in a certain network resource configuration of the honeypot, the more the honeypot deception focuses on that network resource configuration. By the proportion of vulnerabilities in different network resource configurations, the honeypot can be divided into operating system honeypot (OS HoneyTrap), process honeypot (Process HoneyTrap) and service honeypot (Service HoneyTrap). Different types of honeypots have different defense effects against different attack means, and the honeypot chain defense strategy is to deploy different types of correlated honeypots in the network environment.

[0077] S102, constructing a security game model for describing the attack and defense confrontation process of the honeynet according to the network scale and defense requirements, wherein the security game model uses attack and defense participants, current network environment state, attack and defense action set, discount factor, network state transition probability and attack and defense payoff function to represent, wherein the current network environment state includes the number of honeypots in the current network, the location of the deployed honeypot nodes by the defender, the location of the penetrated nodes by the attacker and the location of the vulnerabilities.

[0078] Specifically, constructing a security game model for describing the attack and defense confrontation process of the honeynet according to the network scale and defense requirements can include:

[0079] The network environment is represented as a graph data composed of a node set and an edge set, and the connectivity relationship of the network is represented using the adjacency matrix of the network topology;

[0080] The path that the attacker takes to move laterally in the network and reach the target node is considered the attack path. The attacker's state is represented by the attacker's current node, the set of nodes the attacker can access, and the degree of proximity to the target node. The attacker's actions are represented by the current state of the network node and the attack target.

[0081] The defender's state is obtained based on the attacker's known attack actions and the honey bait information deployed on the network. The defender's actions are represented by the location of the honey bait deployed in the network. The defender's state and actions are used to represent the attacker's attack strategy. The defender's current network state is represented by the attacker's attack strategy to select the defense strategy for deploying the honey bait node.

[0082] Construct a profit function for both attackers and defenders based on the rewards they receive from taking corresponding offensive and defensive actions under the current network state.

[0083] Based on the attack path, attack and defense status, attack and defense actions, attack and defense strategies and attack and defense benefit functions, a security game model is established to describe the honeychain attack and defense confrontation process.

[0084] like Figure 4 As shown, honey baits are designed as individual deception defense resources tailored to specific attacker exploits. To present a false view to attackers, diverse and relevant honey baits must be deployed throughout the network topology at different attack stages. Honey bait deployment strategies require careful consideration of both quantity and location. Too few honey baits will fail to attract attackers, rendering the honeychain ineffective. Excessive honey baits are costly and may alert attackers, exposing the existence of the honeychain. Optimizing honey bait placement is also crucial. Placing honey baits at key nodes in the network topology or along key routes along deep attack paths can significantly improve the timeliness of attack detection. Therefore, honeychain construction requires game modeling based on network scale and defense requirements.

[0085] In this example, we analyze attack behavior and honey bait defense mechanisms based on a security game model. Participants in the model are divided into two distinct roles: leader and follower. The game proceeds sequentially, with the leader making decisions first and the followers subsequently reacting accordingly. When making decisions, the leader anticipates the followers' reactions to maximize their own interests; followers then choose the optimal strategy based on the leader's decisions. The model can be represented by the sextuple SGMDP-S = (N, S, A, γ, T, R), as follows:

[0086] 1) N = (N p ,N Q ) represents the attacking and defending sides, where N p represents the attacker, N QIn this setting, we assume that there is a single attacker and a single defender, so N=2.

[0087] 2) Describes the current state of the network environment, including the number of honey baits m, the node location np where the defender deploys the honey baits, the node location nq where the attacker penetrates, and the vulnerability location n.

[0088] 3)A(A p ,A q ) represents the set of actions of the attacker and the defender, where is the action selected by the attacker from his infiltration strategy library, is the action selected by the defender from its library of honey bait position change strategies.

[0089] 4)γ is the discount factor, which reflects that the returns of both the attacker and the defender depend not only on the immediate gains, but also on the potential future gains.

[0090] 5) T represents the transition probability, which represents the probability of transitioning from the current state s to the next state s′ in matrix form. Specifically, T is the probability of transitioning to state s′ after performing action a in state s. The function is expressed as follows:

[0091] T=T(S t+1 =s'|S t =s,A t =a)

[0092] 6) R is the reward function, which defines the reward obtained by the defender for taking action A in state S.

[0093] To simulate the attack-defense game more realistically, we first adopt single-step reward scoring. We first discretize the game process, consider the impact of each step on the rewards of both parties, and then accumulate these impacts to construct a state-value function.

[0094] When the attacker attacks different ordinary hosts, he will receive immediate rewards. However, due to the different number of known and unknown vulnerabilities in each host, the difficulty and cost of breaking through are also different. If the attacker successfully attacks the target host, he will receive a large number of points and the experiment will be terminated. If the attacker attacks a host covered by honey bait, points will be deducted, indicating that the defender has discovered the attacker and the experiment will be terminated. In addition, the attacker is subject to a step size limit during the attack. The initial score is set to, and the incremental score is subtracted for each attack step. The shorter the attack step, the better the attack effect; exceeding the limit is considered an attack failure and the experiment is terminated. So the attacker reward function is as follows:

[0095]

[0096] Deploying and switching baits incurs costs for the defender, and the number of baits varies across the game. To simplify the calculation, the process is discretized, calculating the cost of deploying baits in a single step. If the bait captures the attacker or the attacker exceeds the time limit, the defense is considered successful, points are awarded, and the experiment ends. If the target host is compromised, points are deducted. The step length required to capture the attacker is also used to judge the effectiveness of the defense. The initial score is set to , and the defender's score decreases with each step. This gives the defender's reward function:

[0097]

[0098] To evaluate the agent's strategy, the state-value function of the attacker and defender is constructed, which can be expressed as follows:

[0099]

[0100] In the initial state S, the attacker and defender use their respective strategies π to achieve maximum rewards.

[0101] Following the principle of Occam's razor, it is advocated to pursue a simple and effective method in model construction. Previous studies have generally been committed to building game models by meticulously simulating the real network attack and defense process, aiming to design attack and defense strategies, states, and benefits by simulating the real network environment, and to establish a target benefit function, and then use reinforcement learning to solve it. However, the attack and defense process in the real network environment is extremely complex and difficult to simulate completely. Therefore, it is necessary to reasonably abstract the entire process and environment. Although there are limitations in using Occam's razor to simplify the model in the process of inverse reinforcement learning (IRL) and irrational agent training, studies have shown that when the model is known and solvable, complex reinforcement learning methods can be avoided. In this embodiment, Occam's razor is applied for specification modeling, and the simplest features are selected to design the reward function, with the aim of simplifying the attack and defense process into an easy-to-handle model.

[0102] First, the network environment and policy rules simulated by the model can be abstracted. Specifically, we consider a network environment based on a microservice container architecture, where nodes are not only interconnected but also capable of communication. Each node is equipped with unique network resources and potential security vulnerabilities. In the attack strategy, after scanning the network, the attacker exploits these vulnerabilities to move laterally between nodes, ultimately aiming to escalate privileges on nodes storing sensitive data and achieve their attack objectives. It is worth noting that the probability of an attacker successfully exploiting a vulnerability varies, and the specific details of the exploitation are not a decisive factor in the attacker's lateral movement. Therefore, when simplifying the model, the focus is on how the attacker moves laterally through the network and reaches the target node. In this process, the specific vulnerability properties of the network nodes and the attacker's exploitation process can be ignored, focusing only on the path the attacker takes to reach the target node. This facilitates honeychain deployment. In complex security defense environments, this approach can significantly reduce computational complexity and improve solution efficiency while maintaining the accuracy of the solution.

[0103] The network environment can be represented as a graph G = (V, E), where V is a set of nodes and E is a set of edges, which represent the connection relationship between nodes. The adjacency matrix E of the network topology is used to represent the connectivity of the network. If there is a connection between node i and node j, then E ij =1, otherwise E ij = 0. Assume that an attacker starts from a node N located in the outer layer of the topology and enters the network through a vulnerability. The path R from node N to the target node T is defined as the potential attack path R = {P|N→T}.

[0104] Based on the simplified attack model based on lateral movement path described above, we will now consider the specification of attack and defense state and action. The attacker's state can be simplified as Sa = (va, Ia), where v a is the node where the attacker is currently located, I a is the attacker’s understanding of the network topology, including the set of accessible nodes and the proximity to the target node. The attacker’s actions are expressed as A a ={v'|v'∈V∧v'≠v a}, that is, based on the current state and the attack target - that is, the node storing sensitive data - the next attack node is selected, specifically, the attacker decides to move to a new node. The state of the defender can be expressed as S d =(v a ,P), indicating that the defender knows the attacker's attack activities, P is the location information of the deployed honey bait. The defender's action is represented by A d= {v'|v'∈V\P}, that is, the location where the honey bait is deployed in the network. Through this simplification, the attacker's strategy can be expressed as a function π d (S d ):S d →A d , which selects nodes to deploy honey baits based on the current state.

[0105] This results in a honey-bait deployment plan corresponding to a single penetration phase. This means analyzing the attack state under different network strategies yields different honey-bait deployment actions, thereby constructing different honey-bait deployments. For multi-stage attack strategies, attack and defense experiments require further integration with reinforcement learning algorithms for intelligent honeychain scheduling.

[0106] S103. Optimize and solve the security game model using a reinforcement learning algorithm to obtain the optimal number and location of honey bait deployment in the spatial dimension, deploy the honey bait into the network based on the optimal number and location, and dynamically adjust the status and location of the honey bait in the network according to a hopping strategy. The hopping strategy is used to adjust the honey bait in the time dimension for network defense.

[0107] Specifically, the security game model is optimized and solved using reinforcement learning algorithms, which can be designed to include:

[0108] Randomly select a network state and input it into the Actor network. The Actor network selects an action based on the input network state and executes the action in the network environment, so that the network environment outputs the next network state and reward. The selected network state, selected action, and output next network state and reward are used as training data and put into the experience replay pool until the number of training data in the experience replay pool reaches the threshold.

[0109] The critic network is trained using the training data in the experience replay pool, the Q-value function is used to carefully evaluate the current state-action strategy, and the critic network parameters are updated by minimizing the mean square error loss.

[0110] Until the model converges to Nash equilibrium, and based on the Nash equilibrium state, the optimal number and position of honey bait deployment in the spatial dimension are obtained.

[0111] Assuming that there are n nodes in the network topology, the order of the adjacency matrix E is n. Among them, a sequence S = [S1, S2, ..., S n ] represents the state, where s i =1 means that the honey bait is placed at the i-th position. A sequence of length n A=[a1,a2,...,a n ] represents the state, where a i= 1 indicates that the honey bait is placed at the i-th position. The action space is the set consisting of the sequence A = [0, 0, ..., 0] to [1, 1, 1, ..., 1]. There are m attack paths in the network topology. In each attack round, the attacker randomly selects a path based on different attack strategies. When the honey bait is placed on the attack path, the attacker cannot avoid the deception resource, and the agent receives a reward of 10. Otherwise, the attacker avoids the deception, and the agent receives a very low reward of -100. When the action is executed, the agent receives a reward of -1, which represents the cost of deploying the honey bait.

[0112] In order to optimize the defender's strategy, in this embodiment, the DDPG algorithm based on the Actor-Critic framework can be used to perform reinforcement learning solutions for honeylink deployment. Figure 5 As shown, the Critic target network is used to approximate the Q value function Q of the state-action pair at the next moment ω' (S t+1 ,π θ' (S t+1 )). The Actor target network is used to approximate the next action value. The target Q value function in the current state can be expressed as:

[0113] y i =r i +γQ ω' (S t+1 ,π θ' (S t+1 ))

[0114] The critic training network outputs the Q-value function Q of the current state-action ω (S t ,a t ) is used to evaluate the current strategy. DDPG also adds a Gaussian noise function to the behavior strategy to expand the agent's exploration of the environment. The critic network can update the parameters of the critic network by minimizing the mean squared error loss. Its loss function is:

[0115]

[0116] Among them, is the target Q value, and Q ω (S i ,a i ) is the Q value estimate of the current state-action pair by the Critic network, a i =π θ (S i )+ε, where ε represents the exploration noise on the behavior strategy.

[0117] The Actor target network is used to provide the strategy for the next state, while the Actor training network provides the strategy for the current state. Combined with the Q-value function of the Critic training network, the policy gradient of the Actor during parameter update can be obtained:

[0118]

[0119] For the update of the target network parameters ω' and θ', DDPG adopts a soft update mechanism, that is, only some parameters are updated during each learning to ensure slow parameter updates, thereby enhancing the stability of learning:

[0120] ω'←ξω+(1-ξ)ω'

[0121] θ'←ξθ+(1-ξ)θ'

[0122] The DDPG algorithm combines the advantages of value-based and policy-based methods, enabling deep reinforcement learning to effectively handle problems in continuous action spaces while maintaining a certain level of exploration capability. For honeychain deployments, it can dynamically adjust strategies to respond to attackers' actions in a constantly changing network environment, effectively confusing and deterring attackers and protecting servers on the network.

[0123] Among them, IP hopping can be adopted as a hopping strategy. Based on the hopping strategy, the host addresses in the network subnet are randomly remapped and the honey bait network configuration in the network is dynamically changed. The hopping strategy includes: setting a time interval hyperparameter for triggering IP hopping. When the time interval is reached, the target host is triggered to randomly generate a new IP address and check whether the new IP address has been used by other hosts. If it has been used, a new IP address is regenerated. If it has not been used, the randomly generated new IP address is assigned to the target host and the host address mapping relationship in the network is updated.

[0124] Furthermore, based on the above method, an embodiment of the present invention also provides a multi-step attack trapping system with dynamic deception defense capability, comprising: a honey bait configuration module, a model setting module and an attack defense module, wherein:

[0125] A honey bait configuration module is used to generate different types of honey baits based on the resources required by microservices and the correlation between microservices, so as to use the honey baits to perceive network attack behaviors. The honey baits include operating system honey baits, process honey baits, and service honey baits, and the honey baits are correlated with each other in terms of network resource configuration;

[0126] A model setting module is configured to construct a security game model for describing a honey chain attack and defense confrontation process according to network scale and defense requirements, the security game model is represented by using attack and defense participants, a current network environment state, an attack and defense action set, a discount factor, a network state transition probability and an attack and defense payoff function, wherein the current network environment state includes a current network honey quantity, a defender deployed honey node position, an attacker penetration node position and a vulnerability position;

[0127] An attack and defense module is configured to obtain an optimal number and position of the honey deployment in a spatial dimension by using a reinforcement learning algorithm to optimize and solve the security game model, deploy the honey to the network according to the optimal number and position, and dynamically adjust the honey state and position in the network according to a hop strategy, the hop strategy is used to adjust the honey in a time dimension to perform network defense.

[0128] Further, the application also provides a multi-level defense system, comprising a defense platform and a protected server, wherein,

[0129] The defense platform is configured to generate the honeys by using the above method, and configure the honeys as deception resources to disguise as nodes in a target network, wherein the honeys are connected in series into a space-time honey chain and distributed at different positions in an attacker intrusion depth direction, and the network address of the honeys is set as a random hop mode to present a dynamic false network view to the attacker.

[0130] The protected server comprises a legal user in a network service and / or a network server.

[0131] According to Figure 4 According to the principle model shown, the space-time honey chain can continuously switch the target network topology over time, and traps are set in different subnet spaces of the target network, and the bypass arrangement can realize the perception of multi-step attacks in a user transparent and inapparent manner, and then trace the threat based on attack traces, and realize attack trapping. This strategy is like a “net of heaven and earth” arranged in the network environment, so that the attacker has no way to go, thereby achieving the purpose of protecting the network security.

[0132] Based on the dynamic deception resource deployment in the space-time dimension, the multi-level defense system can be divided into three parts: an attacker, a defense system and a protected server. First, the attacker represents individuals or organizations that try to carry out illegal activities through the network. They may use various techniques and methods, such as social engineering, vulnerability exploitation and malicious software, to break through the defense layers of the network and select target hosts for attack. Second, the defense system is the core part of the whole system. For example, Figure 6As shown in the figure, it includes spatiotemporal honeylinks as dynamic deception defense resources. Honeylink hosts are configured as deception resources, disguised as nodes in the network, and connected in chains at different locations along the attacker's intrusion depth. This attracts the attacker's attention, captures the attacker's behavioral characteristics and attack methods, and provides important intelligence support for subsequent defense strategies. By randomly changing the network address of the honeylink host, a dynamic, false network view is presented to the attacker. Finally, there are protected servers, representing legitimate users of network services and servers within the network. Through this system's protection framework, spatiotemporal honeylinks are deployed in a bypass system for normal communication, allowing legitimate users to securely access network resources and services without interference from attackers.

[0133] For the system's defense direction, such as Figure 7 As shown, it can be divided into four parts: network configuration, system boundary, functional application, and result testing. These four parts can coordinate and control the implementation of defense actions.

[0134] Network configuration focuses on implementing basic network functionality. This includes assigning IP addresses to devices and laying out the network topology to define network connectivity. Furthermore, managing host information involves adding, removing, and updating devices on the network to maintain accurate and up-to-date network configurations.

[0135] The system boundary is the network's front line of defense, encompassing firewalls and intrusion detection mechanisms. As the first line of defense, firewalls block known malicious traffic and unauthorized access attempts, protecting the internal network from external threats. Intrusion detection mechanisms monitor network activity, identifying and reporting potential attacks, providing early warning for subsequent defensive measures. Together, these two components form the system boundary, ensuring network security and integrity.

[0136] The functional application part includes the deployment of dynamic deception defense resources in honeychain time and space, as well as the application of dynamic deception defense strategy behavior to enhance network security and efficiency. These functional applications play a key role in the network defense system and improve the overall security performance of the network.

[0137] Results testing is a crucial step in evaluating and optimizing network defense strategies. This evaluation verifies the effectiveness of existing security measures, identifies potential weaknesses, and addresses them. Optimization models are constructed based on historical data and current trends to create an optimal security model. Finally, results testing is conducted in a real-world environment to verify the effectiveness of the optimized model and ensure that the network defense system can adapt to the ever-changing threat landscape.

[0138] Moving Target Defense (MTD) is a security measure that increases the difficulty for attackers to identify and exploit network vulnerabilities by dynamically changing certain elements in the network (such as IP addresses and service locations). By negotiating a hopping mechanism to pseudo-randomly change IP addresses, not only does it increase the concealment of communications, it also greatly improves network security. Therefore, in this embodiment, IP hopping in the time dimension can be used as a honeylink time dynamic link resource.

[0139] Initialize host address: The default_mapping method is a key step in the network environment initialization process. It is responsible for establishing the initial address mapping relationship for each subnet. In this process, the Internet subnet (i.e., subnet 0) is specifically ignored because it has special properties and uses and usually does not require regular address mapping. For subnets other than the Internet subnet, the default_mapping method will traverse each subnet and create an address mapping entry for each host in it. These mapping entries not only record the location information of the host in the current subnet, but also contain other necessary network attributes, such as IP address, MAC address, etc. These mapping information will be frequently referenced and updated in subsequent MTD operations to ensure the dynamic and security of the network environment. The algorithm pseudocode can be shown as Algorithm 2.

[0140]

[0141] Honeylink Time Hopping: The moving_target method is the core logic for implementing Moving Target Defense (MTD). The key to this method is to increase the difficulty of attack by dynamically changing the network configuration. Specifically, it randomly remaps host addresses within a subnet, thereby continuously changing the network's attack surface, making it difficult for attackers to identify specific targets.

[0142] When calling the moving_target method, a new state object, new_state, is passed as a parameter. This state object typically contains complete information about the current network, such as subnet structure and host addresses. During execution, the method traverses each subnet and randomly remaps the host addresses within it. However, it is worth noting that this process intentionally ignores Internet subnets and subnets containing only Penboxes. This is because these subnets are special and may involve communication with external networks or specific security policies, making random address remapping unsuitable. For other subnets containing multiple hosts, the method first randomly scrambles the addresses of these hosts. This scrambling is completely random and is intended to disrupt any potential attack patterns or predictive models. After the scrambling is complete, the method updates the subnet-to-host address mapping to ensure the network state is up-to-date. The pseudocode for this algorithm can be seen in Algorithm 3.

[0143]

[0144] Finally, the moving_target method returns an updated state object. This state object now contains the latest network configuration after random remapping and can be used for subsequent network operations or security analysis. In this way, the moving_target method effectively implements dynamic network defense.

[0145] The temporal deployment strategy for mobile target resources in the honeychain focuses on applying appropriate transition behaviors at the right time to execute defensive measures, improving system security and attack resistance. Dynamic detection of honey baits is a temporal defense deployment resource that can be considered in the next step.

[0146] Within the framework of mobile target defense, a key factor in ensuring system security is the careful selection of hopping timing. In this embodiment, a time-based hopping strategy can drive the automatic change of IP addresses in the network through a hopping period manually set by the experimenter. The implementation of this strategy not only simplifies the operation process, but also, by regularly evaluating the effectiveness of the hopping strategy, the timing parameters can be adjusted in a timely manner to adapt to changing security threats. By having an automated system change IP addresses within a preset period, the network can be provided with a dynamically changing environment, making it difficult for potential attackers to identify targets, thereby improving overall network security.

[0147] To implement this strategy, the first step is to determine which IP addresses in the network will serve as the hop parameters. Then, the experimenter manually sets an appropriate hop period. Once the period is determined, the system automatically executes the IP address hop. At the end of each period, the system selects a new IP address for hop based on pre-set rules or algorithms, thus keeping the network configuration continuously updated and changing. Furthermore, the system records detailed information about each hop, including the time, old IP address, and new IP address. These logs are crucial for subsequent security analysis and policy adjustments.

[0148] To ensure the effectiveness of the hopping strategy, regular evaluation and optimization are essential. By analyzing system logs, experimenters can assess the actual impact of the hopping strategy on network defenses, identify any potential security vulnerabilities, and adjust the hopping period and hopping rules based on the evaluation results.

[0149] Building on existing functionality, dynamic deployment of detection baits will further enhance the strategic and effective nature of network deception defenses. Dynamic deployment means adjusting the status and location of detection baits based on real-time network conditions and attack behavior, thereby selecting the optimal deception opportunity.

[0150] In this embodiment, IP hopping is used to implement address hopping of the detection bait, thereby making the honeylink space and time dynamic. The ingenuity of this strategy lies in transforming the originally static detection bait into dynamically changing network nodes, making it difficult for attackers to track and locate these baits. When the detection baits are strategically deployed in their number in the attacker's attack direction, forming a deep and complete spatial chain, IP hopping technology gives these baits dynamic temporal transformation. This means that the IP address of the detection bait will change periodically or randomly, making the attacker face a constantly moving and changing target. This flexibility is like waging a "guerrilla war" against a multi-step attacker. By constantly changing the location and identity of the detection bait, the attacker greatly increases uncertainty about the host type. Attackers cannot determine whether they are interacting with the real target or a disguised honey bait. This uncertainty not only increases the difficulty of the attack but also increases the adaptability of network defenders' strategies to counter the attacker's behavior. Furthermore, IP hopping technology enhances the concealment of the detection bait. Because the honey bait's address constantly changes, attackers find it difficult to detect its presence, thereby increasing the success rate of deception. This strategy makes it more difficult for attackers to operate in the network, forcing them to spend more time and resources trying to understand and adapt to this dynamically changing network environment.

[0151] As a result, a periodic "chain" is formed in the time dimension, and a "chain-like" IP jump is performed in response to each step of the attacker's vulnerability scanning, vulnerability exploitation and covert infiltration attack behavior, thereby building a honeylink time dynamic chain and implementing effective interference (Disrupt) and guidance (Channel) tactical defense. It presents the attacker with a false network view that jumps over time, consumes the attacker's gradual intrusion time, and provides support for evaluating the defense effect of mobile targets.

[0152] In simulation testing, the following method can be used to define an address change strategy to deploy mobile targets. First, a time interval hyperparameter, adjustable according to actual needs, is set to trigger IP hopping. When the specified time interval is reached, the IP hopping behavior is executed. For the host to be hopped, a new IP address is randomly generated. To ensure the uniqueness of the new address, it is necessary to check whether the newly generated IP address is already in use by other hosts. If so, a new IP address is generated. The newly generated IP address is then assigned to the host to be hopped, and the host-address mapping in the network is updated. After the IP hop, the attacker does not know which hosts have been exploited. Because address hopping changes the host's location in the network, the attacker's knowledge base needs to be updated to reflect the new network state. This may include updating the attacker's known information such as the network topology and host addresses. To increase the number of addresses available for mutating, "null" hosts are introduced. Null hosts cannot be detected or attacked. Their sole purpose is to introduce additional addresses into the subnet. Therefore, they serve as unused addresses on the network to which existing hosts can move. This increases the address space and makes the simulation more realistic. After the address hop occurs, its impact on the attacker is evaluated. Since the attacker no longer knows which hosts have been successfully exploited, the attacker may need to re-probe the network to identify the target hosts.

[0153] By scheduling honeylinks using a secure game model and reinforcement learning algorithm, we can accurately calculate the optimal number of honey baits to deploy on the target network, reducing deployment costs. By designing deceptive honey baits, we can enhance the simulation, relevance, and diversity of the honey baits, thereby increasing their attractiveness to attackers. By utilizing spatiotemporal honeylink defense resource deployment strategies, we combine network deception and mobile target defense strategies, increasing the difficulty for attackers to identify attacks and further improving the security protection of cloud-native environments.

[0154] Unless otherwise specifically stated, the relative steps, numerical expressions and values ​​of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0155] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0156] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.

[0157] Those skilled in the art will appreciate that all or part of the steps in the above method can be performed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disk. Alternatively, all or part of the steps in the above embodiment can be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or software functional modules. The present invention is not limited to any specific combination of hardware and software.

[0158] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A multi-step attack trapping method with dynamic deception defense enablement, characterized in that: Include: Different types of honey baits are generated based on the resources required by microservices and the correlation between microservices. The honey baits are used to perceive network attack behaviors. The honey baits include operating system honey baits, process honey baits, and service honey baits. The honey baits are correlated with each other in terms of network resource configuration. A security game model is constructed based on the network scale and defense requirements to describe the honeylink attack and defense process. The security game model is represented by the attack and defense participants, the current state of the network environment, the attack and defense action set, the discount factor, the network state transition probability, and the attack and defense payoff function. The current state of the network environment includes the number of honey baits in the current network, the location of the defender's honey bait nodes, the location of the attacker's penetration nodes, and the location of the vulnerability. A reinforcement learning algorithm is used to optimize and solve the security game model to obtain the optimal number and location of honey bait deployment in the spatial dimension. The honey bait is deployed in the network based on the optimal number and location, and the status and location of the honey bait in the network are dynamically adjusted according to the hopping strategy. The hopping strategy is used to adjust the honey bait in the time dimension for network defense.

2. The multi-step attack trapping method with dynamic deception defense enablement according to claim 1 is characterized in that: Different types of honey bait are generated based on the resources required by microservices and the correlation between microservices, including: Creating a honey bait network resource based on microservice resource elements, wherein the microservice resource elements include an operating system, a service, and a process; The network resource type and quantity in the honey bait are configured according to the microservice correlation parameters, and the honey bait configuration elements are integrated into the host to generate the corresponding honey bait. The correlation parameters include a probability parameter for indicating that the microservice selects a new configuration, an association parameter for indicating whether there is an association between microservice network resources, and a Poisson distribution parameter for indicating the configuration quantity system. The host is a virtual or physical network environment entity consistent with an ordinary network host.

3. The multi-step attack trapping method with dynamic deception defense enablement according to claim 1 is characterized in that: A security game model is constructed based on the network scale and defense requirements to describe the attack and defense process of the honeychain, including: The network environment is represented as a graph data consisting of a set of nodes and a set of edges, and the network topology adjacency matrix is ​​used to represent the connectivity of the network. The path that the attacker takes to move laterally in the network and reach the target node is considered the attack path. The attacker's state is represented by the attacker's current node, the set of nodes the attacker can access, and the degree of proximity to the target node. The attacker's actions are represented by the current state of the network node and the attack target. The defender's state is obtained based on the attacker's known attack actions and the honey bait information deployed on the network. The defender's actions are represented by the location of the honey bait deployed in the network. The defender's state and actions are used to represent the attacker's attack strategy. The defender's current network state is represented by the attacker's attack strategy to select the defense strategy for deploying the honey bait node. Construct a profit function for both attackers and defenders based on the rewards they receive from taking corresponding offensive and defensive actions under the current network state. Based on the attack path, attack and defense status, attack and defense actions, attack and defense strategies and attack and defense benefit functions, a security game model is established to describe the honeychain attack and defense confrontation process.

4. The multi-step attack trapping method with dynamic deception defense enablement according to claim 1 or 3, characterized in that: A security game model is constructed based on the network scale and defense requirements to describe the attack and defense process of the honeychain, which also includes: Divide the roles of participants in the honeylink attack and defense process into leaders and followers; According to the game order of the attack and defense confrontation process, the leader makes the decision first, and the followers respond accordingly based on the leader's decision. When the leader makes a decision, he maximizes his own interests by predicting the followers' action responses. The followers choose the best strategy to perform the corresponding actions based on the leader's decision.

5. The multi-step attack trapping method with dynamic deception defense enablement according to claim 1 is characterized in that: Utilize reinforcement learning algorithms to optimize and solve security game models, including: Randomly select a network state and input it into the Actor network. The Actor network selects an action based on the input network state and executes the action in the network environment, so that the network environment outputs the next network state and reward. The selected network state, selected action, and output next network state and reward are used as training data and put into the experience replay pool until the number of training data in the experience replay pool reaches the threshold. The critic network is trained using the training data in the experience replay pool, the Q-value function is used to carefully evaluate the current state-action strategy, and the critic network parameters are updated by minimizing the mean square error loss. Until the model converges to Nash equilibrium, and based on the Nash equilibrium state, the optimal number and position of honey bait deployment in the spatial dimension are obtained.

6. The multi-step attack trapping method with dynamic deception defense enablement according to claim 1 is characterized in that: Dynamically adjust the status and location of honey baits in the network based on the hopping strategy, including: IP hopping is adopted as the hopping strategy. Based on the hopping strategy, the host addresses in the network subnet are randomly remapped and the honey bait network configuration in the network is dynamically changed.

7. The multi-step attack trapping method with dynamic deception defense enablement according to claim 1 or 6, characterized in that: The hopping strategy includes: setting a time interval hyperparameter for triggering IP hopping. When the time interval is reached, the target host is triggered to randomly generate a new IP address and check whether the new IP address has been used by other hosts. If it has been used, a new IP address is regenerated. If it has not been used, the randomly generated new IP address is assigned to the target host and the host address mapping relationship in the network is updated.

8. A multi-step attack trapping system with dynamic deception defense capability, characterized in that: Contains: honey bait configuration module, model setting module and attack defense module, among which, A honey bait configuration module is used to generate different types of honey baits based on the resources required by microservices and the correlation between microservices, so as to use the honey baits to perceive network attack behaviors. The honey baits include operating system honey baits, process honey baits, and service honey baits, and the honey baits are correlated with each other in terms of network resource configuration; The model setting module is used to construct a security game model for describing the honeychain attack and defense confrontation process based on the network scale and defense requirements. The security game model is represented by the attack and defense participants, the current state of the network environment, the attack and defense action set, the discount factor, the network state transition probability, and the attack and defense benefit function. The current state of the network environment includes the number of honey baits in the current network, the location of the defender's honey bait nodes, the location of the attacker's penetration nodes, and the location of the vulnerability. The attack and defense module is used to optimize and solve the security game model using a reinforcement learning algorithm, obtain the optimal number and location of honey bait deployment in the spatial dimension, deploy the honey bait into the network based on the optimal number and location, and dynamically adjust the status and location of the honey bait in the network according to the hopping strategy. The hopping strategy is used to adjust the honey bait in the time dimension for network defense.

9. A multi-layered defense system comprising: Defense platform and protected servers, where A defense platform for generating honey baits using the method of claim 1, and configuring the honey bait hosts as deceptive resources to disguise themselves as nodes in the target network, wherein the honey bait hosts are connected in series to form a spatiotemporal honey chain and are distributed at different locations in the depth direction of the attacker's invasion, and the network addresses of the honey bait hosts are set to a random hopping mode to present a dynamic false network view to the attacker; Protected servers include: legitimate users and / or network servers in network services.

10. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Verifying the suitability of a computer lure for an objective for deployment on a computer system

    EP4488860A1