Honeypot deployment method, device and equipment based on attack and defense income, medium and product

By modeling the network environment and determining the Stackelberg game reward function, a hybrid honeypot deployment strategy was formulated, which solved the problem of lack of flexibility and low security in the existing honeypot deployment methods, and achieved the effect of dynamic adjustment and multi-type honeypot deployment.

CN120034354APending Publication Date: 2025-05-23BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510017835.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing honeypot deployment methods lack flexibility, cannot dynamically adjust defense strategies, and cannot deploy software-defined honeypots and non-software-defined honeypots at the same time, resulting in low overall system security.

Method used

By modeling the network environment, generating a network topology, and modeling attackers and defenders based on the topology, determining the reward function of Stackelberg game, combining historical interaction experience, number of system vulnerabilities, system defense strength and other factors, a hybrid honeypot deployment strategy is formulated, including the deployment of software-defined honeypots and non-software-defined honeypots.

Benefits of technology

Dynamic adjustment of honeypot deployment strategy has been achieved, the flexibility and defense effect of honeypot deployment has been improved, the phenomenon of attackers using the same vulnerabilities to break through other nodes of the system, and the overall security of the system has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034354A_ABST
    Figure CN120034354A_ABST
Patent Text Reader

Abstract

The invention provides a honeypot deployment method and device based on attack and defense income, equipment, a medium and a product, and relates to the technical field of network security, and the method comprises the steps: carrying out the modeling of a network environment, and generating a network topology; based on network topology, modeling is carried out on an attacker and a defender, and an attacker action space and a defender action space are obtained; determining a reward function of the Stackelberg game based on historical interaction experience, the number of system vulnerabilities, the system defense strength, the attacker action space and the defender action space; determining a mixed honeypot deployment strategy based on the reward function; the mixed honeypot deployment strategy comprises a software-defined honeypot deployment strategy and a non-software-defined honeypot deployment strategy. By means of the mode, dynamic adjustment of the honeypot deployment strategy can be achieved, the flexibility of honeypot deployment is improved, the defense effect of the honeypot is optimized, and the overall safety of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to a honeypot deployment method, device, equipment, medium and product based on attack and defense benefits. Background Art

[0002] Cyber ​​deception is an advanced active technology in the field of network security, which aims to provide credible but misleading information to attackers to lead them astray. At present, the application of deception technology has expanded to the field of cyberspace and has become a means of intrusion detection and defense. Taking proactive measures can capture attackers and closely monitor their actions. Honeypots play an important role in this process. Honeypots can be used as simulated entities in systems or networks to deceive attackers. By using honeypots to study the strategies and intentions of attackers, defenders can improve their understanding of attacks and develop more effective deception schemes. In real power grid scenarios, honeypots also have many applications. For example, in distribution network scenarios, honeypot technology can be used to simulate critical infrastructure, trap and analyze the behavior of potential network attackers, thereby enhancing the security and defense capabilities of the system.

[0003] Software-defined honeypot (SD-Honeypot) is a honeypot system based on software-defined networking (SDN). It can be deployed on software-defined wide area networks (SD-WAN) to identify and monitor malicious activities in the network. Software-defined honeypot uses SDN controllers to create virtual network environments and deploy simulated targets to attract attackers. It helps relevant personnel understand network threats by recording attackers' behaviors and technical means, and can provide useful information about potential security vulnerabilities and attack techniques, which helps to enhance network security and provide timely responses to potential threats.

[0004] Most of the current honeypot deployment methods use game theory or reinforcement learning models to develop an active deceptive honeypot allocation strategy. However, the existing honeypot deployment methods can often only decide a static single honeypot deployment strategy, and cannot develop corresponding defense strategies according to the different actions of attackers in the system. In addition, the existing honeypot deployment methods cannot deploy traditional honeypots (i.e. non-software-defined honeypots) and software-defined honeypots in the same system. Deploying only a single type of honeypot may cause attackers to reuse the same vulnerability on each node, which may lead to the system being compromised and cause serious losses.

[0005] Therefore, the existing honeypot deployment methods lack flexibility and the honeypot's defense effect is poor, resulting in low overall security of the system. Summary of the invention

[0006] The present invention provides a honeypot deployment method, device, equipment, medium and product based on attack and defense benefits, which are used to solve the defects of the honeypot deployment method in the prior art that the honeypot lacks flexibility, the defense effect of the honeypot is not good, and the overall security of the system is not high.

[0007] The present invention provides a honeypot deployment method based on attack and defense benefits, comprising: modeling a network environment to generate a network topology; based on the network topology, modeling an attacker and a defender respectively to obtain an attacker action space and a defender action space; determining a reward function of a Stackelberg game based on historical interaction experience, the number of system vulnerabilities, the system defense strength, the attacker action space and the defender action space; determining a hybrid honeypot deployment strategy based on the reward function; wherein the hybrid honeypot deployment strategy is a deployment strategy including software-defined honeypots and non-software-defined honeypots.

[0008] According to a honeypot deployment method based on attack and defense benefits provided by the present invention, a reward function of the Stackelberg game is determined based on historical interaction experience, the number of system vulnerabilities, the system defense strength, the attacker's action space, and the defender's action space, including: determining the probability of a node in the network topology being hacked based on a logistic regression model according to the historical interaction experience, the number of system vulnerabilities, and the system defense strength; determining the system damage after the node is hacked based on the probability of the node being hacked; and determining the reward function of the Stackelberg game based on the system damage after the node is hacked, the attacker's action space, and the defender's action space.

[0009] According to a honeypot deployment method based on attack and defense benefits provided by the present invention, the calculation formula for the probability of a node being hacked is: ; in, is the probability of a node being hacked; is a natural constant; are the parameters of the logistic regression model; for historical interactive experience; is the number of system vulnerabilities; The system defense strength.

[0010] According to a honeypot deployment method based on attack and defense benefits provided by the present invention, the calculation formula for the system damage after the node is breached is: ; in, System damage after the node is compromised; is the probability of a node being hacked; For Node The weighted value of is the preset coefficient.

[0011] According to a honeypot deployment method based on attack and defense benefits provided by the present invention, the calculation formula of the reward function is: ; in, is the reward function; The cost of deploying a honeypot for the defender; The cost of carrying out an attack for the attacker; Capture rewards for defenders; Reward attackers for success; For Node The weighted value of is the preset coefficient; System damage after a node is compromised; is the defensive actions contained in the defender action space, Indicates that the defensive action is the defender on the side < >Deploy honeypots, It means that the defender does not deploy honeypots; is the attack action contained in the attacker's action space, Indicates that the attack action is the attacker's decision to attack the node , Indicates that the attack action is the attacker's decision to attack the node , It means that the attacker does not commit an attack; A set of nodes representing the network topology; For Node The weighted value of .

[0012] According to a honeypot deployment method based on attack and defense benefits provided by the present invention, a hybrid honeypot deployment strategy is determined based on a reward function, including: determining a defender linear equation based on the reward function; solving the defender linear equation to obtain a target defense strategy of the Stackelberg game; determining an attacker linear equation based on the reward function; solving the attacker linear equation to obtain a target attack strategy of the Stackelberg game; and determining a hybrid honeypot deployment strategy based on the target defense strategy and the target attack strategy.

[0013] The present invention also provides a honeypot deployment device based on attack and defense benefits, including: a network topology modeling submodule, used to model the network environment and generate a network topology; a participant modeling submodule, used to model the attacker and the defender respectively based on the network topology to obtain the attacker's action space and the defender's action space; a reward function determination submodule, used to determine the reward function of the Stackelberg game based on historical interaction experience, the number of system vulnerabilities, the system defense strength, the attacker's action space and the defender's action space; a game submodule, used to determine a hybrid honeypot deployment strategy based on the reward function; wherein the hybrid honeypot deployment strategy is a deployment strategy including software-defined honeypots and non-software-defined honeypots.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, any of the above-mentioned honeypot deployment methods based on attack and defense benefits is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the computer program implements any of the above-mentioned honeypot deployment methods based on attack and defense benefits.

[0016] The present invention also provides a computer program product, including a computer program, which implements any of the above-mentioned honeypot deployment methods based on attack and defense benefits when executed by a processor.

[0017] The honeypot deployment method, device, equipment, medium and product based on attack and defense benefits provided by the present invention comprehensively consider various factors such as historical interaction experience, number of system vulnerabilities, system defense strength, attacker actions and defender actions during the honeypot deployment process, and derive a hybrid honeypot deployment strategy through game theory, which can realize dynamic adjustment of the honeypot deployment strategy, improve the flexibility of honeypot deployment, and optimize the defense effect of the honeypot; at the same time, the hybrid honeypot deployment strategy requires the deployment of software-defined honeypots and non-software-defined honeypots at the same time. The deployment of multiple types of honeypots can avoid the phenomenon that attackers break into a single type of honeypot and then use the same vulnerability to continue to break into other nodes of the system, thereby improving the overall security of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 This is one of the flow charts of the honeypot deployment method based on attack and defense benefits provided by the present invention.

[0020] Figure 2 This is the second flow chart of the honeypot deployment method based on attack and defense benefits provided by the present invention.

[0021] Figure 3 This is one of the system benefit comparison diagrams between the honeypot deployment method based on attack and defense benefits provided by the present invention and the traditional method.

[0022] Figure 4 This is the second system benefit comparison chart between the honeypot deployment method based on attack and defense benefits provided by the present invention and the traditional method.

[0023] Figure 5 This is the third system benefit comparison chart between the honeypot deployment method based on attack and defense benefits provided by the present invention and the traditional method.

[0024] Figure 6 This is the fourth system benefit comparison chart between the honeypot deployment method based on attack and defense benefits provided by the present invention and the traditional method.

[0025] Figure 7 It is a structural schematic diagram of a honeypot deployment device based on attack and defense benefits provided by the present invention.

[0026] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] See also Figure 1 , Figure 1 This is one of the flow charts of the honeypot deployment method based on attack and defense benefits provided by the present invention. In this embodiment, the honeypot deployment method based on attack and defense benefits is applied to a network system, and the honeypot deployment method based on attack and defense benefits includes steps S110 to S140, and each step is as follows: S110: Model the network environment and generate a network topology.

[0029] See also Figure 2 , Figure 2 This is the second flow chart of the honeypot deployment method based on attack and defense benefits provided by the present invention.

[0030] This embodiment proposes a software-defined honeypot deployment method for attack and defense benefits based on Stackelberg game, which mainly includes two parts: one is network topology modeling and participant modeling of Stackelberg game, and the other is the Stackelberg game honeypot deployment method for attack and defense benefits. In the honeypot deployment process, the probability of an attacker successfully bypassing the honeypot is calculated through a machine learning algorithm and the corresponding loss function is calculated. The probability of an attacker successfully breaking through is added to the calculation process of the cost and benefit of honeypot deployment. Finally, the costs and benefits of both the attacker and the defender when deploying different hybrid honeypot strategies are quantified, and an optimal hybrid honeypot deployment strategy is obtained through game.

[0031] like Figure 2 As shown, the network environment is first modeled using Petri nets to generate network topology.

[0032] Specifically, in order to solve the problem that it is difficult to topologically represent the real network environment, this embodiment uses Petri nets to model the real network environment. Petri nets can represent system behaviors through states and transitions, and combine the nodes and edges of the network topology graph to model attack paths and dependencies, capture attack steps and system state changes, and are a reliable formal analysis and verification tool that can improve the accuracy of network security assessments and achieve visualization.

[0033] S120: Based on the network topology, the attacker and the defender are modeled respectively to obtain an attacker action space and a defender action space.

[0034] First, based on the network topology, the defender is modeled to obtain the defender action space .

[0035] Specifically, the network topology includes multiple nodes and multiple edges, and the defender's goal is to protect the target node in the system by selecting the deployment location of the honeypot. represents the total number of honeypots allowed to be allocated (i.e., budget). In order to maintain the structure of the network topology, it can be assumed that honeypots will be placed as interfaces of each node. To compromise any node, an attacker must first compromise its direct parent node. The defender needs to consider the weighted value of the node, the total budget, and the path that the attacker may follow before making a decision.

[0036] Therefore, the defender action space of each honeypot contains all edge sets. Assume is the edge set of the network topology, that is, the set of edges of the network topology, assuming is the defender’s action space. Considering B honeypots, we have ,in, is the inner product of two vectors, Is the length A binary vector with entries set to 1 if a honeypot is assigned to an edge and 0 otherwise. The inner product condition between the unit vector e and all vectors with entries set to 1 ensures that the total number of assigned honeypots does not exceed the defender's budget. Therefore, a feasible deception can be represented as a binary vector with length A binary vector such that To avoid trivial deception scenarios, this embodiment assumes a limited budget that does not allow covering the entire network system through honeypots, so strategic allocation is of great significance to capture attackers and protect critical and high-value nodes.

[0037] The defender's profit function needs to consider the number of successful honeypots and wasted honeypots: if the attacker passes through the edge where the honeypot is placed, it means that the defender's honeypot allocation is successful, otherwise it is considered a honeypot loss because the attacker is able to avoid the honeypot. Each specific node that is successfully protected is multiplied by the node weight to obtain the defense profit of the node. Similarly, the defense loss is compensated by the node that is successfully destroyed multiplied by the node weight.

[0038] Furthermore, based on the network topology, the attacker is modeled to obtain the attacker's action space .

[0039] Specifically, the attacker's goal is to reach the target node that the defender cannot accurately specify. The attacker can try to choose a path from the network entry node to the target node and avoid paths that may have honeypots. Assume that P is the set of all feasible paths between each pair of nodes in the network topology. The network topology may allow multiple paths between the same two nodes in P. From the entry node where the attacker initially compromises the network, Initially, the attacker's action space is a subset of P, where is the starting node. The attacker's attack action is a vector that contains all nodes on the attack path, starting from Start until a target node.

[0040] The attacker needs to choose a complete path between the entry node and one of the target nodes (i.e., leaf nodes), and the attacker obtains attack benefits for each successful move (i.e., attack success), which enables it to avoid the honeypot and compromise a new node.

[0041] S130: Determine the reward function of the Stackelberg game based on historical interaction experience, the number of system vulnerabilities, the system defense strength, the attacker action space, and the defender action space.

[0042] After completing the modeling of the network topology, attackers, and defenders, the probability of a node being compromised in the network topology can be calculated using a logistic regression model based on historical interaction experience, the number of system vulnerabilities, system defense strength, attacker action space, and defender action space. This probability can then be added to the reward function of the Stackelberg game to derive the final hybrid honeypot deployment strategy.

[0043] It should be noted that the existing process of quantifying system benefits and costs has never taken into account factors such as the experience gained by the defender in the process of interacting with the attacker (i.e., historical interaction experience), system vulnerabilities, and defense strength. In order to add the above factors to the influencing variables of the game, this embodiment uses a logistic regression algorithm to calculate the probability of a node in the system being hacked, and adds the above factors to the quantification process in the form of logistic regression function variables, simulating real network attack and defense scenarios, so as to achieve the purpose of determining the final reward based on multiple factors in the network attack and defense scenarios.

[0044] S140: Determine a hybrid honeypot deployment strategy based on the reward function.

[0045] Among them, the hybrid honeypot deployment strategy is a deployment strategy that includes software-defined honeypots and non-software-defined honeypots.

[0046] The honeypot deployment method based on attack and defense benefits provided in this embodiment comprehensively considers various factors such as historical interaction experience, the number of system vulnerabilities, the system defense strength, the attacker's actions and the defender's actions during the honeypot deployment process, and derives a hybrid honeypot deployment strategy through game theory, which can realize dynamic adjustment of the honeypot deployment strategy, improve the flexibility of honeypot deployment, and optimize the defense effect of the honeypot; at the same time, the hybrid honeypot deployment strategy requires the deployment of software-defined honeypots and non-software-defined honeypots at the same time. The deployment of multiple types of honeypots can avoid the phenomenon that attackers break into a single type of honeypot and then use the same vulnerability to continue to break into other nodes of the system, thereby improving the overall security of the system.

[0047] In some embodiments, a reward function of the Stackelberg game is determined based on historical interaction experience, the number of system vulnerabilities, the system defense strength, the attacker's action space, and the defender's action space, including: determining the probability of a node in the network topology being hacked based on a logistic regression model according to the historical interaction experience, the number of system vulnerabilities, and the system defense strength; determining the system damage after the node is hacked based on the probability of the node being hacked; determining the reward function of the Stackelberg game based on the system damage after the node is hacked, the attacker's action space, and the defender's action space.

[0048] Specifically, after modeling the attacker action space and the defender action space, the probability of a node in the network topology being breached can be determined using a logistic regression model based on historical interaction experience, the number of system vulnerabilities, and the strength of system defense.

[0049] Among them, the logistic regression model is a model widely used in binary classification problems. It can estimate the relationship between input features and output labels and output a probability value in the range of [0, 1].

[0050] In this embodiment, the mathematical form of the logistic regression model, that is, the probability of a node being hacked in the network topology, is: ; in, is the probability of a node being hacked, that is, the probability of the attacker's attack being successful, and the probability of the defender's defense failing; is a natural constant; are the parameters of the logistic regression model; for historical interactive experience; is the number of system vulnerabilities; The system defense strength.

[0051] , , , These are the parameters of the logistic regression model and need to be learned through training data.

[0052] is the historical interaction experience, which represents the experience gained by the defender in interacting with the attacker at a certain node, that is, the number of times the attacker attacks the node. The mathematical form of is: ; in, Indicated in At this moment, the number of attacks by the attacker on the node.

[0053] The number of system vulnerabilities can be detected by using Petri nets to model the network environment. If a member does not meet the conditions required for the transition but actually achieves the transition, it means that the node has a vulnerability, so The mathematical form of is: ; in, is a custom Boolean function that indicates whether there is a vulnerability on node n that allows transition t. If the activities m and t detected on node n do not satisfy the legal transition condition C , it means that node n has a vulnerability.

[0054] To measure the defense strength of the system, Petri nets can be used to simulate the process of an attacker attacking a LAN twice. First, the probability of an attacker breaking into the LAN when there is no honeypot in the network is simulated. Then, Petri nets are used to simulate the probability of an attacker breaking into the LAN when the current honeypot deployment strategy is deployed. The difference between the two probabilities represents the defense strength of the current honeypot deployment strategy.

[0055] Assuming that an enterprise network contains multiple servers, each server corresponds to a node in the Petri net. Attackers can try to break into the server through different attack paths (transitions in the Petri net). Then two business scenarios can be set: one is a business scenario without a honeypot system, where attackers can directly attack the real server; the other is a business scenario with a honeypot system, where honeypots are deployed in the network to trap attackers and reduce the probability of successful attacks. On this basis, it can be assumed that the node set of the network topology is , represents each server in the enterprise network, assuming the transition , represents different attack paths of the attacker, assuming that condition C contains the triggering conditions of each transition t, such as the vulnerabilities required for the attack.

[0056] Assume there is a set of attack paths , the probability of success of each path is recorded as , then in the business scenario without the honeypot system, the probability of the system (i.e., LAN) being hacked is for: ; In a business scenario with a honeypot system, the honeypot can detect and block some attack paths. Assume that the set of attack paths that the honeypot can detect and block is , then the probability of the system (i.e. LAN) being hacked is for: ; in, is the success probability of each path in the system, usually .

[0057] System defense strength It can be expressed as: .

[0058] Furthermore, based on the probability of the node being hacked, the system damage after the node is hacked is determined.

[0059] Specifically, the probability of a node being hacked calculated above is added to the reward function of the game. The factors previously considered, such as historical interaction experience, number of system vulnerabilities, and system defense strength, are involved in the game process of the final deployment strategy. Network defenders will incur a fixed cost when placing new honeypots at the edge of the network. This cost is the average cost of each honeypot. Assuming that the cost of deploying a honeypot is is fixed, this cost prevents defenders from placing honeypots anywhere in the network. The monetary and operational costs of honeypots need to be considered, including the overhead of system performance. If the defender places a honeypot on the same edge that the attacker exploits, the defender will receive a defense reward. Therefore, the trade-off faced by the defender is to reduce the cost of placing a honeypot while increasing the chance of catching the attacker by placing more honeypots. The game process also needs to accommodate the additional constraint on the maximum number of honeypots to be allocated.

[0060] On the attacker side, the cost of an attack is recorded as , represents the risk taken by the attacker. If the attacker exploits the security margin, the attacker will receive a successful attack reward. Assume Capture rewards for defenders, A reward for a successful attacker.

[0061] This embodiment adopts a reward function that takes into account the importance of nodes in the network topology. Both the capture reward and the successful attack reward are determined by the weighted value of the protected or attacked node. Determine, among which , then the general reward matrix can be expressed as: ; Attacker Reward Matrix The loss function calculated based on the probability of a node being hacked obtained from the logistic regression model is added to the reward function. The reward function can be extended to any number of possible edges, so the probability of a node being hacked can be included in the cost and reward calculation process.

[0062] Cost part: The defender can regard the probability of a node being breached as the risk cost of a defense failure. When calculating the cost of a defense strategy, the probability of a defense failure can be multiplied by the corresponding cost factor and added to the total cost.

[0063] Reward part: The defender can regard the probability of successful defense as a reward for successful defense. When calculating the benefits of the defensive strategy, the probability of successful defense can be multiplied by the reward factor and added to the total benefit.

[0064] Based on this, the calculation formula for the system damage after the node is compromised is: ; in, System damage after the node is compromised; is the probability of a node being hacked; For Node The weighted value of It is a custom preset coefficient.

[0065] It is understandable that since the importance of each node in the system is different, for example, the damage caused to the system after some database servers are hacked is definitely greater than that of other nodes, each node can be assigned an importance value to distinguish nodes of different importance.

[0066] Furthermore, after determining the system damage after the node is hacked, the reward function of the Stackelberg game is determined based on the system damage after the node is hacked, the attacker's action space, and the defender's action space. The calculation formula of the reward function is: ; in, is the reward function; The cost of deploying a honeypot for the defender; The cost of carrying out an attack for the attacker; Capture rewards for defenders; Reward attackers for success; For Node The weighted value of is the preset coefficient; System damage after the node is compromised; is the defensive actions contained in the defender action space, Indicates that the defensive action is the defender on the side < >Deploy honeypots, It means that the defender does not deploy honeypots; is the attack action contained in the attacker's action space, Indicates that the attack action is the attacker's decision to attack the node , Indicates that the attack action is the attacker's decision to attack the node , It means that the attacker does not commit an attack; A set of nodes representing the network topology; For Node The weighted value of .

[0067] In some embodiments, the probability of a node being compromised is calculated as: ; in, is the probability of a node being hacked; is a natural constant; are the parameters of the logistic regression model; for historical interactive experience; is the number of system vulnerabilities; The system defense strength.

[0068] In some embodiments, the calculation formula for the system damage after a node is compromised is: ; in, System damage after the node is compromised; is the probability of a node being hacked; For Node The weighted value of is the preset coefficient.

[0069] In some embodiments, the reward function is calculated as: ; in, is the reward function; The cost of deploying a honeypot for the defender; The cost of carrying out an attack for the attacker; Capture rewards for defenders; Reward attackers for success; For Node The weighted value of is the preset coefficient; System damage after the node is compromised; is the defensive actions contained in the defender action space, Indicates that the defensive action is the defender on the side < >Deploy honeypots, It means that the defender does not deploy honeypots; is the attack action contained in the attacker's action space, Indicates that the attack action is the attacker's decision to attack the node , Indicates that the attack action is the attacker's decision to attack the node , It means that the attacker does not commit an aggressive act; A set of nodes representing the network topology; For Node The weighted value of .

[0070] In some embodiments, based on the reward function, a hybrid honeypot deployment strategy is determined, including: based on the reward function, determining the defender linear equation; solving the defender linear equation to obtain the target defense strategy of the Stackelberg game; based on the reward function, determining the attacker linear equation; solving the attacker linear equation to obtain the target attack strategy of the Stackelberg game; determining the hybrid honeypot deployment strategy based on the target defense strategy and the target attack strategy.

[0071] In this embodiment, the two parties involved in the Stackelberg game are the defender and the attacker in the system. By adding factors such as the defense experience accumulated in the system and the defense strength of the system into the calculation process of the reward function, a hybrid honeypot deployment strategy with a high benefit-cost ratio and good defense effect can be finally selected.

[0072] For the formulated Stackelberg game Γ, it admits at least one equilibrium point in the mixed strategy.

[0073] Assuming a mixed strategy is distributed in the defender action space The probability vector on , then 0 ≤Γ≤ 1 represents the probability that the defender chooses a specific defense action. On the attacker side, assuming The attacker's action space For any mixed strategy chosen by the attacker, if the point Satisfy the conditions , then the point is considered to be an equilibrium point; similarly, for any mixed strategy , if the condition is met , the defender adopts a mixed strategy .

[0074] Therefore, there is Theorem 1: For a finite game Γ, there exists at least one point Mixed balance.

[0075] For every possible joint action and , we can calculate the reward function of the matrix game A The value of , and then solve the game for the two players. To find the equilibrium point of the finite game Γ , a linear equation (LP) needs to be solved for each participant.

[0076] Specifically, based on the reward function, the defender's linear equation is determined, and the defender's linear equation is solved to obtain the target defense strategy of the Stackelberg game; based on the reward function, the attacker's linear equation is determined, and the attacker's linear equation is solved to obtain the target attack strategy of the Stackelberg game.

[0077] Among them, the defender linear equation is: ; ; ; ; The goal of solving the defender linear equation is to solve the defender mixed strategy The first condition ensures that the mixed strategy for each action is the action taken by the attacker. The best response to this condition represents a set of conditions, because the defender needs to consider all possible attack operations; the remaining two conditions are used to ensure a mixed strategy is a valid probability vector.

[0078] Similarly, the attacker linear equation attempts to minimize the same expected payoff by ensuring that it is the best response to all actions of the defender, then the attacker linear equation is: ; ; ; .

[0079] Solving the above defender linear equation and attacker linear equation, we can get the target defense strategy and targeted attack strategies .

[0080] Assume that there are two nodes in the network topology connected by an edge. The attacker can To enter the network, the attacker must decide whether to attack the node Or retreat, the defender must decide whether to allocate honeypots to existing edges or save the cost of honeypot allocation. In order to derive the game equilibrium with mixed strategies, let the attacker attack the weighted value with probability y The node v of the defense party is assigned a honeypot with probability x, then we have the following formula:

[0081] For each attack strategy adopted by the attacker, the defender’s reward is:

[0082] From this, we can see that the defender's expected reward can form a linear equation. If the slope of this linear equation is negative, then the defender's best response to the attacker's strategy y is , that is, no honeypot is assigned on this edge; if the slope of this linear equation is positive, the best response of the defender is , that is, at the edge Assign honeypots; if the slope of this linear equation is 0, then the defender responds to strategy y, and therefore, it satisfies the equilibrium point of the game.

[0083] Furthermore, based on the target defense strategy and the target attack strategy, a hybrid honeypot deployment strategy is determined.

[0084] Lemma 1: For the Stackelberg game Γ, if the allocation cost If the following conditions are met, the defender will definitely allocate a honeypot: .

[0085] Lemma 2: For the Stackelberg game Γ, if the attacker successfully rewards The attacker decides to back off if the following conditions are met:

[0086] Based on this, there is Theorem 2: The formulated game has a NE (equilibrium point) in the mixed strategy, satisfying: ; .

[0087] This NE (balance point) is the final hybrid honeypot deployment strategy.

[0088] Furthermore, the game admits NE in pure strategy, when the defender allocates a honeypot, if Lemma 1 and , the attacker retreats.

[0089] The honeypot deployment method based on attack and defense benefits provided in this embodiment calculates the probability of an attacker successfully bypassing a honeypot through machine learning and calculates the corresponding loss function, adds the probability of an attacker successfully breaking through a node to the calculation process of the cost and benefit of honeypot deployment, and finally quantifies the costs and benefits of both the attacker and the defender when deploying different hybrid honeypot strategies, and gambles out an optimal hybrid honeypot deployment strategy. This hybrid honeypot deployment method can dynamically adjust the strategy as the interaction experience accumulates, and ultimately outperforms the traditional honeypot deployment method in terms of expected benefits, and shows stronger adaptability and higher average benefits in long-term games. The software-defined honeypot deployment method for attack and defense benefits based on Stackelberg game not only improves the overall hybrid honeypot deployment defense effect of the system, reduces the cost of hybrid honeypots in the system, but also improves the system's benefit-cost ratio. At the same time, due to the deployment of multiple honeypots, it also avoids the phenomenon that the system is attacked by attackers who break into a honeypot and then use the same vulnerability to break into other nodes.

[0090] In order to verify the effectiveness of the honeypot deployment method based on attack and defense benefits provided in this embodiment, this embodiment adopts different levels of attacks, different scales of network environments, different levels of defense measures, etc., to ensure that the software-defined honeypot deployment method based on Stackelberg game for attack and defense benefits is universal.

[0091] See also Figure 3 to Figure 4 , Figure 3 This is one of the system benefit comparison diagrams of the honeypot deployment method based on attack and defense benefits provided by the present invention and the traditional method. Figure 4 This is the second system benefit comparison chart between the honeypot deployment method based on attack and defense benefits provided by the present invention and the traditional method.

[0092] It should be noted that Figure 3 and Figure 4 The only difference is: Figure 3 For a line chart, Figure 4 For a bar chart. Figure 3 and Figure 4 As shown in the figure, when the attacker adopts a relatively conservative strategy, in most cases, the system benefit of the honeypot deployment method based on attack and defense benefits (blue part) is higher than the system benefit of the traditional honeypot deployment method (red part).

[0093] See also Figures 5 and 6 , Figure 5 This is the third system benefit comparison diagram between the honeypot deployment method based on attack and defense benefits provided by the present invention and the traditional method. Figure 6 This is the fourth system benefit comparison chart between the honeypot deployment method based on attack and defense benefits provided by the present invention and the traditional method.

[0094] It should be noted that Figure 5 and Figure 6 The only difference is: Figure 5 For a line chart, Figure 6 For a bar chart. Figure 5 and Figure 6 As shown in the figure, when the attacker adopts a relatively aggressive strategy, in most cases, the system benefit of the honeypot deployment method based on attack and defense benefits (blue part) is higher than the system benefit of the traditional honeypot deployment method (red part).

[0095] The present invention also provides a honeypot deployment device based on attack and defense benefits. Figure 7 , Figure 7 7 is a schematic diagram of the structure of the honeypot deployment device based on attack and defense benefits provided by the present invention. In this embodiment, the honeypot deployment device based on attack and defense benefits includes a network topology modeling submodule 710, a participant modeling submodule 720, a reward function determination submodule 730 and a game submodule 740.

[0096] The network topology modeling submodule 710 is used to model the network environment and generate the network topology.

[0097] The participant modeling submodule 720 is used to model the attacker and the defender respectively based on the network topology to obtain the attacker action space and the defender action space.

[0098] The reward function determination submodule 730 is used to determine the reward function of the Stackelberg game based on historical interaction experience, the number of system vulnerabilities, the system defense strength, the attacker action space and the defender action space.

[0099] The game submodule 740 is used to determine the hybrid honeypot deployment strategy based on the reward function.

[0100] Among them, the hybrid honeypot deployment strategy is a deployment strategy that includes software-defined honeypots and non-software-defined honeypots.

[0101] In some embodiments, the reward function determination submodule 730 is used to determine the probability of a node in the network topology being hacked based on a logistic regression model according to historical interaction experience, the number of system vulnerabilities and the strength of system defense; determine the system damage after the node is hacked based on the probability of the node being hacked; and determine the reward function of the Stackelberg game based on the system damage after the node is hacked, the attacker's action space and the defender's action space.

[0102] In some embodiments, the probability of a node being compromised is calculated as: ; in, is the probability of a node being hacked; is a natural constant; are the parameters of the logistic regression model; for historical interactive experience; is the number of system vulnerabilities; The system defense strength.

[0103] In some embodiments, the calculation formula for the system damage after a node is compromised is: ; in, System damage after a node is compromised; is the probability of a node being hacked; For Node The weighted value of is the preset coefficient.

[0104] In some embodiments, the reward function is calculated as: ; in, is the reward function; The cost of deploying a honeypot for the defender; The cost of carrying out an attack for the attacker; Capture rewards for defenders; Reward attackers for success; For Node The weighted value of is the preset coefficient; System damage after a node is compromised; is the defensive actions contained in the defender action space, Indicates that the defensive action is the defender on the side < >Deploy honeypots, It means that the defender does not deploy honeypots; is the attack action contained in the attacker's action space, Indicates that the attack action is the attacker's decision to attack the node , Indicates that the attack action is the attacker's decision to attack the node , It means that the attacker does not commit an attack; A set of nodes representing the network topology; For Node The weighted value of .

[0105] In some embodiments, the game submodule 740 is used to determine the defender linear equation based on the reward function; solve the defender linear equation to obtain the target defense strategy of the Stackelberg game; determine the attacker linear equation based on the reward function; solve the attacker linear equation to obtain the target attack strategy of the Stackelberg game; determine the hybrid honeypot deployment strategy based on the target defense strategy and the target attack strategy.

[0106] The invention also provides an electronic device. Figure 8 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the honeypot deployment method based on attack and defense benefits.

[0107] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0108] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the honeypot deployment method based on attack and defense benefits provided by the above methods.

[0109] The present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the honeypot deployment method based on attack and defense benefits provided by the above methods.

[0110] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0111] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A honeypot deployment method based on attack and defense benefits, characterized in that: include: Model the network environment and generate network topology; Based on the network topology, the attacker and the defender are modeled respectively to obtain the attacker action space and the defender action space; Determine a reward function of the Stackelberg game based on historical interaction experience, the number of system vulnerabilities, the system defense strength, the attacker action space, and the defender action space; Based on the reward function, determining a hybrid honeypot deployment strategy; The hybrid honeypot deployment strategy is a deployment strategy that includes software-defined honeypots and non-software-defined honeypots.

2. The honeypot deployment method based on attack and defense benefits according to claim 1 is characterized in that: The reward function of the Stackelberg game is determined based on historical interaction experience, the number of system vulnerabilities, the system defense strength, the attacker action space and the defender action space, including: Determine the probability of a node in the network topology being breached based on the historical interaction experience, the number of system vulnerabilities, and the system defense strength based on a logistic regression model; Based on the probability of the node being hacked, determining the system damage after the node is hacked; Based on the system damage after the node is compromised, the attacker's action space and the defender's action space, a reward function of the Stackelberg game is determined.

3. The honeypot deployment method based on attack and defense benefits according to claim 2 is characterized in that: The calculation formula for the probability of the node being breached is: ; in, is the probability of the node being hacked; is a natural constant; are the parameters of the logistic regression model; For said historical interactive experience; is the number of system vulnerabilities; is the defense strength of the system.

4. The honeypot deployment method based on attack and defense benefits according to claim 2 is characterized in that: The calculation formula for the system damage after the node is compromised is: ; in, The system damage after the node is compromised; is the probability of the node being hacked; For Node The weighted value of is the preset coefficient.

5. The honeypot deployment method based on attack and defense benefits according to claim 2 is characterized in that: The calculation formula of the reward function is: ; in, is the reward function; The cost of deploying a honeypot for said defender; the cost of carrying out an attack for said attacker; Capture rewards for defenders; Reward attackers for success; For Node The weighted value of is the preset coefficient; The system damage after the node is compromised; is the defensive actions contained in the defender action space, Indicates that the defensive action is that the defender is at the edge < >Deploy honeypots, Indicates that the defender does not deploy a honeypot; is the attack action contained in the attacker action space, Indicates that the attack action is the attacker's decision to attack the node , Indicates that the attack action is the attacker's decision to attack the node , Indicates that the attacker does not commit an attack; A set of nodes representing the network topology; For Node The weighted value of .

6. The honeypot deployment method based on attack and defense benefits according to claim 1 is characterized in that: Determining the hybrid honeypot deployment strategy based on the reward function includes: Based on the reward function, determining a defender linear equation; Solving the defender linear equation to obtain the target defense strategy of the Stackelberg game; Based on the reward function, determining an attacker linear equation; Solving the attacker linear equation to obtain the target attack strategy of the Stackelberg game; Based on the target defense strategy and the target attack strategy, the hybrid honeypot deployment strategy is determined.

7. A honeypot deployment device based on attack and defense benefits, characterized in that: include: The network topology modeling submodule is used to model the network environment and generate the network topology; A participant modeling submodule is used to model the attacker and the defender respectively based on the network topology to obtain the attacker action space and the defender action space; A reward function determination submodule, used to determine the reward function of the Stackelberg game based on historical interaction experience, the number of system vulnerabilities, the system defense strength, the attacker action space and the defender action space; A game submodule, used to determine a hybrid honeypot deployment strategy based on the reward function; The hybrid honeypot deployment strategy is a deployment strategy that includes software-defined honeypots and non-software-defined honeypots.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the honeypot deployment method based on attack and defense benefits as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the honeypot deployment method based on attack and defense benefits as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the honeypot deployment method based on attack and defense benefits as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Dynamic honey array defense strategy generation method, system and device based on Stackelberg game and medium

    CN121664448A