Dynamic honey array defense strategy generation method, system and device based on Stackelberg game and medium
By deploying multiple types of honeypot devices in the network, real-time attack behavior data is collected, and the configuration of honeypots is optimized using Stackelberg game theory and Markov decision processes. This solves the problem that defense strategies in existing technologies cannot adapt to changes in the network environment, and enables dynamic response to attack behavior and continuous optimization of defense strategies, thereby improving network security protection capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-03-13
AI Technical Summary
Existing deception defense technologies lack the ability to adapt to real-time changes in the network environment and attack patterns, making it difficult to conduct in-depth analysis and accurate prediction of attacker behavior. As a result, defense strategies cannot achieve optimal results and cannot maximize the effectiveness of network security defense.
By deploying multiple types of honeypot devices to collect attack behavior data in real time, the honeypot array is automatically changed based on a preset triggering mechanism. A game model between the defender and the attacker is constructed using a multi-round Stackelberg game architecture. The honeypot deployment strategy is solved by combining Markov decision process and mixed integer linear programming, and the honeypot configuration is dynamically adjusted. The system is optimized by calculating quantitative indicators such as attack detection rate, honeypot failure rate, value of the protected system, and system stability.
It enables precise response to attacks and continuous improvement of defense strategies, enhances the intelligence level of network security protection, strengthens the ability to respond to complex network attacks, and ensures maximum defense effectiveness with limited resources.
Smart Images

Figure CN121664448A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security technology, and in particular to a method, system, device, and medium for generating dynamic honeycomb defense strategies based on Stackelberg game theory. Background Technology
[0002] In the field of cybersecurity, deception defense technology, as a proactive defense method, has made significant progress in recent years. By deploying deception and decoy devices in the defender's network information system, it is possible to effectively interfere with and mislead attackers' perception and judgment of the protected system, inducing attackers to make decisions and actions favorable to the defender, thereby achieving the purpose of detecting, delaying, or blocking attacker activities.
[0003] However, existing deception defense technologies still have some limitations: On the one hand, most traditional deception defense systems are statically configured, lacking the ability to adapt to real-time changes in the network environment and attack patterns. This makes these devices easily identifiable by experienced attackers and gradually render them ineffective. On the other hand, existing systems often fail to fully consider the potential reactions and strategies of attackers, making it difficult to conduct in-depth analysis and accurate prediction of attacker behavior. Consequently, defense strategies cannot achieve optimal results and cannot maximize the effectiveness of network security defense. Furthermore, while some deception defense technologies can record attack behavior, they still lack effective solutions for dynamically adjusting defense strategies based on attacker behavior to cope with constantly evolving network threats. Summary of the Invention
[0004] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for generating dynamic honeypot defense strategies based on Stackelberg game theory, including real-time collection of attack behavior data through multiple types of honeypot devices deployed in the network; When the attack behavior data reaches the preset trigger mechanism, the honey array will automatically change; A target game model of defenders and attackers is constructed based on a multi-round Stackelberg game architecture, and the honeypot deployment strategy is solved by Markov decision process and mixed integer linear programming. Based on the honeypoint deployment strategy, the honeypoint repository resources are invoked, and dynamic adjustments to the honeypoint configuration are automatically performed. Based on the adjusted honeypot configuration and the attack behavior data, the target quantitative index is calculated. The target game model is optimized based on the calculation results of the target quantitative index.
[0005] As a preferred embodiment of the dynamic honeycomb defense strategy generation method based on Stackelberg game theory described in this invention, wherein: the construction of the target game model based on a multi-round Stackelberg game architecture includes, Network topology definition, honeypot policy space and reward function definition, and attacker policy space and reward function definition; The network topology definition represents the network as a directed graph (G = (N, E)), where the set of nodes N represents the nodes in the network, the set of edges E represents the connection relationship between nodes, and the set of entry nodes is determined as the starting point for the attacker's penetration.
[0006] As a preferred embodiment of the dynamic honeycomb defense strategy generation method based on Stackelberg game theory described in this invention, wherein: the calculation of the target quantification index includes, Calculate the attack detection rate, honeypot rejection rate, value of the protected system, and system stability; among which, The attack detection rate refers to the proportion of attacks that are successfully detected. The honey trigger rate refers to the proportion of times an attacker triggers a honey spot.
[0007] As a preferred embodiment of the dynamic honeycomb defense strategy generation method based on Stackelberg game theory described in this invention, the preset triggering mechanism includes a time-slice triggering mechanism and a system event triggering mechanism; wherein, The time slice triggering mechanism refers to triggering honey array transformation when a preset threshold number of game iterations is reached. The system event triggering mechanism refers to triggering honey array transformation when the frequency of honey stepping exceeds the dynamic threshold or a high-risk attack mode is identified.
[0008] As a preferred embodiment of the dynamic honeypot defense strategy generation method based on Stackelberg game theory described in this invention, the attack detection rate is calculated as the ratio of the number of times the honeypot set is triggered to the total number of attack events. The formula for calculating the honey-triggering rate is the ratio of the number of paths that trigger honey points to the total number of paths; The value of the protected system is calculated based on the sum of the values of high-value hosts that have not been attacked and the value of attack behaviors that are honeypot traps; The system stability is measured by the degree of difference between the honeypot deployment strategy in the k-th iteration and the average strategy.
[0009] As a preferred embodiment of the dynamic honeycomb defense strategy generation method based on Stackelberg game theory described in this invention, the high-risk attack mode includes at least one of SQL injection, brute force attack, and lateral movement.
[0010] As a preferred embodiment of the dynamic honeycomb defense strategy generation method based on Stackelberg game theory described in this invention, the optimization of the target game model based on the calculation results of the target quantification index includes: The target quantitative index is input into the target game model, and the decision parameters are updated through a strategy iteration algorithm.
[0011] Secondly, the present invention provides a dynamic honeypot defense strategy generation system based on Stackelberg game theory, comprising: a data acquisition module, used to collect attack behavior data in real time through multiple types of honeypot devices deployed in the network; The change module is used to automatically change the honey array when the attack behavior data reaches a preset trigger mechanism; A solution module is constructed to build a target game model between the defender and the attacker based on a multi-round Stackelberg game architecture, and to solve the honeypot deployment strategy through Markov decision process and mixed integer linear programming. The adjustment module is used to call the honeypoint repository resources based on the honeypoint deployment strategy and automatically perform dynamic adjustments to the honeypoint configuration; The calculation module is used to calculate the target quantitative index based on the adjusted honeypot configuration and the attack behavior data; An optimization module is used to optimize the target game model based on the calculation results of the target quantitative index.
[0012] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.
[0013] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0014] Compared with existing technologies, the beneficial effects of this invention are as follows: By deploying multiple types of honeypot devices in the network security field to collect attack behavior data in real time, and realizing the automatic change of the honeypot array based on a preset trigger mechanism, a game model between defenders and attackers is constructed using a multi-round Stackelberg game architecture. This is combined with Markov decision processes and mixed-integer linear programming to solve the honeypot deployment strategy, achieving dynamic adjustment of honeypot configuration and precise response to attack behavior. Simultaneously, by calculating quantitative indicators such as attack detection rate, honeypot failure rate, protected system value, and system stability, and feeding these quantitative indicators back into the model optimization process, continuous improvement of the defense strategy is achieved. Furthermore, this method also covers the identification and response to high-risk attack patterns, effectively improving the intelligence level of network security protection, enhancing the ability to respond to complex network attacks, and ensuring maximum defense effectiveness with limited resources. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating the method for generating dynamic honeycomb defense strategies based on Stackelberg game theory.
[0017] Figure 2 This is a functional flowchart illustrating the method for generating dynamic honeycomb defense strategies based on Stackelberg game theory.
[0018] Figure 3 This is a schematic diagram of the system architecture for a dynamic honeycomb defense strategy generation method based on Stackelberg game theory. Detailed Implementation
[0019] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0020] Example 1, referring to Figure 1 This is the first embodiment of the present invention, which provides a method for generating dynamic honeycomb defense strategies based on Stackelberg game theory, including: S100: Collects attack behavior data in real time through multiple types of honeypot devices deployed in the network; S200: When the attack behavior data reaches the preset trigger mechanism, the honey array is automatically changed; S300: Based on a multi-round Stackelberg game architecture, a target game model between defenders and attackers is constructed, and the honeypot deployment strategy is solved by Markov decision process and mixed integer linear programming. S400: Based on the honeypot deployment strategy, call the honeypot repository resources and automatically perform dynamic adjustments to the honeypot configuration; S500: Calculate the target quantification index based on the adjusted honeypot configuration and the attack behavior data; S600: Optimize the target game model based on the calculation results of the target quantitative index.
[0021] It should be noted that in modern network security defense systems, the confrontation between attackers and defenders is constantly intensifying. Attackers often exploit the complexity and dynamism of the network environment, continuously trying new attack methods and strategies to bypass traditional defense mechanisms. For example, attackers will frequently scan and probe to find weaknesses in the network, such as vulnerable servers or misconfigured devices. Simultaneously, attackers will utilize various attack tools and scripts to launch attacks rapidly, such as SQL injection, brute-force attacks, and lateral movement, posing serious security threats to network systems. Existing defense systems, facing these complex and ever-changing attacks, often lack the ability to deeply analyze attacker behavior and dynamically respond, making it difficult to detect and prevent attacks in a timely manner, thus severely jeopardizing network system security.
[0022] Therefore, to address the aforementioned network security defense issues, the S100-S600 approach first involves deploying multiple types of honeypot devices to collect attack behavior data in real time, providing a foundation for subsequent analysis and decision-making. Then, when attack behavior data reaches a preset trigger mechanism, the honeypot array automatically changes to confuse and interfere with attackers. Next, a target game model between defenders and attackers is constructed based on a multi-round Stackelberg game architecture, and the honeypot deployment strategy is solved using Markov decision processes and mixed-integer linear programming to optimize the defense strategy. Based on the solved honeypot deployment strategy, honeypot repository resources are invoked to automatically and dynamically adjust the honeypot configuration, enhancing the uncertainty and adaptability of the defense system. Then, target quantitative indicators are calculated using the adjusted honeypot configuration and attack behavior data to evaluate the current defense effectiveness. Finally, the target game model is optimized based on the quantitative indicator calculation results, continuously improving the accuracy and effectiveness of the defense strategy. This achieves dynamic monitoring, intelligent analysis, and proactive defense against network attacks, effectively improving the security and anti-attack capabilities of the network system.
[0023] Example 2, refer to Figures 1-3 As an embodiment of the present invention, based on the above embodiment, a method for generating dynamic honeycomb defense strategies based on Stackelberg game theory is provided.
[0024] In this embodiment of the application, step S100 involves collecting attack behavior data in real time using various types of honeypot devices deployed in the network, including the following step A1: A1: Attack behavior data includes information such as the number of times the attacker tried to access the honeypot, the IP address of the attacker, the attack service, and the attack type.
[0025] It should be noted that by deploying log collectors at key honeypots, the system can efficiently capture attacker behavior patterns (such as frequent scanning and file access attempts) and transmit them to a central database via a Kafka cluster, ensuring the real-time nature of log data and the efficiency of analysis. Regularly updating attack types further improves data availability and analytical accuracy.
[0026] It should be further clarified that critical honeypots refer to honeypots deployed in a network system that have high strategic value or high decoy potential. They are typically located in key positions within the network topology or near high-value hosts, designed to maximize the probability of attackers triggering the attack and to collect high-quality security logs. Configuring honeypots automatically records all relevant security events, including but not limited to the number of times the honeypot was accessed, the IP address of the accessing honeypot, the attacking service, and the attack type.
[0027] In an optional implementation, the collection of attack behavior data in step S100 can also be based on software-defined networking. That is, by utilizing the centralized control and flexible management capabilities of software-defined networking (SDN), traffic data in the network is copied to the honeypot data collection system by deploying traffic mirroring and sampling functions on SDN switches. This system performs deep packet inspection (DPI) on the traffic to identify attack characteristics and behavior patterns, thereby collecting attack behavior data.
[0028] In another optional implementation, the collection of attack behavior data in step S100 can also be based on a combination of active probing and deception. That is, active probing technology is used to simulate the behavior of real users or devices in the network to actively scan and probe the network. When suspicious network nodes or behaviors are detected, they are actively guided to interact with different types of honeypots deployed in the network, thereby collecting attack behavior data.
[0029] In this embodiment of the application, when the attack behavior data reaches a preset trigger mechanism, step S200 involves automatic changes to the honeycomb array, including the following steps B1-B3: Understandably, honeypot transformation aims to enhance deception and reduce the probability of attackers bypassing honeypots by adjusting the type, location, and decoy content of honeypots.
[0030] B1: The preset triggering mechanism includes a time-slice triggering mechanism and a system event triggering mechanism; B2: The time slice triggering mechanism refers to triggering honey array transformation when the preset game iteration number threshold is reached; the system event triggering mechanism refers to triggering honey array transformation when the frequency of honey stepping exceeds the dynamic threshold or a high-risk attack mode is identified.
[0031] It should be noted that triggering the honeypot transformation when the preset game iteration threshold is reached means that the honeypot transformation will be automatically triggered after 5 to 8 game iterations. The purpose of this is to ensure that the honeypot layout is unpredictable, thereby increasing the difficulty for attackers to identify the real system. Furthermore, when a honeypot detects abnormal behavior (such as the number of honeypot visits exceeding the preset threshold, the detection of high-frequency attack patterns or specific attack types), the honeypot transformation will be triggered immediately. Abnormal behavior includes, but is not limited to: the same IP triggering the honeypot multiple times in a short period of time, the detection of known attack patterns, and abnormal network traffic reported by the system monitoring module.
[0032] B3: The high-risk attack modes include at least one of SQL injection, brute force, and lateral movement.
[0033] Specifically, in step 200, the honeypot change rules are based on network topology, host value rating, and attacker behavior analysis. The specific change rules are executed by the dynamic honeypot adjustment module, based on the strategy iteration results of the game theory module and real-time data from the log collection module. The system uses automated scripts and a predefined rule base to dynamically call resources in the honeypot repository according to trigger conditions, quickly completing honeypot deployment and revocation. It should be noted that all change operations are recorded in the database, and the honeypot evaluation module analyzes the change effects and feeds them back to the game theory module to optimize subsequent strategies.
[0034] Even better, this dynamic strategy triggering mechanism enables the honeycomb system to flexibly adjust according to real-time network conditions and attack behavior, thereby improving the intelligence level of network security defense.
[0035] In an optional implementation, the automatic change of the honeycomb array in step S200 when the attack behavior data reaches a preset trigger mechanism can also be based on anomaly detection triggered by machine learning. This involves collecting attack behavior data from the network, including datasets of normal and attack behaviors, and performing preprocessing operations such as cleaning and normalization on the data to suit the input requirements of the machine learning model. A suitable machine learning algorithm, such as Support Vector Machine (SVM), Random Forest, or neural networks in deep learning, is selected to train the collected data. The training objective is to enable the model to distinguish between normal and anomalous attack behaviors and to assess the severity of attack behaviors. Finally, the trained model is deployed in the network environment to detect real-time attack behavior data. When the model detects that the attack behavior data reaches a preset anomaly threshold, such as when certain features indicate a potential Advanced Persistent Threat (APT) attack, the automatic change of the honeycomb array is triggered, adjusting the deployment strategy of honey spots to cope with such complex attack scenarios.
[0036] In another optional implementation, the automatic change of the honey array in step S200 when the attack behavior data reaches the preset trigger mechanism can also be based on predictive triggering of attacker behavior patterns. This involves in-depth analysis of the attacker's historical behavior data to identify common attack patterns and paths. While monitoring attack behavior in the network in real time, the current attack behavior data is matched with known attack pattern templates (using algorithms such as Dynamic Time Warping (DTW) to measure the similarity between the current behavior and the template). Based on the matching results and similarity thresholds, the attacker's next possible behavior is predicted. If it is predicted that an attacker is about to enter a critical area of the network or launch an attack on important assets, the automatic change of the honey array is triggered in advance to enhance the defense capabilities of critical areas and induce the attacker to enter the pre-set honey spot area, thereby effectively protecting real assets.
[0037] In this embodiment of the application, step S300 involves constructing a target game model between the defender and the attacker based on a multi-round Stackelberg game architecture, and solving the honeypot deployment strategy using Markov decision processes and mixed-integer linear programming, including the following steps C1-C2: Understandably, the model, built upon a multi-round Stackelberg game, simulates the dynamic strategic interaction between the defender (honeycomb system) and the attacker. In this model, the defender, as the leader, first selects honeycomb configurations and changes strategies; the attacker, as the follower, selects attack paths and methods based on the defender's strategies, optimizing the defense effect through multiple rounds of iteration.
[0038] C1: The target game model constructed based on the multi-round Stackelberg game architecture includes, The definitions of network topology, honeycomb strategy space and reward function, and attacker strategy space and reward function are as follows: The network topology is defined as representing the network as a directed graph (G=(N,E)), where the set of nodes N represents the nodes in the network, the set of edges E represents the connection relationship between the nodes, and the set of entry nodes is determined as the attacker's penetration starting point.
[0039] It is understandable that the set of nodes N = {1, 2, ..., i} has value for each node i ∈ N.
[0040] It should be noted that the strategy space for the honeypot (defender) refers to the choice of which nodes to deploy honeypots on. The specific decision set for honeypot deployment is defined as: ∈{0,1}, ∀i∈N, where =1 indicates that a honeypot is deployed at node i.
[0041] Rewards are distributed between the honeycomb array and the attacker based on their respective strategies. The goal of the honeycomb array is to maximize network security after honeypot deployment, and its reward function is defined as follows: ; In the formula, A is the set of nodes that an attacker might attack; H is the set of nodes that the defender successfully defends against; V i C represents the value of node i; i It is the cost of deploying a honeypot on node i; x i ∈{0,1} indicates whether the defender deploys a honeypot at node i; β is a coefficient used to measure the additional value or reward brought by successfully protecting the node. Furthermore, This represents the net revenue generated from deploying honeypots on all nodes. If a honeypot is deployed on a specific node, the revenue is the difference between the node's value and the deployment cost. This demonstrates the additional reward derived from deploying honeypots on nodes that attackers might target. This is likely because successfully deploying honeypots on these critical nodes can more effectively confuse or block attackers, thus gaining an extra value bonus. This takes into account the losses incurred by nodes that were actually successfully attacked by the attacker. If the attacker successfully attacks a node that was not successfully protected (i.e., a node that belongs to the attack set but not to the successful protection set), then the defender needs to bear the value loss of these nodes.
[0042] Furthermore, regarding the attacker's policy space: it involves choosing an attack path from the entry point to the final target. The decision variable for path selection is: y (i,j) ∈{0,1},∀(i,j)∈E. When y (i,j) =1 indicates that the attacker chooses the path from node i to node j to launch the attack; when y (i,j) When =0, it means that the path is not selected.
[0043] As a follower in the game, the attacker reacts to the defender's honeypot deployment strategy by choosing to attack node j and maximizing their attack gains even if the defenses are bypassed. Their objective is to select a path that maximizes their gains, and the reward function is defined as follows: ; In the formula: P is the set of node pairs on the attack path; V j It is the value of node j; C j γ is the cost of deploying the honeypot on node j; γ is the cost ratio for the attacker. Furthermore, This means that when an attacker attacks node j, if the defender has not deployed a honeypot on that node (i.e., x), the attacker will succeed. j If V = 0, then the attacker can obtain the full value V of node j. j If the defender has already deployed a honeypot at this node (i.e., x) j If the value is 1, then the attacker cannot obtain the value of the node; This indicates that attackers need to consider the cost of their attack during the attack process, γC j It reflects the cost incurred by the attacker when attacking node j, and γ is used to adjust the weight of the cost; This indicates that if an attacker encounters a node j with a honeypot deployed during an attack, they will suffer a certain penalty or loss of value. α is used to measure the degree of this penalty.
[0044] C2: The model training uses a policy iteration algorithm, which combines Markov decision process (MDP) and mixed integer linear programming (MILP) to solve the Stackelberg equilibrium, ensuring the optimality of the defense strategy under the attacker's optimal response.
[0045] It should be noted that the training process specifically includes the following steps: 1) Initialization: In the initialization phase, we set the initial state of the network system and the initial strategy of the defenders.
[0046] 2) Policy Evaluation: Based on the current policy, calculate the value function of state s, reflecting the long-term utility of the policy. The value function is defined based on the Bellman expectation equation: ; In the formula: For immediate rewards, γ is the discount factor. This represents the state transition probability. It is calculated through iterative updates. Until convergence.
[0047] 3) Strategy Improvement: Based on the current value function Update strategy Choose the action that maximizes the expected value: ; 4) Status Updates and Feedback: Defender Real-Time Strategy The attacker observes and selects the optimal path p. The system then... Transition to a new state And update the attack behavior log D.
[0048] Preferably, compared with traditional deception defense techniques (such as static honeypots and single Stackelberg games), the advantages of this model include: 1) capturing the evolution of the attacker's strategy through multiple rounds of game playing, dynamically adjusting the honeypot deployment, and enhancing the long-term defense effect; 2) optimizing the long-term defense utility and balancing current and future benefits by defining states, actions, transition probabilities, and reward functions through MDP; 3) solving the bi-level optimization problem through MILP to ensure optimal honeypot deployment under resource constraints.
[0049] It should be further explained that a set of defense strategies for defenders and a set of attack strategies for attackers are constructed. The construction of the defense strategy set and the attack strategy set is based on a multi-round Stackelberg game model and Markov decision process (MDP), combined with network topology, host value rating, attack behavior logs and external threat intelligence. The aim is to optimize honeypot deployment to maximize defense effectiveness, while simulating attacker response behavior to predict their strategy evolution. The design of the strategy set comprehensively considers network security status, resource constraints and attacker dynamic behavior. Through multi-round game iteration optimization, it is ensured that the defense strategy can dynamically adapt to complex attack scenarios. Furthermore, the defense strategy includes decisions such as selecting different types of honeypots and which server to deploy the honeypots near. Attackers can choose attack paths and methods based on the system's response and the honeypot configuration they know. The attack strategy set is based on the assumption of the attacker's rational behavior. It combines the system response and honeypot configuration to simulate the attacker's path selection and attack methods in the network. Specifically, attackers observe the log feedback triggered by the honeypot (such as latency, error response) or system service characteristics (such as port open status). Attackers will prioritize the shortest path or high-value nodes (nodes with important assets) and attempt to bypass the honeypot.
[0050] Ideally, constructing strategy sets for both defenders and attackers allows defenders to proactively engage in strategic maneuvering with attackers, achieving dynamic and intelligent defense. Compared to traditional static defense methods, this is more effective in resisting cyberattacks, reducing the risk of network breaches, and improving the overall security level of the network. Furthermore, the continuous confrontation and evolution of the strategy sets between defenders and attackers can prompt both sides to discover and summarize lessons learned, driving the continuous improvement and optimization of security strategies. This helps to form a virtuous cycle, continuously improving network security protection capabilities and better addressing increasingly complex cybersecurity threats.
[0051] In this embodiment of the application, step S400, which calls the honeypot repository resources based on the honeypot deployment strategy and automatically performs dynamic adjustments to the honeypot configuration, includes the following steps D1-D2: It is understandable that dynamic adjustments to honeypot configuration include dynamic adjustments to the honeypot deployment location and type.
[0052] D1: Adjustment rules are based on network topology, host value rating, attack behavior patterns, and system performance indicators to ensure that the honeycomb array achieves maximum defense effectiveness with limited resources.
[0053] D2: The adjustment process is implemented through automated scripts, combining Markov Decision Process (MDP) and Mixed Integer Linear Programming (MILP) to dynamically optimize honeypot deployment.
[0054] In this embodiment of the application, step S500 calculates the target quantification index based on the adjusted honeypot configuration and the attack behavior data, including the following steps E1-E2: E1: Calculate the attack detection rate, honeypot detection rate, protected system value, and system stability; wherein, the attack detection rate refers to the proportion of attacks successfully detected; and the honeypot detection rate refers to the proportion of times an attacker triggers a honeypot.
[0055] In an optional implementation, the calculation of the value of the protected system in step E1 can also be based on a weighted calculation of key network assets. This involves comprehensively identifying key assets such as systems, data, and devices in the network and classifying them according to their importance and sensitivity. For example, core databases, critical business servers, and sensitive user information storage devices are classified as high-value assets. Corresponding weight coefficients are determined based on the importance of the assets. These weight coefficients can be determined through expert evaluation, historical security event analysis, or business impact analysis. After the honeycomb defense strategy is implemented, the protection status of each key asset over a certain period is statistically analyzed, such as no attack, minor attack, and severe attack. The asset value is then discounted according to different statuses. For example, the asset value is 100% of its original value when there is no attack, 80% during a minor attack, and 50% during a severe attack. The discounted value of each key asset is multiplied by its corresponding weight coefficient, and then summed to obtain the value of the protected system.
[0056] In another optional implementation, the calculation of the protected system value in step E1 can also be based on system value assessment of attack paths. This involves using security analysis tools and threat intelligence to identify potential attack paths within the network. An attack path refers to the path an attacker takes from an external network or initial penetration point, gradually advancing towards critical assets. The value of each node (system, device, etc.) along each attack path is assessed. Value assessment can consider factors such as the node's business importance, data sensitivity, and dependence on other systems. After the honeycomb defense strategy is implemented, the number of actually protected attack paths and the number of protected nodes on each path are counted. The actual protected values of all attack paths are summed, and the importance weight of each path in the overall network security is considered (which can be determined based on historical attack data, business relevance, etc.), ultimately yielding the value of the protected system.
[0057] E2: The attack detection rate is calculated as the ratio of the number of times the honeypot set is triggered to the total number of attack events; The formula for calculating the honey-triggering rate is the ratio of the number of paths that trigger honey points to the total number of paths; The value of the protected system is calculated based on the sum of the values of high-value hosts that have not been attacked and the value of attack behaviors that are honeypot traps; The system stability is measured by the degree of difference between the honeypot deployment strategy in the k-th iteration and the average strategy.
[0058] Specifically, in step E2, the attack detection rate reflects the honeycomb system's ability to identify attack behaviors, and its specific manifestation is as follows: ; In the formula: H is the set of honey spots, For honey spots The number of times it is triggered. For attack events The total number of times.
[0059] Furthermore, the honey-touching rate reflects the trapping effect of the honey array, specifically manifested in the following ways: ; in, H represents the set of paths chosen by the attacker, and H represents the set of honeypots. The number of paths that trigger honeypots. This represents the total number of paths.
[0060] Furthermore, the value of the protected system measures the honeycomb system's ability to protect high-value hosts. It is based on the total value of high-value hosts that have not been attacked and the value of attack behaviors captured by honeypots. Specifically, it is expressed as follows: ; In the formula: T is the set of real hosts, and A is the set of attacked nodes. For nodes The value of H is the set of honey spots. This is the trap reward coefficient.
[0061] Furthermore, system stability refers to the convergence and robustness of policy adjustments, reflecting the stability of the honeycomb array during dynamic transformations. Specifically, it manifests as follows: ; In the formula: Honeycomb deployment strategy for the t-th iteration , For the average strategy, T is the number of iterations.
[0062] In this embodiment of the application, step S600, which optimizes the target game model based on the calculation results of the target quantitative index, includes the following step F1: F1: Input the target quantitative index into the target game model and update the decision parameters through the strategy iteration algorithm.
[0063] In an optional implementation, the optimization of the target game model in step S600 can also be based on a reinforcement learning algorithm. Specifically, a reinforcement learning algorithm is used, with the target quantification index as the reward signal, to train an agent to optimize the parameters of the target game model. The agent continuously explores different parameter adjustment strategies through interaction with the environment, and learns the optimal parameter adjustment strategy based on feedback from the target quantification index (such as an increase in attack detection rate or honeypot detection rate).
[0064] In another optional implementation, the optimization of the target game model in step S600 can also be based on a genetic algorithm, specifically, using a genetic algorithm to optimize the network topology weights and reward function coefficients of the target game model. First, a set of model parameter populations is initialized. Then, genetic operations such as selection, crossover, and mutation are used to continuously evolve the individual parameters in the population. The fitness of each individual is determined by the calculated target quantification index; individuals with higher fitness are more likely to be retained and propagated. After multiple generations of evolution, the genetic algorithm can converge to a set of model parameters that optimize the target quantification index, thereby optimizing the target game model.
[0065] In summary, this invention deploys multiple types of honeypot devices in the network security field to collect attack behavior data in real time, and automatically changes the honeypot array based on a preset trigger mechanism. It utilizes a multi-round Stackelberg game architecture to construct a game model between defenders and attackers, and combines Markov decision processes and mixed-integer linear programming to solve the honeypot deployment strategy, achieving dynamic adjustment of honeypot configuration and precise response to attack behavior. Simultaneously, by calculating quantitative indicators such as attack detection rate, honeypot failure rate, protected system value, and system stability, and feeding these quantitative indicators into the model optimization process, continuous improvement of the defense strategy is achieved. Furthermore, this method also covers the identification and response to high-risk attack patterns, effectively improving the intelligence level of network security protection, enhancing the ability to respond to complex network attacks, and ensuring maximum defense effectiveness with limited resources.
[0066] Example 3 illustrates a schematic scheme for a dynamic honeycomb defense strategy generation method based on Stackelberg game theory. It should be noted that the technical solution of this system for generating dynamic honeycomb defense strategies based on Stackelberg game theory is based on the same concept as the aforementioned method for generating dynamic honeycomb defense strategies based on Stackelberg game theory. Details not described in detail in the system for generating dynamic honeycomb defense strategies based on Stackelberg game theory in this embodiment can be found in the description of the aforementioned method for generating dynamic honeycomb defense strategies based on Stackelberg game theory.
[0067] This embodiment also provides a dynamic honeycomb defense strategy generation system based on Stackelberg game theory, including: The data collection module is used to collect attack behavior data in real time through various types of honeypot devices deployed in the network; The change module is used to automatically change the honey array when the attack behavior data reaches a preset trigger mechanism; A solution module is constructed to build a target game model between the defender and the attacker based on a multi-round Stackelberg game architecture, and to solve the honeypot deployment strategy through Markov decision process and mixed integer linear programming. The adjustment module is used to call the honeypoint repository resources based on the honeypoint deployment strategy and automatically perform dynamic adjustments to the honeypoint configuration; The calculation module is used to calculate the target quantitative index based on the adjusted honeypot configuration and the attack behavior data; An optimization module is used to optimize the target game model based on the calculation results of the target quantitative index.
[0068] This embodiment also provides an electronic device applicable to the generation of dynamic honeycomb defense strategies based on Stackelberg game theory, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for generating dynamic honeycomb defense strategies based on Stackelberg game theory as proposed in the above embodiment.
[0069] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the method for generating a dynamic honeycomb defense strategy based on Stackelberg game theory as proposed in the above embodiments. The storage medium proposed in this embodiment and the method for generating a dynamic honeycomb defense strategy based on Stackelberg game theory proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0070] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0071] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for generating dynamic honeycomb defense strategies based on Stackelberg game theory, characterized in that: include, Attack behavior data is collected in real time through various types of honeypot devices deployed in the network; When the attack behavior data reaches the preset trigger mechanism, the honey array will automatically change; A target game model of defenders and attackers is constructed based on a multi-round Stackelberg game architecture, and the honeypot deployment strategy is solved by Markov decision process and mixed integer linear programming. Based on the honeypoint deployment strategy, the honeypoint repository resources are invoked, and dynamic adjustments to the honeypoint configuration are automatically performed. Based on the adjusted honeypot configuration and the attack behavior data, the target quantitative index is calculated. The target game model is optimized based on the calculation results of the target quantitative index.
2. The method for generating a dynamic honeycomb defense strategy based on Stackelberg game theory as described in claim 1, characterized in that: The objective game model constructed based on the multi-round Stackelberg game architecture includes, Network topology definition, honeypot policy space and reward function definition, and attacker policy space and reward function definition; The network topology definition is to represent the network as a directed graph G = (N, E), where the set of nodes N represents the nodes in the network, the set of edges E represents the connection relationship between the nodes, and the set of entry nodes is determined as the starting point for the attacker's penetration.
3. The method for generating a dynamic honeycomb defense strategy based on Stackelberg game theory as described in claim 2, characterized in that: The quantitative index of the calculation target, include, Calculate the attack detection rate, honeypot rejection rate, value of the protected system, and system stability; among which, The attack detection rate refers to the proportion of attacks that are successfully detected. The honey trigger rate refers to the proportion of times an attacker triggers a honey spot.
4. The method for generating a dynamic honeycomb defense strategy based on Stackelberg game theory as described in claim 3, characterized in that: The preset triggering mechanism includes a time-slice triggering mechanism and a system event triggering mechanism; wherein... The time slice triggering mechanism refers to triggering honey array transformation when a preset threshold number of game iterations is reached. The system event triggering mechanism refers to triggering honey array transformation when the frequency of honey stepping exceeds the dynamic threshold or a high-risk attack mode is identified.
5. The method for generating a dynamic honeycomb defense strategy based on Stackelberg game theory as described in claim 3, characterized in that: The attack detection rate is calculated as the ratio of the number of times the honeypot set is triggered to the total number of attack events. The formula for calculating the honey-triggering rate is the ratio of the number of paths that trigger honey points to the total number of paths; The value of the protected system is calculated based on the sum of the values of high-value hosts that have not been attacked and the value of attack behaviors that are honeypot traps; The system stability is measured by the degree of difference between the honeypot deployment strategy in the k-th iteration and the average strategy.
6. The method for generating a dynamic honeycomb defense strategy based on Stackelberg game theory as described in claim 4, characterized in that: The high-risk attack modes include at least one of SQL injection, brute force, and lateral movement.
7. A method for generating dynamic honeycomb defense strategies based on Stackelberg game theory as described in any one of claims 1-6, characterized in that: The optimization of the target game model based on the calculation results of the target quantitative index includes, The target quantitative index is input into the target game model, and the decision parameters are updated through a strategy iteration algorithm.
8. A dynamic honeycomb defense strategy generation system based on Stackelberg game theory, employing the method described in any one of claims 1-7, characterized in that, include: The data collection module is used to collect attack behavior data in real time through various types of honeypot devices deployed in the network; The change module is used to automatically change the honey array when the attack behavior data reaches a preset trigger mechanism; A solution module is constructed to build a target game model between the defender and the attacker based on a multi-round Stackelberg game architecture, and to solve the honeypot deployment strategy through Markov decision process and mixed integer linear programming. The adjustment module is used to call the honeypoint repository resources based on the honeypoint deployment strategy and automatically perform dynamic adjustments to the honeypoint configuration; The calculation module is used to calculate the target quantitative index based on the adjusted honeypot configuration and the attack behavior data; An optimization module is used to optimize the target game model based on the calculation results of the target quantitative index.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Honeypot deployment method, device and equipment based on attack and defense income, medium and product
CN120034354A
Optimal strategies in security games
WO2013176784A1