Iot deception trapping strategy evaluation method and system based on ddpg security game
By constructing a network attack and defense security game model and the Deep Deterministic Policy Gradient Algorithm (DDPG) in the Internet of Things (IoT) environment, the problem of the lag in the evaluation of deception defense strategies in the IoT environment is solved, and the effective evaluation and rapid optimization of deception defense strategies are realized, thereby improving the security and defense effectiveness of the IoT environment.
Patent Information
- Application Number
- CN202410647742.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-23
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2044-05-23
AI Technical Summary
Existing technologies are insufficient to effectively evaluate the merits of deception defense strategies in the Internet of Things (IoT) environment, and cannot quickly iterate and evaluate the effectiveness of deception defense implementation, resulting in lag and waste of resources in defense methods.
This paper proposes an evaluation method for IoT deception and trapping strategies based on DDPG security game theory. By constructing a network attack and defense security game model, the paper quantifies the gains of both the attacker and defender using CVE vulnerabilities, and optimizes the deception asset deployment strategy using the Deep Deterministic Policy Gradient Algorithm (DDPG), thereby achieving an effective evaluation of deception defense strategies.
It enables accurate evaluation of deception defense strategies, improves the security and defense effectiveness of IoT environments, and allows for rapid iteration and optimization of defense strategies in real-world network environments, reducing resource consumption.
Smart Images

Figure CN118400169B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Internet of Things (IoT) security technology, and in particular to an evaluation method and system for IoT deception and trapping strategies based on DDPG security game theory. Background Technology
[0002] In recent years, the continuous iterative development of information and communication technologies has facilitated the emergence of the Internet of Things (IoT). Due to its ability to sense, transmit, and process intelligent information, it is now widely used in hospitals, transportation, power grids, and other fields, resulting in an exponential growth in the number of devices connected to the IoT environment. However, the increasing complexity, diversity, and heterogeneity of these connected devices have broadened the attack surface of the IoT environment, leading to more complex and diverse security threats. Simply deploying passive defense methods such as firewalls and intrusion detection systems to address IoT security threats incurs significant performance overhead, and these methods are also lagging, unable to defend against attacks targeting unknown vulnerabilities in the IoT environment. Therefore, there is an urgent need to design a proactive security protection technology for the IoT environment. Deception defense technology, as a proactive defense technology, dynamically deploys deception resources such as honeypots and honeynets in the network environment to attract attacks. When malicious attackers attack these deception resources, the resources can detect the attack and take effective measures for tracing and countermeasures. Currently, various deception technologies are used in IoT security protection, such as IoTPOT, IoTCMal, and U-pot.
[0003] Evaluation of deception effectiveness in the Internet of Things (IoT) environment is a crucial issue. By assessing the effectiveness of deception defense facilities and optimizing deception strategies, the practical significance of deception defense implementation can be significantly enhanced. Currently, methods for evaluating deception effectiveness mainly fall into three categories: Real-world experiments, which involve deploying the designed deception defense system in a live network environment and verifying its effectiveness through red-blue team exercises; Simulated environment experiments, which use simulation tools to simulate network attack and defense behaviors and evaluate the effectiveness of deception defense strategies; Game theory-based evaluation, where the effects of attack and defense strategies are quantified by the payoffs, and the quality of deception defense strategies can be assessed by appropriately setting game payoffs and solving for Nash equilibrium; and Probabilistic model-based evaluation, which uses probabilistic models to analyze indicators such as attack success rate and system security protection time to evaluate the effectiveness of deception defense strategies. Real-world deception effectiveness evaluation requires conducting real red-blue team exercises, which incurs significant resource costs. Therefore, simulated environment-based evaluation is a good alternative. While it cannot practically evaluate the merits of deception strategies, it can assess the level of deception defense to some extent. However, both game-theoretic and probabilistic model-based methods for evaluating deception effectiveness only assess the deception effect at the theoretical level, thus limiting their usability. Therefore, there is an urgent need for an effective solution to evaluate the merits of deception defense strategies to meet the application needs of the current Internet of Things (IoT) environment. Summary of the Invention
[0004] To this end, the present invention provides an evaluation method and system for IoT deception and trapping strategies based on DDPG security game theory. By quantifying the gains of both the attacker and defender through vulnerability scoring, it is more in line with the actual gains of both the attacker and defender in the process of attack and defense confrontation, and can improve the evaluation effect of security strategies from the perspective of game gains.
[0005] According to the design scheme provided by this invention, on the one hand, a method for evaluating IoT deception and trapping strategies based on DDPG security game theory is provided, comprising:
[0006] A network attack and defense security game model is constructed based on attack strategies and security strategies in the Internet of Things environment. The attack strategy includes the CVE vulnerabilities used by the attacker against the network host, and the security strategy includes the number and / or location of deceptive assets deployed by the defender to prevent the attacker from exploiting the CVE vulnerabilities.
[0007] By utilizing CVE vulnerabilities, the payoffs of both attackers and defenders in a network attack and defense security game model are quantified, and the payoff functions of both sides are obtained.
[0008] Based on the payoff functions of both attackers and defenders, a security strategy optimization problem is constructed. The Deep Deterministic Policy Gradient Algorithm (DDPG) is used to solve the security strategy optimization problem, and the optimal deception asset deployment strategy is obtained based on the solution results.
[0009] As an evaluation method for IoT deception and trapping strategies based on DDPG security game theory in this invention, a network attack and defense security game model is further constructed based on attack strategies and security strategies in the IoT environment, including:
[0010] Information on known CVE vulnerabilities in the Internet of Things (IoT) environment is obtained through network probing and scanning, and a set of attacker attack strategies and attack behaviors is constructed.
[0011] Acquire deceptive assets deployed to detect attackers' offensive activities, and construct a set of defender security policies based on the information on deployed deceptive assets;
[0012] The network state of the IoT environment can be represented by deceiving the location of asset deployments and by attackers exploiting CVE vulnerabilities to attack hosts.
[0013] This model represents a network attack and defense security game based on a set of attack strategies, a set of security strategies, and network states.
[0014] As an evaluation method for IoT deception and trapping strategies based on DDPG security game theory in this invention, the method further quantifies the gains of both attackers and defenders in the network attack and defense security game model by utilizing CVE vulnerabilities, including:
[0015] CVE vulnerabilities are quantified using a base score, which includes a score representing the impact of an attacker's attack on the integrity, confidentiality, and availability of the network system, and an exploitability score representing the complexity of an attacker's exploitation of the CVE vulnerability.
[0016] Construct payoff functions for both the attacker and defender using the influence score and the available score.
[0017] As an evaluation method for IoT deception and trapping strategies based on DDPG security game theory in this invention, the attacker's payoff function is further expressed as: The defender's payoff function is expressed as: Among them, ES a BS is the exploitability score for CVE vulnerabilities in an attack. a This represents the base score for CVE vulnerabilities in the attack. These represent the attack gains when the attack is captured by the deceitful asset and the attack gains when the attack is not captured by the deceitful asset, respectively. a For defensive action overhead, IS a The impact of CVE vulnerabilities on the score during defensive actions. These represent the defensive gains when the deceptive asset can capture the attack, and the defensive gains when the deceptive asset cannot capture the attack, respectively.
[0018] As an evaluation method for IoT deception and trapping strategies based on DDPG security game theory in this invention, the security strategy optimization problem constructed according to the payoff functions of both attackers and defenders is further expressed as:
[0019] Where a is the attack strategy, A is the set of attack strategies, and p a p represents the probability that attack strategy a is captured by any one of the k deployed deception assets. a,k Let a represent the probability that attack strategy a is captured by one of the k deceptive assets, specifically the deterministic deceptive asset. * For deterministic attacks that deceive asset capture, α represents a metric for network system security and performance.
[0020] As a method for evaluating IoT deception and trapping strategies based on DDPG security game theory, this invention further utilizes the Deep Deterministic Policy Gradient Algorithm (DDPG) to solve the security strategy optimization problem, including:
[0021] The current network environment state is input into the Actor network by the intelligent agent. The Actor network outputs behavior according to the current network environment state, and adds noise to the output behavior before inputting it into the network environment. This allows the current network environment to obtain the next network environment state and reward under the common effect of the output behavior and noise, generating a trajectory composed of state, behavior and reward, and storing the trajectory in the experience replay pool.
[0022] The intelligent agent repeatedly executes the trajectory generation process until the number of trajectories in the experience revisit pool reaches a preset threshold.
[0023] A specified number of trajectories are sampled from the experience revisit pool as training samples for the Actor network and the Critic network. The Actor network and the Critic network are trained using the training samples, and the target Actor network and the target Critic network are obtained by updating the network parameters.
[0024] A deception asset optimization model is constructed based on the target Actor network and the target Critic network, and the optimal security strategy is obtained through the deception asset optimization model.
[0025] As an evaluation method for IoT deception and trapping strategies based on DDPG security game theory according to the present invention, it further includes:
[0026] We use simulators to build network scenario instances, verify the effectiveness of the best deception asset deployment strategy through network attack simulation, and display the attack and defense simulation process in the network scenario instances in real time through data visualization.
[0027] Furthermore, this invention also provides an evaluation system for IoT deception and trapping strategies based on DDPG secure game theory, comprising: a model building module, a profit quantification module, and an optimization solution module, wherein...
[0028] The model building module is used to build a network attack and defense security game model based on attack strategies and security strategies in the Internet of Things environment. The attack strategy includes the CVE vulnerabilities used by the attacker against the network host, and the security strategy includes the number and / or location of deceptive assets deployed by the defender to prevent the attacker from exploiting the CVE vulnerabilities.
[0029] The revenue quantification module is used to quantify the revenue of both the attacker and defender in a network attack and defense security game model by exploiting CVE vulnerabilities, and obtain the revenue function of both the attacker and defender.
[0030] The optimization and solution module is used to construct a security strategy optimization problem based on the profit functions of both the attacker and the defender, and to solve the security strategy optimization problem using the Deep Deterministic Policy Gradient Algorithm (DDPG). Based on the solution results, the optimal deception asset deployment strategy is obtained.
[0031] The beneficial effects of this invention are:
[0032] This invention constructs a network attack and defense security game model based on attack and security strategies in the Internet of Things (IoT) environment. It utilizes CVE vulnerabilities to quantify the payoffs of both attackers and defenders in the model, obtaining their payoff functions. Based on these payoff functions, it constructs a security strategy optimization problem and solves it using the Deep Deterministic Policy Gradient Algorithm (DDPG). The optimal deception asset deployment strategy is obtained based on the solution, effectively evaluating the effectiveness of deception defense implementation strategies. The advancement and effectiveness of the selected strategy are verified through the NASim-DEE simulation platform. This invention can be deployed and implemented in the IoT environment, demonstrating promising application prospects. Attached image description:
[0033] Figure 1 This is a schematic diagram of the evaluation process for the IoT deception and trapping strategy based on DDPG security game in the embodiment;
[0034] Figure 2 This is a schematic diagram of an attack and defense scenario in an IoT environment, as illustrated in the example.
[0035] Figure 3 This is a schematic diagram illustrating the principle of the DDPG-based deception effect evaluation algorithm in the embodiment;
[0036] Figure 4 This is a schematic diagram of the network platform deception simulation experiment framework in the embodiment;
[0037] Figure 5 This is a schematic diagram of network scene generation in the embodiment;
[0038] Figure 6 This is a schematic diagram of attack and defense confrontation in the embodiment;
[0039] Figure 7 This is a visualization of data acquisition and analysis in the example.
[0040] Figure 8 This example illustrates algorithm convergence under different Actor network learning rates.
[0041] Figure 9 This example illustrates the convergence of attacker strategy acquisition under different parameters 'a'.
[0042] Figure 10 This example illustrates the convergence of the defender strategy under different parameters 'a' in the embodiment. Detailed implementation method:
[0043] To make the objectives, technical solutions, and advantages of this invention clearer and more understandable, the invention will be further described in detail below with reference to the accompanying drawings and technical solutions.
[0044] To address the shortcomings of existing game theory models for evaluating the effectiveness of deception defense in the Internet of Things (IoT) environment—namely, the inability to accurately assess the merits of deception defense implementation strategies and the inability to rapidly iterate and evaluate the effectiveness of deception defense implementation—this invention provides an embodiment, see [link to embodiment]. Figure 1 As shown, a method for evaluating IoT deception and trapping strategies based on DDPG security game theory is provided, comprising:
[0045] S101. Construct a network attack and defense security game model based on attack strategies and security strategies in the Internet of Things environment. The attack strategy includes the CVE vulnerabilities used by the attacker against the network host, and the security strategy includes the number and / or location of deceptive assets deployed by the defender to prevent the attacker from exploiting the CVE vulnerabilities.
[0046] Specifically, a network attack and defense security game model, constructed based on attack and security strategies in the Internet of Things (IoT) environment, can be designed to include:
[0047] Information on known CVE vulnerabilities in the Internet of Things (IoT) environment is obtained through network probing and scanning, and a set of attacker attack strategies and attack behaviors is constructed.
[0048] Acquire deceptive assets deployed to detect attackers' offensive activities, and construct a set of defender security policies based on the information on deployed deceptive assets;
[0049] The network state of the IoT environment can be represented by deceiving the location of asset deployments and by attackers exploiting CVE vulnerabilities to attack hosts.
[0050] This model represents a network attack and defense security game based on a set of attack strategies, a set of security strategies, and network states.
[0051] like Figure 2 As shown, in an IoT network environment, IoT access devices connect to cloud servers via SDN networks. The cloud servers store data generated by these devices and perform related calculations. Therefore, cloud servers contain important asset information, making them targets for attackers. Security administrators can use SDN controllers to schedule traffic and create deceptive assets, thereby enhancing network traffic resilience.
[0052] Attackers first infect access devices, then use network probing to obtain information about the cloud server's operating system and application versions. Based on this information, they identify relevant vulnerabilities and exploit them to gain privilege escalation and server access. If a defender deploys a deceptive asset near the attacker's target server, this asset can detect the attacker's behavior through mirrored traffic and take appropriate measures to prevent the attacker from gaining server access. Therefore, the attacker's goal is to exploit vulnerabilities to escalate user privileges while preventing the deceptive asset from detecting its actions, while the defense's goal is to prevent the attacker from successfully exploiting the vulnerabilities.
[0053] against Figure 2 The scenario shown selects n known CVE vulnerabilities to construct an attacker's strategy set. The attacker can exploit these vulnerabilities to perform privilege escaping and gain access to the cloud server. The attacker's actions can be represented as a set A = {a1, a2, a3, ..., a...}. n}, where a1, a2, a3, ..., a n An identifier representing an attack can be represented by <hostname, CVE-ID>.
[0054] In network attack and defense confrontations, defenders incur certain resource overheads when deploying deception assets, including traffic scheduling overhead and overhead incurred by changing the location of deception assets. Furthermore, the deployment of deception assets can impact network performance. Therefore, considering the needs of real-world scenarios, in... Figure 1 In the scenario shown, only k < n deceptive assets can be deployed to prevent attackers from successfully exploiting the vulnerability. The defender's strategy set can be represented as D. k ={D∈A:|D|=k}.
[0055] A Cyber Deception Evaluation Model based on security game (CDEM-SG) is constructed. The defender, as the leader, selects a strategy first, and the attacker, as the follower, selects a strategy based on the defender's strategy. To prevent the attacker from discovering the defender's deployed strategy and thus bypassing the deployed deception assets, the defender's strategy is considered to have mixed probabilities. Where N = (N... A N D () represents the two sides in an offensive and defensive confrontation. N A Indicates the attacker, N D The attacker and defender are represented here, with only one attacker and one defender, i.e., |N| = 2. S represents the current state of the network environment, i.e., the location of the deceptive asset deployment, the host being attacked, and the CVE vulnerability. P = (P A ,P D () represents the strategy set of the attacking and defending sides. P A The attacker's set of policies can be represented as P. A =A = {a1, a2, a3, a4, a5, a6}, which represents the container targeted by the attacker and the CVE vulnerability exploited; P D The set of strategies of the defender can be represented as P. D =D k ={D∈A:|D|=k}, which represents the current locations where the defender deploys k deception assets, in order to enhance the randomness of the deployment of deception assets.
[0056] S102. Quantify the gains of both the attacker and defender in the network attack and defense security game model by using the CVE vulnerability, and obtain the gain function of both the attacker and defender.
[0057] Specifically, quantifying the gains of both attackers and defenders in a network attack-defense security game model using CVE vulnerabilities can be designed to include:
[0058] CVE vulnerabilities are quantified using a base score, which includes a score representing the impact of an attacker's attack on the integrity, confidentiality, and availability of the network system, and an exploitability score representing the complexity of an attacker's exploitation of the CVE vulnerability.
[0059] Construct payoff functions for both the attacker and defender using the influence score and the available score.
[0060] To better reflect the actual gains of both attackers and defenders in a security contest, this embodiment employs the Common Vulnerability Scoring System (CVSS) to evaluate CVE vulnerabilities, thereby guiding the quantification of gains for both sides in the security game. CVSS utilizes the knowledge of global cybersecurity experts to provide a numerical value within the range [0, 10] for each known vulnerability. This value reflects the severity of the vulnerability and the expertise required to exploit it. Therefore, using CVSS to score vulnerabilities and guide the gains for both attackers and defenders can effectively reflect the ease with which attackers can exploit vulnerabilities, the degree of threat they pose to the system, and the gains for defenders in successfully defending against the vulnerability.
[0061] The gains for both attackers and defenders in a security game can be quantified based on the impact score (IS), exploitability score (ES), and base score (BS) of a vulnerability assessed in CVSS. The impact score (IS) represents the impact of an attacker's specific attack on the system's integrity, confidentiality, and availability; the exploitability score (ES) represents the complexity of actually exploiting a specific vulnerability; and the base score (BS) is composed of the impact score (IS) and the exploitability score (ES). Based on these CVSS vulnerability scores, the gains for both attackers and defenders in a security game are defined as follows.
[0062] Since deceptive assets deployed in the network environment can capture attacks, the attack payoff is considered in the following two scenarios: when the attack is captured by the deceptive asset, the attack payoff is defined as follows: The exploitability score of CVE vulnerabilities can be used in ES. a This represents the time and cost required to launch the attack. Since the attack is detected by the deceitful asset, the attack payoff is negative. Conversely, when the attack is not detected by the deceitful asset, the attack payoff is defined as... Due to the base score BS a Taking into account the overhead required for the attack, IS a Impact of vulnerabilities on the system (ES) a Therefore, the attack reward at this time is based on the base score BS for CVE vulnerabilities. a Based on the above analysis, the attack gain can be expressed as:
[0063]
[0064] Similarly, since the number of deceptive assets deployed by the defender is less than the number of CVE vulnerabilities existing in the target network, there is a certain probability that the deceptive assets deployed by the defender will fail to capture the attacker's attack. Regarding the defense benefit, when the deceptive assets fail to capture the attack, the defense benefit is... Define the overhead of the defender deploying deception resources as c. a For any a ∈ A, the cost is proportional to the betweenness centrality of network nodes. The higher the betweenness centrality of the location where the deception asset is deployed, the greater its impact on network communication. Therefore, based on the cost c... a The payoff when a∈A is defined as the defender's inability to capture the attacker. In other words, when the defender fails to capture the attacker, there is a certain cost, and the benefit is negative; when the deceptive asset can capture the attack, the defense benefit is positive. It can be represented as Even if the defender captures the attacker, they still incur certain defense costs, and do not receive additional positive rewards for a successful defense. Based on the above analysis, the defender's payoff can be expressed as:
[0065]
[0066] S103. Construct a security strategy optimization problem based on the profit functions of both the attacker and defender, and use the Deep Deterministic Policy Gradient Algorithm (DDPG) to solve the security strategy optimization problem. Based on the solution results, obtain the optimal deception asset deployment strategy.
[0067] Based on the security game model and the attack and defense payoff function, p is defined. a p represents the probability that attack strategy a is captured by any one of the k deceptive assets deployed. a,k This represents the probability that attack strategy a is captured by one of k deceptive assets, specifically a deterministic deceptive asset, thus yielding the probability that the defender captures a deterministic attack strategy a. * The expected return is Since the overhead of k-1 deceptive assets is not considered here, the expected return of defense is modified to... Similarly, the attack yields... The first term represents the attacker's gain when the attack is successful, and the second term represents the attacker's cost when the attack is unsuccessful. Based on the above analysis, to solve for the optimal deception asset deployment strategy, the optimization problem can be described as follows:
[0068]
[0069] Here, α represents the input parameter for measuring system security and performance. When α = 0, it means that the defender only considers system security and ignores the overhead generated by k deceptive assets. When α = 1, it means that the defender only considers system performance and will not deploy deceptive assets next to cloud servers that affect system performance, even if it affects system security.
[0070] The solution to the security policy optimization problem using the Deep Deterministic Policy Gradient Algorithm (DDPG) can be designed to include:
[0071] The current network environment state is input into the Actor network by the intelligent agent. The Actor network outputs behavior according to the current network environment state, and adds noise to the output behavior before inputting it into the network environment. This allows the current network environment to obtain the next network environment state and reward under the common effect of the output behavior and noise, generating a trajectory composed of state, behavior and reward, and storing the trajectory in the experience replay pool.
[0072] The intelligent agent repeatedly executes the trajectory generation process until the number of trajectories in the experience revisit pool reaches a preset threshold.
[0073] A specified number of trajectories are sampled from the experience revisit pool as training samples for the Actor network and the Critic network. The Actor network and the Critic network are trained using the training samples, and the target Actor network and the target Critic network are obtained by updating the network parameters.
[0074] A deception asset optimization model is constructed based on the target Actor network and the target Critic network, and the optimal security strategy is obtained through the deception asset optimization model.
[0075] To evaluate the effectiveness of deception defense strategies, this embodiment employs the Optimal Cyber Deception Assets Evaluation Method based on Deep Deterministic Policy Gradient Algorithm (DDPG, DAE-DDPG). The algorithm's framework diagram is shown below. Figure 3 As shown. First, the agent sets the state s of the network environment. t The input to the Actor network is the reconnaissance strategy of the attacker at degree k in the current network environment. The Actor network then outputs behavior a based on the current network environment state. t And with the addition of noise N, the behavior a t +N is input into the environment, and the current environment is in behavior a. t The next state s' is obtained under the action of +N, and the reward r is obtained. At this time, a state-behavior-reward pair (s,a,r,s') is formed and stored in the experience replay pool. The above process is repeated until the number of trajectories in the experience replay pool reaches a certain threshold. At this time, the trajectories are sampled to update the network parameters.
[0076] For updating the parameters of the Actor network, the state s in the trajectory is first input into the Actor network, and the Actor network outputs action a. Then, action a and the state s in the trajectory are simultaneously input into the Critic network to obtain the corresponding Q value. The above process is repeated until the sampled trajectories can output the corresponding Q value. At this time, the parameters θ of the Actor network are updated according to formula (4).
[0077]
[0078] For parameter updates of the Critic network, the state s' in the trajectory is first input into the Target Actor network, which outputs action a. Then, action a and the state s' in the trajectory are input into the Target Critic network, which outputs the corresponding Q' value. The Q' value is then summed with the r in the trajectory corresponding to action a and state s' according to formula (5) to obtain Q. target Compare the Q-value output by the current Critic network with Q... target Subtraction yields Repeat the above process until all sampled trajectories have been traversed, and the loss function of the Critic network is obtained as shown in formula (6). Based on this loss function, the parameters of the Critic network can be adjusted. Update.
[0079] Q target =r i +γQ'(s i+1 ,π θ' (s i+1 (5)
[0080]
[0081] For updating the parameters of the target Actor network and the Critic network, after a certain threshold, a small-step update is performed on the target Actor network and the Critic network according to the soft update mechanism using the formula shown below. The specific algorithm flow is shown in Algorithm 1.
[0082]
[0083]
[0084]
[0085] Furthermore, in this embodiment, a simulator can be used to construct network scenario instances, verify the effectiveness of the optimal deception asset deployment strategy through network attack simulation, and display the attack and defense simulation process in the network scenario instances in real time through data visualization.
[0086] Network Attack Simulator (NASim) is a tool for simulating attacks on real computer networks using fast, parallelizable memory abstraction. It allows for easy creation of network environments and simulation of adversarial behavior between attackers and defenders. Users can use NASim's library functions to build network scenarios with different subnets, topologies, host configurations, and firewalls as needed. Based on these created network scenario instances, attackers can customize their actions, including subnet scanning, service scanning, operating system scanning, vulnerability scanning, process scanning, vulnerability exploitation, and privilege escalation. Attackers use these actions to penetrate the constructed network scenario and gain root privileges on critical hosts, incurring corresponding overhead.
[0087] In this solution, the Deceptionevaluation and Optimization Simulation Experimental Environment Based on Network Attack Simulation Platform (NASim-DEE) comprehensively considers the simulation and analysis of the attack and defense process, network environment generation, and visualization of the attack and defense process. Its framework is as follows: Figure 4 As shown, the system mainly consists of four parts: network environment generation, policy learning and generation, attack and defense simulation, and data visualization. Network environment generation provides a user-friendly UI interface, allowing users to easily create network environments and configure host vulnerabilities and network attack behaviors to generate attack tasks. Once the network scenario is successfully configured, policy learning and generation are trained based on the network scenario to adaptively generate the number and placement of deceptive assets in the network topology. Attack and defense simulation simulates the attack and defense process based on the configured network scenario, attack tasks, and the optimal deceptive asset deployment strategy obtained through policy learning and generation. This simulates the probability of an attacker successfully attacking a vulnerable host and the effectiveness of the deception defense. During the attack and defense simulation, the generated simulated data is stored in a database and sent to the data visualization in real time. The data visualization analyzes the data to display the attack and defense process in real time and outputs the attack success rate of multiple attack and defense confrontations after the simulation experiment.
[0088] Network scenario generation mainly includes a scenario building module and an attack task building module. Users can configure network scenarios, such as network topology, attack behaviors, vulnerability types and numbers, through a GUI interface, and then configure this network configuration into the network attack and defense simulation process via the REST API interface. For the scenario building module, the method is consistent with the NASim simulation tool for creating network scenarios, including three methods: For the `make_benchmark` method, users can directly drag and drop to generate a benchmark network scenario through the GUI interface. This method does not require configuring other network parameters; after dragging and dropping, all the configurations for network attack and defense simulation are already available. For the `load.yaml` method, users configure information such as subnet, topology, firewall, and host configurations through the GUI interface. This configuration is then written to the `host.yaml` file via the REST API, allowing the attack and defense simulation to load this configuration for simulation. For the `generate` method, users can configure parameters such as the number of hosts, host operating systems, and attack overhead required during the attack and defense simulation through the GUI interface. After configuration, the parameters of the `generate` function are assigned via the API interface to the NASim simulation platform. The attack task construction module primarily handles the configuration and initialization of attack behaviors. It evaluates and constructs attacker strategies through strategy selection and initializes the selected strategies. The program flow of the scenario construction module and the attack task construction module is as follows: Figure 5 As shown.
[0089] In the NASim deception simulation environment, the deception effectiveness of deception defense technologies is evaluated through continuous iteration of the attack and defense confrontation process. The simulation of the attack and defense confrontation is completed through continuous interaction between the attack simulation module and the deception deployment and transformation module and the network scenario. The entire simulation process is consistent with the reinforcement learning training process, that is, firstly, the state of the network scenario is acquired, then the attacker takes an attack behavior and applies this behavior to the network environment, at which point the network scenario enters a new state, and the above process is continuously iterated until the end, and the probability of the attacker's successful attack is obtained to evaluate the deception effectiveness. The attack simulation module launches an attack in a network scenario based on the attack task configuration information. A simulation ends when the attacker successfully exploits vulnerabilities and escalates privileges on any host within a certain number of steps. Conversely, a simulation also ends if the attacker fails to exploit vulnerabilities or escalate privileges on any host within a certain number of steps, or if they exploit vulnerabilities in a deceptive asset within a certain number of steps. The deception deployment module selects host locations in the network topology created by the NASim platform to deploy sensitive hosts for deception defense based on the deception deployment strategy trained and generated. If an attacker exploits vulnerabilities on a sensitive host during the attack, the attacker's behavior is considered captured and the attack fails. The deception asset transformation module changes the deployment location of sensitive hosts according to the deception deployment strategy trained and generated, preventing attackers from discovering the location of sensitive hosts through multiple attacks. The program flow for the attack and defense simulation is as follows: Figure 6 As shown.
[0090] Attackers can be categorized into cautious attackers, standard attackers, and offensive attackers.
[0091] Cautious attackers: Before attacking hosts in the network environment, these attackers first scan the network, then analyze the probability of a successful attack, and target hosts with a high success rate. The specific attack stages of a cautious attacker can be divided as follows:
[0092] (1) Attackers first perform a comprehensive scan of the network environment, including subnet scanning, port scanning, vulnerability scanning and operating system scanning;
[0093] (2) Attackers analyze the scan results and select the host with the highest probability of success to launch an attack, that is, select the best vulnerability to obtain user privileges.
[0094] (3) After the attacker successfully obtains user privileges, the attacker will perform a process scan to obtain root privileges on the host.
[0095] Standard attackers: These attackers do not scan all hosts in the network environment, but instead use vertical scanning to target specific hosts and then exploit vulnerabilities in those hosts. This type of attacker is more consistent with real-world network attackers. The specific attack stages are as follows:
[0096] (1) The attacker first performs a subnet scan to obtain active addresses in the network environment;
[0097] (2) The attacker randomly selects a host to perform service scanning, operating system scanning, vulnerability scanning and process scanning;
[0098] (3) Depending on the host's access permissions, the attacker can exploit vulnerabilities and escalate privileges based on the vulnerabilities obtained from the scan.
[0099] (4) When the attacker completes the privilege escalation according to the third phase, the attack enters the second phase until all hosts are successfully compromised; otherwise, the third phase continues.
[0100] Offensive attackers: These attackers do not perform service scanning, operating system scanning, vulnerability scanning, or process scanning. Instead, they attempt to exploit the same vulnerabilities to launch attacks. The specific attack stages are as follows:
[0101] (1) The attacker first performs a subnet scan;
[0102] (2) The attacker randomly selects an action from the available vulnerabilities and privilege escalation actions to launch the attack;
[0103] The attacker randomly selects hosts to launch the second phase of the attack until all hosts have been compromised.
[0104] Suppose that the defender conceals the real assets by deploying deceptive assets in the network environment, luring attackers to attack these deceptive assets and thus capturing their behavior, reducing the probability of a successful attack. To evaluate the effectiveness of the deceptive assets, a sensitive host is introduced to replace the deceptive assets in the NASim simulation environment. When an attacker launches an attack in the constructed network environment, they will indiscriminately scan and exploit vulnerabilities on the hosts deployed in NASim. Once the attacker penetrates the sensitive host, i.e., the attacker compromises the deceptive assets, their behavior is captured by the deceptive assets, and the attack fails.
[0105] Based on the above analysis, the NASim simulation platform can be used to verify the merits of the selected deception strategy. When the deep reinforcement learning algorithm is used to output the best deception strategy deployed in the network environment in real time, the number and location of the deception assets determined by the NASim platform in the network scenario can be deployed to evaluate the effectiveness of the deception decision.
[0106] like Figure 7As shown, data visualization mainly includes a data acquisition and analysis module and a visualization module. The data acquisition and analysis module primarily collects and stores data generated during the attack and defense confrontation process, performs further analysis on the stored data, and inputs the analyzed data into the visualization module for display. To output the confrontation process between the attacker and defender in real time, this experimental environment utilizes the subscription and push function of the Redis database to process the data generated during the attack and defense confrontation process in real time, including network topology information, the attacker's location, and the actions taken by the attacker. During the experimental simulation, when the attacker's attack behavior changes, data is pushed to the Redis database in real time, and the Redis database pushes the stored data to the subscribed front-end GUI display interface. During the simulation experiment, the attack and defense confrontation process iterates multiple times. Attackers exhibit two behaviors during the confrontation: either failing to defraud assets or successfully exploiting vulnerabilities in the real host. Therefore, the data after each iteration can be collected and analyzed to further evaluate the effectiveness of the deception defense. After the attack and defense simulation experiment ends, the attacker's attack success rate, attack success time, and attack steps are stored in the Redis database. After the experiment, the stored information is analyzed and displayed.
[0107] Furthermore, based on the above method, this embodiment of the invention also provides an evaluation system for IoT deception and trapping strategies based on DDPG secure game theory, comprising: a model building module, a profit quantification module, and an optimization solution module, wherein,
[0108] The model building module is used to build a network attack and defense security game model based on attack strategies and security strategies in the Internet of Things environment. The attack strategy includes the CVE vulnerabilities used by the attacker against the network host, and the security strategy includes the number and / or location of deceptive assets deployed by the defender to prevent the attacker from exploiting the CVE vulnerabilities.
[0109] The revenue quantification module is used to quantify the revenue of both the attacker and defender in a network attack and defense security game model by exploiting CVE vulnerabilities, and obtain the revenue function of both the attacker and defender.
[0110] The optimization and solution module is used to construct a security strategy optimization problem based on the profit functions of both the attacker and the defender, and to solve the security strategy optimization problem using the Deep Deterministic Policy Gradient Algorithm (DDPG). Based on the solution results, the optimal deception asset deployment strategy is obtained.
[0111] To verify the effectiveness of this solution, the following explanation is based on experimental data:
[0112] Based on such Figure 2 The experimental environment shown is used to evaluate deception strategies. Based on the CVE vulnerabilities listed in Table 1... Figure 2The experiment is conducted in the experimental environment shown. During the experiment, the attacker attacks any CVE vulnerability, and the defender defends against the attacker's attack by deploying a honeypot. When the attacker attacks the honeypot, the attack is considered to have failed.
[0113] Table 1 Attacker Strategies
[0114]
[0115] To determine the optimal strategy for deceiving asset deployment, five CVE vulnerabilities were selected, and CVSS was used to calculate the impact score (IS), exploit score (ES), and base score (BS) for each CVE vulnerability, as shown in Table 2.
[0116] Table 2 Indicators for different CVE vulnerabilities
[0117]
[0118]
[0119] Based on the influence score IS, usable score ES, and basic score BS obtained above, the payoff matrix for both the attacker and defender can be further obtained as shown in Table 3:
[0120] Table 3: Gains for both attacking and defending sides
[0121]
[0122] Based on the profit values of both the attacker and defender, the experimental parameters of the DAE-DDPG algorithm are set as shown in Table 4 to obtain the optimal deception asset deployment strategy:
[0123] Table 4 Experimental Parameter Design
[0124]
[0125] To verify the convergence of the DAE-DDPG algorithm in this scheme and the impact of the Actor network learning rate on the algorithm, a 1×10⁻⁶ test was performed on the DAE-DDPG algorithm. 6 During training, the learning rate of the Actor network was selected as 1×10⁻⁶. -4 2×10 -4 3×10 -4 4×10 -4 5×10 -4 Training was conducted, and the training results were as follows: Figure 8 As shown in the figure. The experimental results show that the DAE-DDPG algorithm can converge under different Actor network learning rates, effectively verifying the algorithm's stability. Furthermore, the experimental results also show that when the Actor network learning rate is 3×10... -4At this point, the DAE-DDPG algorithm reaches its maximum benefit; therefore, in this proposed solution, a learning rate of 3×10⁻⁶ can be selected. -4 .
[0126] Based on the selected learning rate of the Actor network, and assuming the convergence of the DAE-DDPG algorithm, the strategies of the defender and attacker are analyzed under different input parameters α. The experimental results are as follows: Figure 9 and 10 As shown in the figure. The experimental results show that the strategies obtained by the DAE-DDPG algorithm for both attackers and defenders converge, verifying the effectiveness of the algorithm in solving the optimal deception asset deployment strategy. From the attacker's strategy shown, it can be seen that the attacker's optimal attack strategy is to exploit vulnerabilities CVE-2011-2523 and CVE-2022-0492. Figure 10 The defender's strategy shown illustrates that the probability of deploying deceptive assets near five different CVE vulnerabilities varies depending on the performance and security impact factor α. When α is small, the defender deploys deceptive assets near hosts with high CVE vulnerability scores, without considering the performance overhead of deploying deceptive assets. As α increases, the performance impact of deploying deceptive assets is gradually taken into account, thus sacrificing some security to balance the performance impact.
[0127] Based on the above experimental data, it can be further explained that the proposed solution quantifies the benefits of attack and defense by vulnerability scoring and introduces influencing factors for measuring security performance. It uses the DDPG algorithm to generate deceptive asset deployment strategies, which can reasonably predict and select IoT security strategies. This approach is more in line with actual network confrontation scenarios and can effectively improve the security performance of the IoT environment, showing good application prospects.
[0128] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of the invention.
[0129] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0130] The units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations are not considered to be beyond the scope of this invention.
[0131] Those skilled in the art will understand that all or part of the steps in the above methods can be implemented by a program instructing related hardware, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk. Optionally, all or part of the steps in the above embodiments can also be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiments can be implemented in hardware or as a software functional module. This invention is not limited to any particular combination of hardware and software.
[0132] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An Internet of Things fraud luring strategy evaluation method based on DDPG security game, characterized in that, Comprise: According to the attack strategy and security strategy under the Internet of Things environment, a network attack and defense security game model is constructed, wherein the attack strategy comprises the CVE vulnerability adopted by the attacker against the network host, and the security strategy comprises the number and / or location of the deception assets deployed by the defender to prevent the CVE vulnerability attack of the attacker; The benefits of the attack and defense sides in the network attack and defense security game model are quantified by using CVE vulnerabilities, and the benefit functions of the attack and defense sides are obtained, wherein the benefits of the attack and defense sides in the network attack and defense security game model are quantified by using CVE vulnerabilities, including: quantifying CVE vulnerabilities by using basic scores, wherein the basic scores include impact scores representing the influence of attack behavior on the integrity, confidentiality and availability of a network system, and availability scores representing the complexity of the attack behavior using CVE vulnerabilities; constructing the benefit functions of the attack and defense sides by using the impact scores and the availability scores; and the benefit function of the attack side is represented as: The benefit function of the defense side is represented as: ES a BS is the availability score of the CVE vulnerability in the attack behavior a is the basic score of the CVE vulnerability in the attack behavior respectively represent the attack benefit when the attack behavior is captured by a deception asset, and the attack benefit when the attack behavior is not captured by a deception asset a IS is the defense behavior overhead a is the impact score of the CVE vulnerability in the defense behavior respectively represent the defense benefit when the attack behavior can be captured by a deception asset, and the defense benefit when the attack behavior cannot be captured by a deception asset According to the attack and defense side benefit function, a security strategy optimization problem is constructed, and a deep deterministic policy gradient algorithm DDPG is used to solve the security strategy optimization problem, and the best deception asset deployment strategy is obtained according to the solving result, wherein the security strategy optimization problem constructed according to the attack and defense side benefit function is represented as: a is an attack strategy, A is an attack strategy set, p a represents the probability that any one of the k deception assets deployed by the attack strategy a is captured, p a,k represents the probability that the attack strategy a is captured by a certain deception asset in the k deception assets, a * is a deterministic attack of deception asset capture, and alpha represents a measurement parameter of network system security and performance.
2. The DDPG security game based Internet of Things deception decoy strategy evaluation method according to claim 1, characterized in that, According to the attack strategy and security strategy under the Internet of Things environment, a network attack and defense security game model is constructed, comprising: Obtain the known CVE vulnerability information in the Internet of Things environment through network detection scanning, and construct the attack strategy set and attack behavior set of the attacker; Obtain the deception assets deployed for detecting the attack behavior of the attacker, and construct the security strategy set of the defender according to the deployment information of the deception assets; The deployment location of the deception assets and the behavior of the attacker attacking the host using the CVE vulnerability represent the network state of the Internet of Things environment; Based on the attack strategy set, the security strategy set and the network state, the network attack and defense security game model is represented.
3. The DDPG security game based Internet of Things deception decoy strategy evaluation method according to claim 1, characterized in that, Solve the security strategy optimization problem by using the deep deterministic policy gradient algorithm DDPG, comprising: The agent inputs the current network environment state into the Actor network, the Actor network outputs the behavior according to the current network environment state, and adds noise to the output behavior and inputs it into the network environment, so that the current network environment obtains the next network environment state and the income under the common action of the output behavior and the noise, generates a trajectory composed of state, behavior and income, and stores the trajectory in the experience replay pool; The agent repeatedly executes the trajectory generation process until the number of trajectories in the experience replay pool reaches a preset threshold; Sample a specified number of trajectories from the experience replay pool as training samples of the Actor network and the Critic network, train the Actor network and the Critic network using the training samples, and obtain the target Actor network and the target Critic network through network parameter updating; Based on the target Actor network and the target Critic network, a deception asset optimization model is constructed, and the optimal security strategy is obtained through the deception asset optimization model.
4. The DDPG security game based Internet of Things deception decoy strategy evaluation method according to claim 1, characterized in that, Further comprising: Use the simulator to construct a network scenario instance, verify the effectiveness of the best deception asset deployment strategy through network attack simulation, and real-time display the attack and defense simulation process in the network scenario instance through data visualization.
5. An Internet of Things deception luring strategy evaluation system based on DDPG security game, characterized in that, Comprise: model construction module, income quantification module and optimization solving module, wherein, The model construction module is used to construct a network attack and defense security game model according to the attack strategy and security strategy under the Internet of Things environment, wherein the attack strategy comprises the CVE vulnerability adopted by the attacker against the network host, and the security strategy comprises the number and / or location of the deception assets deployed by the defender to prevent the CVE vulnerability attack of the attacker; The benefit quantification module is configured to quantify benefits of both attack and defense sides in the network attack-defense security game model by using the CVE vulnerability, and obtain benefit functions of both attack and defense sides. The quantification of the benefits of both attack and defense sides in the network attack-defense security game model by using the CVE vulnerability comprises: quantifying the CVE vulnerability by using a basic score, wherein the basic score comprises an impact score representing an impact on integrity, confidentiality and availability of a network system when an attacker carries out an attack behavior, and a vulnerability score representing a vulnerability of the attacker to the complexity of the CVE vulnerability; constructing the benefit functions of both attack and defense sides by using the impact score and the vulnerability score; and the attacker benefit function is represented as: The defender benefit function is represented as: ES a BS is the vulnerability score of the CVE vulnerability in the attack behavior. a is the basic score of the CVE vulnerability in the attack behavior. respectively represent attack benefits when the attack behavior is captured by a deception asset and attack benefits when the attack behavior is not captured by the deception asset, and c a IS is a defense behavior overhead. a is the impact score of the CVE vulnerability in the defense behavior. respectively represent defense benefits when the deception asset can capture the attack behavior and defense benefits when the deception asset cannot capture the attack behavior. The optimization solving module is used to construct a security strategy optimization problem according to the income functions of the attack and defense sides, solve the security strategy optimization problem by using the deep deterministic policy gradient algorithm DDPG, and obtain the best deception asset deployment strategy according to the solving result, wherein the security strategy optimization problem constructed according to the income functions of the attack and defense sides is represented as: a is an attack strategy, A is a set of attack strategies, p a denotes the probability that an attack strategy a is captured by any one of the k deceptive assets deployed, p a,k denotes the probability that an attack strategy a is captured by a deterministic deceptive asset among the k deceptive assets, a * is a deterministic attack for deceptive asset capture, a denotes a measurement parameter of network system security and performance.
6. An electronic device, comprising: Include: at least one processor, and a memory coupled with the at least one processor; wherein the memory stores a computer program, and the computer program is capable of being executed by the at least one processor to implement the method according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is capable of being executed to implement the method according to any one of claims 1-4.
Citation Information
Patent Citations
Behavior imitation training method for air intelligent game
CN113221444A
Network space intelligent game decision-making method and system
CN117196040A