Network attack and defense strategy selection method and system based on intelligent evolutionary game

By combining evolutionary game and regret minimization algorithms, a network offense and defense evolutionary game decision-making model is constructed, which solves the problem that the optimization process of network offense and defense strategies in the existing technology is inconsistent with the actual process, improves the convergence and learning efficiency of strategy selection, provides the optimal defense decision-making method, and enhances the defense capabilities of network security operation and maintenance personnel.

CN116248335BActive Publication Date: 2025-05-09Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211640495.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-20
Publication Date
2025-05-09
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

In the existing technology, there is a problem that the strategic optimization process is inconsistent with the actual network offensive and defense processes in the network offensive and defense games, resulting in a decrease in application value and practical significance. At the same time, the method based on Markov decision-making process has poor strategy convergence and strategy degradation in the high-dimensional continuous action space.

Method used

Combining evolutionary game and regret minimization algorithm, a network offense and defense evolutionary game decision model is constructed, and an offense and defense strategy set is obtained by analyzing network scenario vulnerability information, setting regret values ​​and constructing a strategy probability equation of offense and defense agent based on the regret minimization RM algorithm, and solving a differential equation system to obtain the optimal strategy of both offense and defense.

Benefits of technology

It improves the convergence and learning efficiency of the strategy selection algorithm, analyzes the evolution laws of different strategies of both offense and defense in different states, provides the optimal defense decision-making method, is suitable for high-dimensional continuous action space, and enhances the defense capabilities of network security operation and maintenance personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116248335B_ABST
    Figure CN116248335B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of network security technology, and particularly relates to a network attack and defense strategy selection method and system based on intelligent evolutionary game, which obtains an attack and defense strategy set by analyzing network scenario vulnerability information, builds a network attack and defense evolutionary game decision model in combination with a limited rational game scenario, and obtains the attack and defense benefits of different strategy combinations of the attack and defense parties based on the model; in the attack and defense game process, the regret value is set according to the benefits of the unimplemented strategies of both parties and the benefits of the currently implemented strategies, and the probability equations of the implementation strategies of the attack and defense agents are constructed based on the regret minimization RM algorithm by using the strategy weights and the expected benefit loss of the strategies, and the probability equations of the attack and defense parties are combined to construct a differential equation group for the decision selection of the game process of the attack and defense parties; and the optimal strategy of the attack and defense parties is obtained by solving the differential equation group through evolutionary equilibrium. The present invention combines evolutionary game with the regret minimization algorithm to improve the correctness and practicality of strategy selection in the attack and defense game process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and in particular relates to a method and system for selecting network attack and defense strategies based on intelligent evolutionary game. Background Art

[0002] The current network security situation is becoming increasingly severe. Network attacks are developing in the direction of intelligence, combination and concealment. More and more security incidents have caused great damage to the security of cyberspace. The situation of network attack and defense game confrontation is becoming more and more intense, and network defense is also evolving from passive defense to active defense. However, the asymmetry of network security situation is still particularly significant. For attackers, they have sufficient information, cost and time advantages, and can use the smallest possible cost to cause the greatest possible attack damage; while for defenders, they are tired of dealing with the inherent advantages of attackers and must use the smallest possible cost to obtain the greatest possible defense benefits. Game theory provides a theoretical tool for analyzing decision-making, which has made great achievements in the field of cyberspace security. Research on network attack and defense decisions based on game theory has become a current research hotspot. Therefore, by analyzing network attack and defense behaviors, it can help network security operation and maintenance personnel improve the protection capabilities of network information systems, help network security operation and maintenance personnel control the network security situation, and implement network defense strategies in a timely and scientific manner, thereby reversing the current asymmetric situation of "easy to attack and difficult to defend" in cyberspace security.

[0003] At present, network attack and defense game decision-making has developed into non-completely rational game decision-making. The current mainstream methods are mainly divided into two categories: one is the network attack and defense decision-making method based on evolutionary game, and the other is the network attack and defense decision-making method based on reinforcement learning. The network attack and defense decision-making method based on evolutionary game has been widely used in wireless sensor networks (WSN). This type of method focuses on providing strategy selection guidance for defenders, but most of them are based on replicating dynamic equations to solve the optimal strategy. The strategy optimization process does not match the actual network attack and defense process, which greatly reduces the application value and practical significance. The network attack and defense decision-making method based on reinforcement learning has made great research progress in scenarios such as Internet of Vehicles, cloud environment, smart grid, and self-organizing network, but most of them are based on Markov decision process, based on the expected discount of future benefits, and deterministic strategy selection based on value function. The decision convergence is poor, there will be strategy degradation, and it is not suitable for high-dimensional continuous action space. Summary of the invention

[0004] To this end, the present invention provides a network attack and defense strategy selection method and system based on intelligent evolutionary game, which combines evolutionary game with regret minimization algorithm to solve the limited situations in the actual application of network attack and defense in the prior art.

[0005] According to the design scheme provided by the present invention, a network attack and defense strategy selection method based on intelligent evolutionary game is provided, which includes the following contents:

[0006] By analyzing the vulnerability information of network scenarios to obtain the set of attack and defense strategies, a network attack and defense evolutionary game decision model is constructed in combination with the limited rational game scenario, and the attack and defense benefits of different strategy combinations of the attack and defense parties are obtained based on the model;

[0007] In the process of attack and defense game, the regret value is set according to the benefits of the unimplemented strategies of both parties and the benefits of the currently implemented strategies. The probability equations of the implementation strategies of the attack and defense agents are constructed based on the regret minimization RM algorithm using the strategy weights and the expected benefit loss of the strategies. The probability equations of the attack and defense parties are combined to construct the differential equation group for the decision selection of the game process between the attack and defense parties.

[0008] The optimal strategies for both the attacker and the defender are obtained by solving the evolutionary equilibrium of the differential equations.

[0009] As the network attack and defense strategy selection method based on intelligent evolutionary game in the present invention, further, before obtaining the attack and defense strategy set by analyzing the vulnerability information of the network scenario, it also includes: using a vulnerability scanning tool to obtain the vulnerability information of the network scenario.

[0010] As a network attack and defense strategy selection method based on intelligent evolutionary game in the present invention, further, a network attack and defense evolutionary game decision model constructed in combination with a limited rational game scenario is represented by a quintuple (N, D, π, S, U), wherein N represents a set of attack and defense game participants, D represents an attack and defense game strategy space, π represents an attack and defense game strategy selection probability set, S represents an attack and defense game state set, and U represents an attack and defense game benefit matrix set.

[0011] As a network attack and defense strategy selection method based on intelligent evolutionary game in the present invention, further, the probability equation of each attack and defense agent implementing the strategy is constructed based on the regret minimization RM algorithm using the strategy weight and the strategy expected benefit loss: first, the strategy weight in the attack and defense game is set according to the strategy expected benefit; then, the strategy selection process is modeled as in, The defender's strategy DS in the attack-defense game at time t is j The weight it has, It means that the defender selects the attack and defense game strategy DS at time t j The probability of It represents the attacker's strategy AS in the attack and defense game at time t j The weight it has, Indicates that the attacker selects the attack and defense game strategy AS at time t j probability.

[0012] As a network attack and defense strategy selection method based on intelligent evolutionary game in the present invention, further, the strategy weight value during the attack and defense game set according to the expected return of the strategy is expressed as Among them, λ is the learning ability parameter, The defender implements strategy DS in the attack and defense game at time t-1 j The loss function when The attacker implements strategy AS in the attack and defense game at time t-1 j The loss function when .

[0013] As a network attack and defense strategy selection method based on intelligent evolutionary game of the present invention, further, the loss function of the attacking and defending parties is represented by the difference between the maximum value of all the expected returns of each single strategy of the attacking and defending parties and the expected returns of implementing their respective corresponding strategies at the moment of the attack and defense game.

[0014] As a network attack and defense strategy selection method based on intelligent evolutionary game of the present invention, further, the differential equation group of the decision selection of the game process between the attacking and defending parties is expressed as Among them, A and B represent the profit matrices of the attacking and defending parties respectively, the probability vector p is a vector composed of probability elements selected from all pure attack strategies, and the probability vector q is a vector composed of probability elements selected from all pure defense strategies. i Indicates the selection of attack strategy AS i The probability of i / dt means selection strategy AS i The rate of change of probability with time, (Aq) i Indicates the policy AS i The expected return, p T Aq represents the average benefit of the attack strategy set; q j Indicates the selection of defense strategy DS j The probability of j / dt means selection strategy DS j The rate of change of probability over time, (Bp) j Defense strategy DS j The expected return, q T Bp represents the average benefit of the defense strategy set, λ is the learning ability parameter, and k represents the maximum strategy mark among all single strategy expected benefits.

[0015] As a network attack and defense strategy selection method based on intelligent evolutionary game of the present invention, further, by solving the evolutionary equilibrium of the differential equation group to obtain the optimal strategy for both the attack and defense, the strategy selection probability and the weight of the strategy in the strategy set are updated by learning the regret value, and the optimal strategy is selected according to the updated weight.

[0016] Furthermore, the present invention also provides a network attack and defense strategy selection system based on intelligent evolutionary game, comprising: a model building module, an attack and defense analysis module and an optimal output module, wherein:

[0017] The model building module is used to obtain the attack and defense strategy set by analyzing the vulnerability information of the network scenario, build a network attack and defense evolutionary game decision model in combination with the limited rational game scenario, and obtain the attack and defense benefits of different strategy combinations of the attack and defense parties based on the model;

[0018] The attack and defense analysis module is used to set the regret value according to the benefits of the unimplemented strategies and the current implemented strategies of both parties during the attack and defense game process, and to construct the probability equations of the implementation strategies of the attack and defense agents based on the regret minimization RM algorithm by using the strategy weights and the expected benefit loss of the strategies. The probability equations of the attack and defense parties are combined to construct the differential equation group for the decision-making selection of the game process between the attack and defense parties.

[0019] The optimal output module is used to obtain the optimal strategies of both the attacker and the defender by solving the evolutionary equilibrium of the differential equations.

[0020] Beneficial effects of the present invention:

[0021] The present invention aims at the differences and limitations of cognitive abilities of both sides of network security attack and defense, combines the limited rationality game scenario, constructs a network attack and defense evolutionary game decision model based on the regret minimization RM algorithm, applies evolutionary game theory to describe the attack and defense evolution process, adopts the RM algorithm to optimize the strategy learning mechanism, expands the static analysis in the traditional game into a dynamic evolution process, ensures the randomness and convergence of strategy learning, analyzes the evolutionary laws of different strategies of both sides of the attack and defense in different states, and effectively improves the convergence and learning efficiency of the strategy selection algorithm; finally, the optimal defense decision method is given by solving the evolutionary stable equilibrium, so as to describe the evolutionary trajectory of the optimal strategies of both sides of the attack and defense, and provide decision support for active network defense under moderate security. And further verified by numerical experimental results, the scheme of this case has better superiority than other game decision methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 The following is a schematic diagram of the process of selecting a network attack and defense strategy based on intelligent evolutionary game in an embodiment;

[0023] Figure 2 This is an illustration of an enterprise network scenario in the embodiment;

[0024] Figure 3 This is a diagram showing the network state transformation in the embodiment;

[0025] Figure 4 It is a schematic diagram of the probability change curve of defense strategy selection in each state in the embodiment;

[0026] Figure 5 It is a schematic diagram of the strategy evolution under different initial defense selection probabilities of the defense strategy in state S1 in the embodiment;

[0027] Figure 6 It is a schematic diagram of the strategy evolution under different initial attack selection probabilities of the attack strategy in state S1 in the embodiment;

[0028] Figure 7 This is a schematic diagram of the probability change curve of selecting the optimal defense strategy under different learning capabilities in the embodiment;

[0029] Figure 8 Schematic diagram of the comparison of convergence rates of the game strategy selection methods in the embodiments. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention is further described in detail below in conjunction with the accompanying drawings and technical solutions.

[0031] Evolutionary game is a game theory for bounded rational players, which can effectively model the bounded rational attack and defense confrontation process. Reinforcement learning is a complex task, and the impact of a decision made by an agent may depend on the decisions made by other agents in the system. Therefore, evolutionary game can be combined with reinforcement learning, so as to use the reinforcement learning mechanism to solve the decision-making problem of network attack and defense evolutionary game. Figure 1 As shown, a method for selecting network attack and defense strategies based on intelligent evolutionary game is provided, comprising:

[0032] S101. Obtain the attack and defense strategy set by analyzing the vulnerability information of network scenarios, build a network attack and defense evolutionary game decision model in combination with the bounded rational game scenario, and obtain the attack and defense benefits of different strategy combinations of the attack and defense parties based on the model;

[0033] S102. During the attack and defense game, the regret value is set according to the benefits of the strategies not implemented by both parties and the benefits of the currently implemented strategies. The probability equations of the strategies implemented by the attack and defense agents are constructed based on the regret minimization RM algorithm by using the strategy weights and the expected benefit loss of the strategies. The probability equations of the attack and defense agents are combined to construct the differential equation group for the decision selection of the game process of the attack and defense agents.

[0034] S103. Obtain the optimal strategies for both the attacker and the defender by solving the differential equations through evolutionary equilibrium.

[0035] The regret minimization algorithm is a policy-based reinforcement learning algorithm that associates the action history of the agent with his current policy decision. The core idea is that after the agent implements the policy, it will review the history of the policy implementation so far and the corresponding rewards, and regret not implementing the optimal policy afterwards. The loss function is constructed based on the expected return of the optimal policy afterwards and the actual return of the current policy to measure the regret value, so as to complete the policy update. The network attack and defense evolutionary game decision model based on the regret minimization algorithm uses the RM algorithm to determine the update rules of future policy selection according to the regret degree of the game history, to describe how the policy evolves over time, that is, by establishing a differential equation group based on the RM algorithm to characterize the dynamic changes of the probability of network defense policy selection, so as to dynamically display the network attack and defense decision process and learning behavior trajectory.

[0036] In the network attack and defense confrontation, since both the attacker and the defender have the characteristics of limited rationality, they cannot fully grasp each other's attack and defense information. Therefore, when facing uncertain attack and defense decisions, it is impossible to select the optimal strategy through a single game. It is often the case that after choosing a strategy, it is found that the effect may be better if another strategy is implemented. For example, the attacker's strategy set is {A1, A2, A3}, and the defender's strategy set is {D1, D2, D3}. Taking the defender as an example, with u d (A i ,D j ) indicates that the attacker chooses strategy A i , the defender chooses strategy D j The defender's gain is Indicates that the defender did not adopt strategy D k The regret value is the profit generated by the strategy not taken minus the current strategy D j The income generated satisfies Assume that when the attacker selects strategy A1, the defender selects D1, D2, and D3, and the game payoffs are -1, 0, and 1 respectively. In the first round, the attacker and the defender play the game with strategy (A1, D1). After the first round (afterwards), the regret value of strategy D2 can be calculated to be 1. Similarly, the regret value of strategy D3 can be calculated to be 2. Then in the second round of the game, the probability of the defender selecting strategies D1, D2, and D3 is 0, 1 / 3, and 2 / 3 respectively. Therefore, in the second round, the defender tends to select strategy D3. This is repeated. After each round, the probability of selecting each strategy is calculated through the regret value, so as to determine the strategy selection for the next round. By continuously updating the probability of strategy selection, the optimal strategy is finally found.

[0037] In the embodiment of this case, the evolutionary game is combined with the regret minimization algorithm, the network attack and defense strategy is parameterized, and the evolutionary game learning mechanism based on the replication dynamic equation is broken through by designing a network attack and defense game decision-making scheme for non-completely rational scenarios to ensure the convergence of the decision; on the other hand, the introduction of a strategy-based regret minimization algorithm to ensure the randomness of strategy learning can provide a scientific and efficient game theory tool for network attack and defense decision-making, and effectively improve the defense capabilities of network security operation and maintenance personnel.

[0038] Vulnerability scanning tools can be used to obtain vulnerability information of network scenarios. The network attack and defense evolutionary game decision model constructed in combination with the bounded rational game scenario is represented by a five-tuple (N, D, π, S, U), where N represents the set of participants in the attack and defense game, D represents the attack and defense game strategy space, π represents the attack and defense game strategy selection probability set, S represents the attack and defense game state set, and U represents the attack and defense game benefit matrix set.

[0039] The Network Attack-Defense Evolutionary Game Making-decision Model based on Regret Minimum can be expressed as: ADEG-RM = (N, D, π, S, U). Where N = (N A ,N D ) represents the set of players in the network attack and defense game, N A is the attacker, N D D=(AS,DS) represents the network attack and defense game strategy space, AS={AS1,AS2,…,AS m} represents the attacker’s strategy set, DS = {DS1, DS2, …, DS n} represents the defender’s strategy set, m and n represent the number of strategies of the attacker and defender, respectively, m, n are positive integers and m, n ≥ 2. π=(p, q) represents the network attack and defense game belief set, p=(p1, p2, …, p m ) represents a probability distribution of the attacker’s strategy set AS, that is, p i ∈p means the attacker has probability p i Random Selection Strategy AS i Carry out the attack, satisfying 1≤i≤m, q=(q1,q2,…,q n ) represents a probabilistic configuration of the defender’s strategy set DS, namely, q j ∈q means the defender takes j Random Selection Strategy DS j Implement defense, satisfying 1≤j≤n, S=(S1,S2,...,S n ) represents the state set of the network attack and defense game, and the attacker’s control over the server is regarded as the network state. A ,U D ) represents the set of profit functions of the network attack and defense game, which refers to the profit obtained by both the network attacker and the defender during the game. Different strategy combinations (AS i ,DS j ) get different benefits. A is the attacker’s payoff matrix, U D is the defender’s payoff matrix.

[0040] The attack and defense profit matrix M is composed of different attack and defense strategies (AS i ,DS j ) The attack and defense benefits generated by the game (a ij , d ij ), where A is the attacker’s strategy payoff matrix, B is the defender’s strategy payoff matrix, and the attack payoff value a ij =U A (AS i ,DS j ), defense benefit value d ij =U D (AS i ,DS j ).

[0041]

[0042] The replication dynamic equation describes that the number of individuals who choose a more successful strategy in a group gradually increases, and the proportion of the strategy selection is constantly adjusted and changed, and finally tends to a stable state. The strategy update rule is that the strategy with an expected return higher than the average return is gradually adopted by more individuals, and then the selection probability of the strategy (the proportion of individuals using this strategy in the group) changes dynamically until it stabilizes. Therefore, it can be used to study how the probability of the attacker and defender choosing their own strategies changes dynamically over time during the attack and defense evolutionary game. Then the attacker attacks with probability p i Select attack strategy AS i , the defender has probability q j Select Defense Strategy DS j The replication dynamic evolution equation of

[0043]

[0044] Among them, A and B are the payoff matrices of the attacking and defending parties respectively, and the probability vector p={p1,p2,...,p m} describes all pure attack strategies {AS1,AS2,...,AS m}, the probability vector q={q1,q2,...,q n} describes all pure defense strategies {DS1,DS2,...,DS n}. For the attacker, p i Indicates the selection of attack strategy AS i The probability of i / dt means selection strategy AS i The rate of change of probability with time, (Aq) i Indicates the policy AS i The expected return, p T Aq represents the average benefit of the attack strategy set; for the defender, q j Indicates the selection of defense strategy DS j The probability of j / dt means selection strategy DS j The rate of change of probability over time, (Bp) j Defense strategy DS j The expected return, q T Bp represents the average return of the defense strategy set. From formula (1), it can be seen that the probability of strategy selection is proportional to the difference between the expected return of a single strategy and the average return of the strategy set.

[0045] The loss function based on expected return can be expressed as:

[0046]

[0047] Since the expected return can better reflect the overall effect of a defense strategy on all attack strategies, the loss function is set based on the expected return. is a measure of regret, where Indicates the implementation of a certain strategy DS j Expected return (Bp) j , r represents the maximum value of the expected returns of all individual strategies, that is, r = max k (Bp) k .

[0048] The Polynomial Weight algorithm is a type of RM algorithm that calculates the regret of the relative optimal strategy after the fact. It defines the value assigned to the strategy DS j The weight of Incurring losses The relationship between the two is to continuously update the strategy DS through the loss j The degree of preference in the strategy set can be expressed as follows:

[0049]

[0050] Among them, λ is the learning ability parameter, which is used to control the speed of weight change. The process of finding the optimal strategy can actually be understood as the process of increasing the weight assigned to the strategy. At the beginning of the game, the weight of each strategy in the strategy set is equal. As the attack and defense game proceeds, the defender continuously enhances his understanding of uncertain information such as the game environment and attack knowledge, and gradually adjusts the weight of each strategy in the strategy set. It can be seen that only when the loss of a certain strategy compared with the optimal strategy is smaller, the weight of the strategy will increase in the next round of the game.

[0051] The probability of selecting a network defense strategy based on the RM algorithm can be expressed as:

[0052]

[0053] As a policy-based learning algorithm, the RM algorithm directly models the strategy, as shown in formula (4), and can better handle the learning of random strategies selected by probability. The defender's strategy DS in the attack-defense game at time t is j The weights it has, which are based on the loss function Update, the larger the weight, the greater the probability that the strategy will be selected, so as to achieve the purpose of updating the probability of strategy selection.

[0054] The process of finding the optimal strategy for both the attacker and defender is a process of continuous learning, exploration, and optimization. During the game, the probability of selecting each strategy is gradually updated. Denotes the defense strategy DS at time t j The update of the selection probability can be expressed by formula (5) as follows.

[0055]

[0056] From the above formula, we can see that the update of strategy selection probability is Depends on the assigned weights and the probability of the strategy being selected.

[0057] By associating equations (3), (4) and (5), we can obtain the network attack and defense evolutionary game decision equation (6) based on the RM algorithm, which describes the exploration of the optimal strategy by both the attacker and the defender and describes the update rules for the selection of the attack and defense strategies.

[0058] The decision equation of the network attack and defense evolutionary game based on the RM algorithm can be expressed as:

[0059]

[0060] The above formula is the expected loss weighted replication dynamic equation obtained based on the RM algorithm, which is used to describe the dynamic evolution of the limited rational strategy selection of both sides in the attack and defense game over time. In the network attack and defense confrontation, both sides update the probability of strategy selection through the regret value to achieve the purpose of optimal strategy selection, that is, in multiple games, the weight of each strategy in the strategy set is continuously updated through the learning of the regret value, so as to find their respective optimal strategies.

[0061] Based on the above method and solution, the algorithm for achieving the optimal network attack and defense evolutionary game decision selection can be designed as shown in Algorithm 1.

[0062]

[0063]

[0064]

[0065] Furthermore, based on the above method, an embodiment of the present invention also provides a network attack and defense strategy selection system based on intelligent evolutionary game, comprising: a model building module, an attack and defense analysis module and an optimal output module, wherein:

[0066] The model building module is used to obtain the attack and defense strategy set by analyzing the vulnerability information of the network scenario, build a network attack and defense evolutionary game decision model in combination with the limited rational game scenario, and obtain the attack and defense benefits of different strategy combinations of the attack and defense parties based on the model;

[0067] The attack and defense analysis module is used to set the regret value according to the benefits of the unimplemented strategies and the current implemented strategies of both parties during the attack and defense game process, and to construct the probability equations of the implementation strategies of the attack and defense agents based on the regret minimization RM algorithm by using the strategy weights and the expected benefit loss of the strategies. The probability equations of the attack and defense parties are combined to construct the differential equation group for the decision-making selection of the game process between the attack and defense parties.

[0068] The optimal output module is used to obtain the optimal strategies of both the attacker and the defender by solving the evolutionary equilibrium of the differential equations.

[0069] In order to verify the effectiveness of this solution, the following is a further explanation based on experimental data:

[0070] A small enterprise network scenario is deployed to verify the effectiveness of the proposed game model. First, the network scenario is set up, and the attack and defense strategy set and the attack and defense strategy benefit matrix are given according to the vulnerability information; secondly, the probability of selecting the optimal defense strategy under different states is calculated, and the evolution trajectory of the defense strategy selection is dynamically characterized; then the stability of the defense strategy selection is verified, that is, the probability of defense strategy selection does not change with the change of the initial state. Finally, the algorithm proposed in this case is compared with the attack and defense evolutionary game algorithm based on the replication dynamic equation to verify the convergence and learning efficiency of the algorithm proposed in this case.

[0071] 1. Experimental setup

[0072] See also Figure 2 As shown in the figure, the network is mainly composed of three types of server clusters: LDAP server, Web server and FTP server. Among them, the LDAP server under Windows system is built based on Apache server and Mysql database, the Web server is built based on PentesterLab, and finally the FTP server is built based on the Docker tool Vulnstudy, and the server vulnerability is set using the vulnerability scanning tool AWVS. The attacker's goal is to invade the server cluster, use the vulnerabilities of each server cluster as a springboard, obtain the control authority of the server, and finally steal the key network data in the FTP server cluster through different attack paths. The defender's goal is to protect the server cluster, monitor and identify the network attack path, and block the attack by deploying an intrusion detection system.

[0073] The attacker has the User privileges on the LDAP server in the initial state, and his purpose is to steal key data from the FTP server. The attacker's attack strategy can be defined as vulnerability scanning and exploitation based on each server, corresponding to Exp-LDAP, Exp-Web, and Exp-FTP. Without loss of generality, Exp-LDAP means that the attacker exploits a specific vulnerability (CVE-2016-5195) to attack the LDAP server, Exp-Web means that the attacker exploits a specific vulnerability (CVE-2017-5095) to attack the Web server, and Exp-FTP means that the attacker exploits a specific vulnerability (CVE-2015-3306) to attack the FTP server. For specific vulnerability information, see Table 1. Set two attack paths:

[0074] Attack path 1: Exp-LDAP—>Exp-FTP;

[0075] Attack path 2: Exp-LDAP—>Exp-Web—>Exp-FTP.

[0076] According to different attack paths, the network status is also changing accordingly, such as Figure 3As shown, the dotted line on the left is attack path 1, and the dotted line on the right is attack path 2. In the initial state S0, Exp-LDAP can be implemented through a specific vulnerability to reach state S1, in which the attacker has the root privilege of the LDAP server, the user privilege of the Web server, and the user privilege of the FTP server; in state S1, the attacker can obtain the root privilege of the FTP server through remote code execution to achieve the ultimate goal, or implement Exp-Web on the Web server through a cross-site scripting attack, and then reach state S2, in which the attacker has the root privilege of the Web server and the user privilege of the FTP server. In state S2, the attacker can implement Exp-FTP to obtain the root privilege of the FTP server to achieve the ultimate goal. Of course, the attacker may also worry about being detected and not implement the No-Exp attack, but continue to stay in the corresponding state.

[0077] At the same time, for specific vulnerability scanning attacks on different server vulnerabilities, the defender monitors the services and traffic running on the host and deploys the corresponding intrusion detection system. The defender's defense strategy can be defined as attack detection and intrusion defense based on each server vulnerability, corresponding to Mon-LDAP, Mon-Web, and Mon-FTP respectively. Without loss of generality, Mon-LDAP means that the defender uses Auditd software to monitor the access rights and traces of sensitive files in the system, Mon-Web means that the defender uses OSSEC HIDS software to detect the log information of the Web server, and Mon-FTP means that the defender uses Snort software to monitor the traffic of the FTP server port in a specific state. The defender may also be limited by resources and performance, and thus choose not to implement monitoring, which can be represented by No-mon.

[0078] Assuming that the resources available to the defender are limited, it is necessary to select the optimal strategy to implement monitoring; and the attacker needs to avoid the attack behavior being detected by the defender, so it is also necessary to implement the optimal strategy to exploit the vulnerability. The vulnerability information of various servers in the experimental network is shown in Table 1. Vulnerabilities are inherent security defects residing on a given port of the server, which can be measured based on confidentiality, integrity, and availability (CIA).

[0079] Table 1 Server vulnerability information

[0080]

[0081] Assuming that the gain of the attacker is the loss of the defender, the attack and defense gains are regarded as zero-sum, that is, the sum of the attack gain and the defense gain is zero. In the quantification method of the attack and defense strategy gain, the benefit matrix of the network attack and defense strategy under different states S0, S1, and S2 can be obtained according to the characteristics of different attack and defense strategies, as shown in Table 2, Table 3, and Table 4.

[0082] Table 2 Profit matrix of attack and defense strategy under state S0

[0083]

[0084] Table 3 Profit matrix of attack and defense strategy under state S1

[0085]

[0086] Table 4 Profit matrix of attack and defense strategies under state S2

[0087]

[0088] 2. Numerical analysis

[0089] 1) Probability of selecting the optimal defense strategy under different states

[0090] Initialize the attack-defense evolutionary game model according to Algorithm 1. The attacker’s strategy space is {No-exp, Exp-LDAP, Exp-Web, Exp-FTP}, and the probability distribution of the attack strategy space is {p1, p2, p3, p4} and satisfies The defender's strategy space is {No-mon,Mon-LDAP,Mon-Web,Mon-FTP}, and its probability distribution is {q1,q2,q3,q4} and satisfies Assuming that both the attacker and the network administrator have certain learning ability, we set λ = 0.3. Next, we establish the attack and defense strategy evolution equation based on the RM algorithm under different states and study the evolution process of the optimal defense strategy under each state.

[0091] The evolution trajectory of each defense strategy under states S0, S1 and S2 is obtained through simulation, such as Figure 4As shown in the figure. The horizontal axis t represents the number of attack and defense games, and the vertical axis represents the probability of selecting a defense strategy. In order to better illustrate the evolution effect of strategy selection, the corresponding strategy is selected with equal probability in the initial state. For two-strategy games, such as states S0 and S2, the initial selection probability of attack and defense strategies is set to 1 / 2; for three-strategy games, such as state S1, the initial selection probability of attack and defense strategies is set to 1 / 3. As can be seen from the figure, the change curve of the optimal defense strategy of the defense strategy {No-mon, Mon-LDAP, Mon-Web, Mon-FTP} in different states. In the process of repeated games with attackers, network administrators have continuously tried and failed, learned and adjusted strategies, and the probability of selecting defense strategies finally reaches a stable state. When the defender faces an attack in state S0, the optimal strategy of the defender is finally implemented by selecting the strategy {No-mon, Mon-LDAP} with a mixed probability of {q1=0.41862, q2=0.58138}; the optimal defense strategy of the defender in state S1 is finally implemented by selecting the strategy {No-mon, Mon-Web, Mon-FTP} with a mixed probability of {q1=0.00006, q3=0.53979, q4=0.46015}; the optimal defense strategy of the defender in state S2 is finally implemented by selecting the strategy {No-mon, Mon-FTP} with a mixed probability of {q1=0.15961, q2=0.84039}, thereby ensuring the maximum defense effect at the minimum cost in each state.

[0092] In the initial state S0, the attacker implements Exp-LDAP. For the defender, the optimal defense strategy is to adopt Mon-LDAP blocking attack to cut off the attacker's attack source to the FTP server, or to temporarily not adopt monitoring considering the limited defense resources and high cost factors; in state S1, the root privilege of the FTP server can be obtained directly and indirectly. Therefore, the defender can block the attack in state S1. The optimal defense strategy is implemented with the probability of {q1=0.00006,q3=0.53979,q4=0.46015}, which can prevent the attacker from attacking the FTP server directly and indirectly attacking the Web server. If the defender mistakenly selects the No-mon strategy to allow the attacker to obtain the root privilege of the Web server and reach state S2, then when the attacker implements Exp-FTP, the defender will select the optimal defense strategy Mon-FTP with a high probability of 0.84039 to block the attack on the FTP server to prevent the FTP server from being compromised and causing the loss of key data.

[0093] 2) Convergence of defense strategy selection

[0094] In order to better illustrate the stability of defense strategy selection, the attack and defense game in state S1 is taken as an example. The following attack and defense scenarios are set. The first case is the evolution of attack and defense strategies under different defense strategy selection probabilities at the initial moment. It is assumed that the attacker randomly selects the attack strategy with an equal probability of 1 / 3, and the defender's strategy selection is changed to observe the evolution trajectory of the optimal defense strategy; the second case is the evolution of attack and defense strategies under different attack strategy selection probabilities at the initial moment. It is assumed that the defender randomly selects the defense strategy with an equal probability of 1 / 3, and the attacker's strategy selection is changed to observe the evolution trajectory of the optimal defense strategy.

[0095] The first case is the probability of selecting different initial defense strategies. For different defense strategies, the attacker randomly implements the attack strategy {No-exp, Exp-Web, Exp-FTP} with an equal probability of 1 / 3. The initial probabilities of selecting the defense strategy {No-mon, Mon-Web, Mon-FTP} correspond to the following three cases: ①{q1=0.1, q3=0.3, q4=0.6}; ②{q1=0.3, q3=0.5, q4=0.2}; ③{q1=0.6, q3=0.1, q4=0.3}. Through experiments, we can obtain the evolution trajectory diagram of the defense strategy of state S1 in the above three cases, as shown in the figure. Figure 5 shown.

[0096] The second case is the probability of selecting different initial attack strategies. For different attack strategies, the defender randomly selects the defense strategy {No-mon, Mon-Web, Mon-FTP} with an equal probability of 1 / 3. The initial probabilities of selecting the attack strategy {No-exp, Exp-Web, Exp-FTP} correspond to the following three cases: ①{p1=0.1, p3=0.3, p4=0.6}; ②{p1=0.2, p3=0.5, p4=0.3}; ③{p1=0.7, p3=0.1, p4=0.2}. At this time, the evolution trajectory diagram of the defense strategy of state S1 in the above three cases can be obtained through experiments, as shown in the figure below: Figure 6 shown.

[0097] It can be seen from the above figure that the decision result of the optimal defense strategy will not change due to the different probabilities of selecting the defense strategy and the attack strategy at the beginning. In the process of the game, it will eventually reach a stable state and always maintain this stable state.

[0098] 3) The impact of changes in learning ability on the selection of defense strategies

[0099] Take state S2 as an example to illustrate the impact of different learning abilities on the selection of the optimal defense strategy. The attacker and defender initially randomly select their own strategies in the attack and defense strategy space with a probability of 1 / 2. In this scenario, by changing the learning ability parameter λ, the impact of the improvement of learning ability on the evolution of the strategies of both attackers and defenders is observed. That is, when λ=0.1, 0.3, 0.5, 0.7, 0.9, the evolution law of the game between the attacker and defender is studied. Using Algorithm 1 to solve the defense strategy evolution equation under state S2, the change curve of the defense decision results under different learning abilities can be obtained as follows Figure 7 As shown in the figure, the decision results of the optimal defense strategy eventually tend to be stable, but the time required for different learning abilities to reach stability is obviously different. The figure shows that as the learning ability λ continues to increase, the time it takes for the probability of selecting the optimal defense strategy to evolve to a stable state becomes shorter, indicating that in the process of attack and defense evolution, as the defender's learning ability improves, the strategy selection has a more accurate understanding, so it can make quick decisions on strategy selection and select the best defense strategy Mon-FTP.

[0100] 4) Comparison of game strategy selection methods

[0101] Considering that the attacker and defender are affected by factors such as attack and defense knowledge and computing power, the attacker and defender only have partial information about the opponent, and the game requires continuous trial and error learning, which is a gradual optimization process. In order to better illustrate the superiority of this method, this method is compared with the strategy selection method based on the traditional replication dynamic equation. The comparison results are as follows Figure 8 As shown in the figure, the x-axis is the number of games t, and the y-axis is the probability of selecting the optimal defense strategy. The dark gray solid line in the figure represents the evolution trajectory of the optimal defense strategy based on the strategy selection method of the traditional replication dynamic equation, and the light gray solid line represents the evolution trajectory of the optimal defense strategy of this case. As can be seen from the figure, this case has found the optimal defense strategy at t = 504, while the strategy selection method based on the traditional replication dynamic equation found the optimal defense strategy at t = 578. Therefore, compared with the strategy selection method based on the traditional replication dynamic equation, this method takes less time and is faster to learn the optimal strategy, and the convergence rate of the optimal strategy is increased by 12.8%. At the same time, the fluctuation range during the learning process is relatively small, and the impact on the defender's judgment is relatively small, which has better convergence and learning efficiency.

[0102] Therefore, based on the above experimental data, it can be better explained that the solution in this case, by combining evolutionary game with regret minimization algorithm and parameterizing the network attack and defense strategy, can improve the correctness and practicality of strategy selection in the attack and defense game process, and facilitate the optimal allocation of resources in network threat defense.

[0103] Unless otherwise specifically stated, the relative steps, numerical expressions and values ​​of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0104] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0105] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.

[0106] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits, and accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of combination of hardware and software.

[0107] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A network attack and defense strategy selection method based on intelligent evolutionary game, characterized in that: Contains the following: By analyzing the vulnerability information of network scenarios to obtain the set of attack and defense strategies, a network attack and defense evolutionary game decision model is constructed in combination with the limited rational game scenario, and the attack and defense benefits of different strategy combinations of the attack and defense parties are obtained based on the model; In the process of attack and defense game, the regret value is set according to the benefits of the unimplemented strategies of both parties and the benefits of the currently implemented strategies. The strategy weights in the attack and defense game are set according to the expected benefits of the strategies. The strategy weights and the expected benefits of the strategies are used to construct the strategy selection probability equations of the strategies implemented by the attack and defense agents based on the regret minimization RM algorithm. The probability equations of the attack and defense parties are combined to construct the differential equations for the decision selection of the game process between the attack and defense parties. Among them, the strategy selection probability equation is expressed as: The defender's strategy DS in the attack-defense game at time t is j The weight it has, It means that the defender selects the attack and defense game strategy DS at time t j The probability of It represents the attacker's strategy AS in the attack and defense game at time t j The weight it has, Indicates that the attacker selects the attack and defense game strategy AS at time t j probability; The optimal strategies for both the attacker and the defender are obtained by solving the evolutionary equilibrium of the differential equations.

2. The method for selecting network attack and defense strategies based on intelligent evolutionary game according to claim 1 is characterized in that: Before obtaining the attack and defense strategy set by analyzing the vulnerability information of the network scenario, it also includes: using vulnerability scanning tools to obtain the vulnerability information of the network scenario.

3. The method for selecting network attack and defense strategies based on intelligent evolutionary game according to claim 1 or 2, characterized in that: The network attack and defense evolutionary game decision-making model constructed in combination with the bounded rational game scenario is represented by the quintuple (N, D, π, S, U), where N represents the set of attack and defense game participants, D represents the attack and defense game strategy space, π represents the attack and defense game strategy selection probability set, S represents the attack and defense game state set, and U represents the attack and defense game payoff matrix set.

4. The method for selecting network attack and defense strategies based on intelligent evolutionary game according to claim 1, characterized in that: The strategy weight in the attack and defense game set according to the expected return of the strategy is expressed as Among them, λ is the learning ability parameter, The defender implements strategy DS in the attack and defense game at time t-1 j The loss function when The attacker implements strategy AS in the attack and defense game at time t-1 j The loss function when .

5. The method for selecting network attack and defense strategies based on intelligent evolutionary game according to claim 4 is characterized in that: The loss function of the attacking and defending parties is represented by the difference between the maximum expected return of all individual strategies of the attacking and defending parties and the expected return of implementing their corresponding strategies at the moment of the attack and defense game.

6. The method for selecting network attack and defense strategies based on intelligent evolutionary game according to claim 1, characterized in that: The differential equation group for the decision-making process of the attacking and defending sides is expressed as Among them, A and B represent the profit matrices of the attacking and defending parties respectively, the probability vector p is a vector composed of probability elements selected from all pure attack strategies, and the probability vector q is a vector composed of probability elements selected from all pure defense strategies. i Indicates the selection of attack strategy AS i The probability of i / dt means selection strategy AS i The rate of change of probability with time, (Aq) i Indicates the policy AS i The expected return, p T Aq represents the average benefit of the attack strategy set; q j Indicates the selection of defense strategy DS j The probability of j / dt means selection strategy DS j The rate of change of probability over time, (Bp) j Defense strategy DS j The expected return, q T Bp represents the average benefit of the defense strategy set, λ is the learning ability parameter, and k represents the maximum strategy mark among all single strategy expected benefits.

7. The method for selecting network attack and defense strategies based on intelligent evolutionary game according to claim 6 is characterized in that: The optimal strategy for both the attacker and the defender is obtained by solving the evolutionary equilibrium of the differential equations. The probability of strategy selection and the weight of the strategy in the strategy set are updated by learning the regret value, and the optimal strategy is selected based on the updated weight.

8. A network attack and defense strategy selection system based on intelligent evolutionary game, characterized in that: The method according to claim 1 is implemented, comprising: a model building module, an attack and defense analysis module and an optimal output module, wherein: The model building module is used to obtain the attack and defense strategy set by analyzing the vulnerability information of the network scenario, build a network attack and defense evolutionary game decision model in combination with the limited rational game scenario, and obtain the attack and defense benefits of different strategy combinations of the attack and defense parties based on the model; The attack and defense analysis module is used to set the regret value according to the benefits of the unimplemented strategies and the current implemented strategies of both parties during the attack and defense game process, and to construct the probability equations of the implementation strategies of the attack and defense agents based on the regret minimization RM algorithm by using the strategy weights and the expected benefit loss of the strategies. The probability equations of the attack and defense parties are combined to construct the differential equation group for the decision-making selection of the game process between the attack and defense parties. The optimal output module is used to obtain the optimal strategies of both the attacker and the defender by solving the evolutionary equilibrium of the differential equations.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory, used to store computer programs; A processor, configured to execute a program stored in a memory and implement the method steps described in any one of claims 1 to 7 when the program is executed.