A multi-stage network defense decision-making method based on an improved evolutionary game model
By constructing an improved evolutionary game model and solving for an evolutionarily stable defense strategy, the problems of complexity and insufficient adaptability of network defense decision-making algorithms are solved, achieving faster response speed and adaptability to observation errors, thereby improving the effectiveness of network defense.
Patent Information
- Application Number
- CN202210949423.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2042-08-09
AI Technical Summary
Existing network defense decision-making algorithms are complex and lack agility and adaptability, resulting in slow response speed and insufficient network defense situational awareness, which cannot effectively solve the problems of false alarms and false negatives in intrusion detection systems.
An improved evolutionary game model is constructed. By initializing parameters and step size, a replicating dynamic equation is built to solve the evolutionary stable defense strategy. The optimal pure defense strategy set is then solved by traversing the set of strategies. The improved evolutionary game model is used to improve the response timeliness and adaptability to observation errors of network defense decisions.
It improves the responsiveness of network defense decisions and its adaptability to observation errors, enabling the output of reliable optimal defense strategies in environments with incomplete information, thereby enhancing the agility and accuracy of network defense.
Smart Images

Figure CN115834100B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer network information security, and particularly relates to a multi-stage network defense decision-making method based on an improved evolutionary game model. BACKGROUND
[0002] With the development of network information technology, new challenges have been brought to network security and privacy protection. The rapid development of new network attack and defense theories and technologies such as advanced persistent threats and dynamic target defense makes the network security situation and attack and defense game behaviors increasingly complex. The current network defense decision-making faces two problems: on the one hand, due to the relatively complex algorithm and the relatively insufficient agility, which directly affects the speed of the network defense party in responding to attack events and the timeliness of network defense decision-making; on the other hand, the network defense situation awareness capability is insufficient, which causes the defense party to be unable to fundamentally solve the false positive and false negative problems of the intrusion detection system when facing incomplete information or false information. Therefore, it is necessary to improve one or more problems in the above technical solutions.
[0003] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0004] The purpose of the embodiments of the present disclosure is to provide a multi-stage network defense decision-making method based on an improved evolutionary game model to improve the response timeliness and adaptability to observation errors of network defense decision-making.
[0005] According to an embodiment of the present disclosure, a multi-stage network defense decision-making method based on an improved evolutionary game model is provided, and the manufacturing method comprises the following steps:
[0006] An improved evolutionary game model is constructed, and parameters and steps in the improved evolutionary game model are initialized;
[0007] According to the attack and defense dynamic game process in the improved evolutionary game model, a replication dynamic equation is constructed and solved to obtain an evolutionary stable defense strategy;
[0008] All the above steps are repeated to traverse and solve all the evolutionary stable defense strategies to obtain an optimal defense pure strategy set.
[0009] In an exemplary embodiment of the present disclosure, in the step of constructing an improved evolutionary game model and initializing parameters and steps in the improved evolutionary game model, the improved evolutionary game model comprises a 7-element ordered group model wherein,
[0010] N = (NA D ), N represents the set of attack-defense game participants; N A = (N A1 , N A2 ,..., N Aj ), N A represents the set of attack game participants, N A1 , N A2 ,..., N Aj represents attack game participant individuals; N D = (N D1 , N D2 ,..., N Di ), N D represents the set of defense game participants, N D1 , N D2 ,..., N Di represents defense game participant individuals;
[0011] S = (S A , S D ), S represents the set of mixed strategies of attack-defense game participants; S A = (S A1 , S A2 ,..., S Aj ), S A represents the set of pure strategies of attack participants, S A1 , S A2 ,..., S Aj represents the pure strategy of attack participant individuals; S D = (S D1 , S D2 ,..., S Di ), S D represents the set of pure strategies of defense participants, S D1 , S D2 ,..., S Di represents the pure strategy of defense participant individuals;
[0012] P = (P A , P D ), P represents the set of actual game beliefs of attack-defense game participants; P A = (P A1 , P A2 ,..., P Aj ), P A represents the set of actual game beliefs of attack, (P A1 , P A2 ,..., P Aj )
[0013] P A ' = (PA1 ′,P A2 ′,…,P Aj ′), P A ′ represents the set of the attacker's experienced game beliefs, (P A1 ′,P A2 ′,…,P Aj ′) represents the defender's judgment of the attacker's chosen pure strategy S. Aj The probability of;
[0014] U=(U A U D U represents the set of game payoffs; U A =(U A1 U A2 ,…,U Aj ), U A Represents the set of payoffs for the attacker, (U A1 U A2 ,…,U Aj ) represents individual player N in the attacking game. Aj Adopting pure strategy S Aj The expected return obtained in a stage of a game; U D =(U D1 U D2 ,…,U Di ), U D Represents the set of payoffs for the defender, (U D1 U D2 ,…,U Di ) represents individual N, the defensive player in the game. Di Adopting pure strategy S Di The expected return obtained in a stage of a game;
[0015] e represents the observation error;
[0016] A set representing short-term forecasts; This represents the set of short-term predictions of the attacker's game-theoretic beliefs. Indicates the game belief P Aj Short-term forecasts; Indicates the game belief P Di A set of short-term forecasts This represents the set of short-term predictions of the attacker's game-theoretic beliefs.
[0017] In an exemplary embodiment of this disclosure,
[0018] When U Di When = 1, the game has no effect on the pure strategy choices of the attacking and defending sides;
[0019] When UDi When U > 1, the game has a positive impact on the selection of pure strategies of both the attacker and the defender.
[0020] When U Di < 1, the game has a negative impact on the selection of pure strategies of both the attacker and the defender.
[0021] In an example embodiment of the present disclosure, the game payoff U D of the defender is calculated according to the following formula:
[0022] U D = φ·V r -D cost (1)
[0023] wherein V r represents the resource payoff; D cost represents the cost of the defender; and φ represents the defense effect.
[0024] In an example embodiment of the present disclosure, the game payoff U A of the attacker is calculated according to the following formula:
[0025] U A = λ·V r -A cost (2)
[0026] wherein V r represents the resource payoff; A cost represents the cost of the attacker; and λ represents the infection probability.
[0027] In an example embodiment of the present disclosure, in the step of initializing the parameters and the step size of the improved evolutionary game model, the step size is the time delay T between the input and the output of the improved evolutionary game model.
[0028] In an example embodiment of the present disclosure, in the step of constructing the replicator dynamic equation and solving it according to the attack-defense dynamic game process in the improved evolutionary game model, the calculation formula of the replicator dynamic equation comprises:
[0029]
[0030] wherein, d(P Di ) / dt represents the rate of change of the game belief P Di over time t; P Aj (t)′ represents the experienced game belief of the attacker at time t; represents the prediction of the game belief P Aj at time t; and β(x) represents the response function. P (t) represents the actual game belief of the attacker at time t, e(t) represents the observation error at time t. Aj and its derivative the joint response of the defense player set at time t; P (t) represents the average income of the defense player set at time t; i and j are one-to-one correspondence.
[0031] In an example embodiment of the present disclosure, the experience game belief P Aj of the attacker at time t is calculated by the following formula:
[0032] P Aj (t)′=P Aj (t)+e(t) (4)
[0033] wherein P Aj (t) represents the actual game belief of the attacker at time t, e(t) represents the observation error at time t;
[0034] The calculation formula of the average income of the defense player set at time t includes:
[0035]
[0036] In an example embodiment of the present disclosure, the value range of the observation error e(t) at time t includes any random number in [-1, 1].
[0037] In an example embodiment of the present disclosure, the set N of attack-defense game participants has a strict Nash equilibrium, and the response function β(x) is a monotonic function.
[0038] The technical solution provided by the present disclosure can include the following beneficial effects:
[0039] The present disclosure provides a multi-stage network defense decision-making method based on an improved evolutionary game model, which can improve the response timeliness of network defense decision-making and the adaptability to observation errors.
[0040] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0041] The drawings herein are incorporated into the specification and form part of the specification, which show embodiments consistent with the present disclosure, and together with the specification, serve to explain the principles of the present disclosure. It is obvious that the drawings in the following description are only some embodiments of the present disclosure, and those skilled in the art can obtain other drawings from these drawings without creative labor.
[0042] Figure 1A schematic diagram showing steps of the multi-stage network defense decision-making method based on the improved evolutionary game model in the exemplary embodiments of the present disclosure is shown.
[0043] Figure 2 A schematic diagram showing the structure of a typical biological control system in the exemplary embodiments of the present disclosure is shown.
[0044] Figure 3 A schematic diagram showing Figure 2 A schematic diagram showing the structure of a typical biological control system combined with the replication dynamic characteristics in the exemplary embodiments of the present disclosure is shown.
[0045] Figure 4 A schematic diagram showing the topological environment of a network information system in the simulation experiment in the exemplary embodiments of the present disclosure is shown.
[0046] Figure 5 A convergence trajectory curve diagram showing the evolutionary stable solution in the game model when e(t)≡0 is verified in the simulation experiment in the exemplary embodiments of the present disclosure is shown.
[0047] Figure 6 A convergence trajectory curve diagram showing the evolutionary stable solution in the game model when e(t)≠0 is verified in the simulation experiment in the exemplary embodiments of the present disclosure is shown.
[0048] Figure 7 A curve diagram showing the difference comparison of the time performance of the algorithms of different game models in the simulation experiment in the exemplary embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0049] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations can, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example implementations to those skilled in the art. Features, structures or characteristics described can be combined in any suitable manner in one or more implementations.
[0050] In addition, the accompanying drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and serve to explain the principles of the present disclosure. Like reference numerals refer to like elements throughout the various drawings. Identical or similar components shown in the drawings are designated with the same reference numerals, and thus repetitive description of them will be omitted. Some block diagrams shown in the drawings are functional entities, and thus do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0051] The present example implementations provide methods, with reference to Figure 1 as shown, comprising the following steps:
[0052] Step S101: Construct an improved evolutionary game model, and initialize the parameters and step size in the improved evolutionary game model;
[0053] Step S102: According to the attack-defense dynamic game process in the improved evolutionary game model, construct a dynamic equation of replication and solve it to obtain an evolutionary stable defense strategy;
[0054] Step S103: Repeat all the above steps to traverse and solve all evolutionary stable defense strategies to obtain an optimal defense pure strategy set.
[0055] In the following, each step of the above method in the present example embodiment will be described in more detail.
[0056] Network attack and defense is a complex dynamic game process, and its characteristics mainly include six aspects: target opposition, strategy dependence, non-cooperative relationship, incomplete information, dynamic evolution, and interest-driven. Specifically, the three characteristics of target opposition, strategy dependence, and non-cooperative relationship reflect the antagonistic and heterogeneous properties of network attack and defense. This non-cooperative antagonistic relationship makes the attack and defense sides try to conceal their decision-making information as much as possible, so in the game process, the game sides can only master their own information and incomplete information of the other side, and these information will also be dynamically updated with the changes of network environment and opponent preferences, causing decision-making difficulties. In this uncertain environment, the internal driving force of the decision-making behavior of the attack and defense sides is the principle of maximizing benefits, that is, regardless of the changes of the opponent and the environment, the strategy with the highest expected return at the moment is always chosen. Here, the basic assumptions of network attack and defense game and the improved evolutionary game model are proposed.
[0057] Assumption 1: The information environment of network attack and defense game is an incomplete information environment. Under the condition of incomplete information, the attack and defense sides can observe the game beliefs of the opponent, but cannot know the benefits of the strategy.
[0058] Assumption 2: Network attack and defense game is a symmetric game, and all game participants are divided into network attack and network defense according to their own attributes.
[0059] Assumption 3: The game participants of network attack and defense game follow the principle of bounded rationality. That is, the purpose of all game participants is to maximize the benefits of the strategy.
[0060] Assumption 4: Network attack and defense game is in the form of heterogeneous group evolutionary game. The group in the game represents the combination of game participants with the same attack and defense attributes; the subgroup represents the set of game participants who choose the same game strategy in the same group.
[0061] Reference Figure 1 In step 1,
[0062] Definition 1: The constructed Improved Evolutionary Game Model (IEGM) can be represented as a 7-tuple ordered set where,
[0063] Combining with Hypothesis 2, N = (N A ,N D ), N represents the heterogeneous group, i.e., the set of attack-defense game participants; N A = (N A1 ,N A2 ,…,N Aj ), N A represents the set of attack game participants, N A1 ,N A2 ,…,N Aj represents attack game participant individuals; N D = (N D1 ,N D2 ,…,N Di ), N D represents the set of defense game participants, N D1 ,N D2 ,…,N Di represents defense game participant individuals.
[0064] S = (S A ,S D ), S represents the mixed strategy set of attack-defense game participants; S A = (S A1 ,S A2 ,…,S Aj ), S A represents the set of pure strategies of attack participants, S A1 ,S A2 ,…,S Aj represents the pure strategy of attack participant individuals; S D = (S D1 ,S D2 ,…,S Di ), S D represents the set of pure strategies of defense participants, S D1 ,S D2 ,…,S Di represents the pure strategy of defense participant individuals.
[0065] P = (P A ,P D ), P represents the set of actual game beliefs of attack-defense game participants; P A = (P A1 ,P A2 ,…,PAj ), P A represents the set of actual beliefs of the attacker, (P A1 , P A2 ,..., P Aj ) represents the probabilities of the pure strategies S Aj actually chosen by the attacker; P D = (P D1 , P D2 ,..., P Di ), P D represents the set of actual beliefs of the defender, (P D1 , P D2 ,..., P Di ) represents the probabilities of the pure strategies S Di actually chosen by the defender;
[0066] P A ' = (P A1 ', P A2 ',..., P Aj '), P A ' represents the set of empirical beliefs of the attacker, (P A1 ', P A2 ',..., P Aj ') represents the probabilities of the pure strategies S Aj chosen by the attacker as judged by the defender;
[0067] U = (U A , U D ), U represents the set of payoffs; U A = (U A1 , U A2 ,..., U Aj ), U A represents the set of payoffs of the attacker, (U A1 , U A2 ,..., U Aj ) represents the expected payoffs of the individual N Aj of the attacker's players who adopts the pure strategy S Aj in a stage game; U D = (U D1 , U D2 ,..., U Di ), U D represents the set of payoffs of the defender, (U D1 , U D2 ,..., U Di ) represents the expected payoffs of the individual N Di of the defender's players who adopts the pure strategy S Di in a stage game; The average income of the set of attack participants, The average income of the set of defense participants, The income in this embodiment refers to the incremental effect of the influence of the game on the adaptability of the game participants. Taking the defense as an example, when U Di = 1, the game has no influence on the selection of the pure strategy of the attack and defense in the game; when U Di > 1, the game has a positive influence on the selection of the pure strategy of the attack and defense in the game; and when U Di < 1, the game has a negative influence on the selection of the pure strategy of the attack and defense in the game.
[0068] e represents an observation error, which refers to the false estimation of the belief P of the opponent in the game due to induction, fraud and other factors.
[0069] e(t) represents the observation error at time t, which appears in the form of probability and is a random number in [-1, 1], that is, |e(t)| ≤ 1.
[0070] represents the set of short-term predictions; represents the set of short-term predictions of the attack game belief, represents the short-term prediction of the game belief P Aj . represents the set of short-term predictions of the game belief P Di . represents the set of short-term predictions of the attack game belief.
[0071] In step S102, the evolutionary game is a multi-stage dynamic game process, and the game result of each stage will have an influence on the game of the next stage. By improving the attack and defense dynamic game process in the evolutionary game model, the derivation process of the replicator dynamic equation is as follows:
[0072] The replicator dynamic equation of the heterogeneous group evolution defense at time t is:
[0073]
[0074] Wherein, d(P Di ) / dt is the change rate of the game belief P Di with time t, which is denoted by the derivative P Di in the following; U Di (t) is the expected income of the individual N Di adopting the pure strategy S Di at time t, is the average income of the set of defense participants at time t.
[0075] The defender's gain is actually a logical mapping of the attacker's attack strategy. Define the response function β(x), where β(x) → U, then U... Di (t)=β(P Aj (t)).
[0076] refer to Figure 2 As shown, the derivative of variable x is introduced into the response function β(x). The replicative dynamic evolution process of evolutionary games can be described as follows: Figure 3 As shown. Figure 2 and Figure 3 There is a time delay T between the system input and output. Figure 2 This is a schematic diagram of a typical biological motion control system, where θ ref The baseline trajectory representing biological movement; θ actual S represents the actual trajectory of biological movement. error Indicates signal error; B c Indicates control of viscosity; K c Indicates control stiffness; F c This indicates negative feedback regulation; "controlled plant" indicates auxiliary control device; and "disturbances" indicates external disturbance.
[0077] refer to Figure 3 As shown, Figure 3 This is a schematic diagram of the hypothetical servo system for the dynamic evolution of replication, which combines the characteristics of a typical biological control system with the dynamic features of replication. Where, P D * This represents the evolutionary stable game belief of the defender in the Nash equilibrium. P represents the defender's game belief about the attacker. Aj and its derivative The joint response, where feedback represents the necessary information feedback during the iteration. Under conditions of incomplete information, the defender can obtain the empirical game belief P at time t through intrusion detection and other means. Aj (t)′, and can be derived from historical records. but
[0078]
[0079] Here, i and j are in a one-to-one correspondence. This indicates that the payoff can be further expressed as a joint response to the attacker's attack strategy and the short-term prediction of the attack strategy's defense rate. Under ideal conditions, the known attacker's game belief P can be considered as... AjWhile (t)′ is true and reliable, in actual network attack and defense, the defender's information acquisition and transmission are subject to various interferences and deceptions. For example, attackers may release false attack information, causing frequent false alarms in the IDS system; they may forge attack traffic similar to the network background traffic, causing the IDS to miss alarms; they may use hybrid attacks to confuse attack characteristics and mislead the judgment of attack types, etc. These attack tactics will all lead to the defender's acquisition of the attacker's experience and game-theoretic beliefs P. Aj (t)′ and the attacker's actual game belief P Aj (t) produces a non-negligible observation error e. Since e is unpredictable, this embodiment assumes that the defender's empirical game belief at any time t is the sum of the attacker's actual game belief and the observation error e(t), i.e.
[0080] P Aj (t)′=P Aj (t)+e(t) (4)
[0081] Among them, P Aj Let e(t) represent the attacker's actual game belief at time t, and e(t) represent the observation error at time t. To reduce the impact of the observation error e on the model's prediction results, this embodiment utilizes the characteristic of iterative evolution in evolutionary game models, which involves repeated iterations to dilute the observation error, allowing the model to approximate a reliable optimal solution. (Combined with the formula...) By combining the observation error with formula (4), the mapping of the replication dynamic equation payoff function is unified, resulting in...
[0082]
[0083] Therefore, the improved replication dynamic equation can be derived as follows:
[0084]
[0085] in, d(P Di ) / dt represents the game belief P Di Rate of change of P over time t; Aj (t)′ represents the attacker's experienced game-theoretic belief at time t; Represents the game theory belief P Aj Prediction within time t; β(x) represents the response function. This represents the defender's game belief P against the attacker at time t. Aj and its derivative A joint response; Let represent the average payoff of the defensive player set at time t; i and j correspond one-to-one. The calculation formulas include:
[0086]
[0087] It is worth mentioning that the credibility of the decision comes from the certainty and stability of the basis of the decision. The sufficient and necessary condition for the evolution stability of the heterogeneous group N is that N exists strict Nash equilibrium, and the sufficient condition for N to exist Nash equilibrium is that the response function β(x) is monotonic. According to the analysis of formulas (3), (4) and (5), it can be seen that the response function β(x) in the improved evolutionary game model with observation error is not a monotonic function, and the effectiveness of the model needs to be further explained as follows:
[0088] In this embodiment, starting from the Cauchy convergence criterion of the function, the stability of the convergence of the game model is proved, and the following definitions are given first:
[0089] Definition 2: Q * ={(P Aj * ,P Di * ) k ,k=1,2,…,N} is the ideal Nash equilibrium solution set of the evolutionary game of the heterogeneous group, and the solution in Q * can be listed.
[0090] Definition 3: P(t0)={(P Aj (t0),P Di (t0)) k ,k=1,2,…,N} is the actual game belief output solution set of the improved evolutionary game model with observation error, and t0 is the time required for evolution.
[0091] Definition 4: The ε-neighborhood is a closed interval with the ideal Nash equilibrium point (P Aj * ,P Di * ) k as the center and ε(ε>0) as the radius. When e(t)≡0, the short-term prediction is adjusted according to the feedback of historical experience, which does not affect the stability of the classical replicator dynamic equation. At this time, the game model converges to the ideal Nash equilibrium solution set Q * . When e(t)≠0, according to the Cauchy convergence criterion of the function, the following theorem is proposed:
[0092] Theorem 1: There is a real number δ, when |e(t)|≤δ, for any k, (P Aj (t0),P Di (t0)) k falls in (P Aj * ,PDi * ) k ε k - Within the neighborhood.
[0093] Proof: For any initial game belief pair (P) Aj ,P Di When there is an observation error, e(t), and e(t)≤δ, after time t0 and the evolution process corresponding to formulas (3), (4) and (5), the actual game belief output solution (P) can be obtained. Aj (t0),P Di (t0)), and (P) Aj (t0),P Di (t0) falls when e(t)≡0 (P) Aj ,P Di Within the ε-neighborhood of the ideal Nash equilibrium solution corresponding to ), then for enumerable (P) Aj * ,P Di * ) k There must be a corresponding (P) Aj ,P Di ) k δ k , (P Aj (t0),P Di (t0)) k and ε k Select δ < m k inδ k Due to δ k and (P) Aj ,P Di ) k , (P Aj * ,P Di * ) k , (P Aj (t0),P Di (t0)) k , ε k One-to-one correspondence, therefore when |e(t)|≤δ, |e(t)| lies in all δ k Inside, at this time (P) Aj (t0),P Di (t0)) k All fall on (P) Aj * ,P Di * ) k ε k -Within the neighborhood, the proof is complete.
[0094] According to the definition of convergence, it can be considered that the model converges at this time, and the obtained actual solution (P A (t0), P D (t0)) k is stable and reliable.
[0095] In step S103,
[0096] The Nash equilibrium solution of the classical evolutionary game generally includes pure strategy Nash equilibrium and mixed strategy Nash equilibrium. In the mixed strategy Nash equilibrium, different defense strategies appear in the form of probability, and such uncertainty makes the operability of the mixed strategy Nash equilibrium solution not strong in the actual network defense action decision. Therefore, based on the principle of deterministic decision, the embodiment designs an optimal defense pure strategy selection algorithm. This algorithm is calculated by means of the ode45 function in MATLAB, the ode45 function adopts the fourth-order-fifth-order Runge-Kutta algorithm, the truncation error is (Δx) 5 , and the time complexity is O(n·T). The overall time complexity of the algorithm is O(n 2 ·T), and in the actual network attack and defense, the number n of optional defense strategies and the step length T of the Runge-Kutta algorithm are the key factors to determine the time complexity of the algorithm. The optimal defense pure strategy set S D (Best) obtained by the algorithm can provide a scientific basis for the deterministic network defense decision and improve the reliability of the defense decision.
[0097] Simulation and analysis
[0098] In order to verify the effectiveness of the multi-stage network defense decision method based on the improved evolutionary game model in the embodiment, the following simulation experiment is carried out, and the simulation experiment results are analyzed in detail.
[0099] Referring to Figure 4 , it is a schematic diagram of the topological environment of the network information system in the simulation experiment. The simulation experiment environment mainly consists of network defense equipment, a WEB server, a database server, a file server and a user terminal. The experimental system is divided into three areas by the network defense equipment: the external network, the isolation area (Demilitarized Zone, DMZ) and the internal network, and the access control policy of the firewall is that the non-internal network users can only access the WEB server in the isolation area, the internal network users can access the database server, the file server and the WEB server. Relying on the specific experimental environment, the benefits of network attack and defense can be quantitatively expressed. The key parameters and calculation formula of benefit quantization in the embodiment are defined as follows:
[0100] Definition 5: Resource value V r, refers to the value of the resources that are contended in the network attack and defense. Generally, the value of the resources is considered to be equally important to both the attacker and the defender. In this embodiment, the fixed V r = 1.
[0101] Definition 6: Defense cost D cost , refers to the cost that the defender needs to pay for making targeted adjustments to make the attack of the attacker invalid. For example, the system overhead is increased, the quality of service is decreased, etc.
[0102] Definition 7: Supply cost A cost , refers to the cost that the attacker needs to pay when performing the attack. For example, the time cost of the attack, the risk cost, etc. In this embodiment, the attack cost is related to the threat level of the vulnerability. The higher the threat level of the vulnerability is, the lower the attack cost is.
[0103] Definition 8: Infection probability λ, refers to the probability that the attacker successfully infects the defender through the virus by exploiting the vulnerability. In this embodiment, the value of λ is defined as the vulnerability damage score provided by the China National Vulnerability Database of Information Security (CNNVD).
[0104] Definition 9: Defense effect φ, refers to the probability that the defender successfully makes the attack invalid by using the defense action.
[0105] Based on the above, in the complete network attack and defense game,
[0106] The income U of the defender D may be specifically represented as: U D = φ·V r -D cost (1),
[0107] The income of the attacker is derived from the income obtained after infecting the platform, which is related to the infection probability. Therefore, the income U of the attacker A may be specifically represented as: U A = λ·V r -A cost (2)
[0108] According to the topological structure characteristics of the network information system, combined with the vulnerability information provided by the CNNVD, this simulation experiment is aimed at two types of attack strategies, namely, the information integrity attack strategy S A1 aiming at the database and the web service, and the hijacking attack strategy S A2 aiming at the user terminal. Each type of attack strategy includes two atomic attack strategies S A1 = (a1, a2) and S A2 = (a3, a4), which are specifically shown in Table 1 as follows:
[0109] Table 1: Atomic attack strategy
[0110]
[0111] As can be seen from Table 1, a1 is a web resource management vulnerability attack, the infection probability is 0.78, and the attack cost is 0.20; a2 is an Oracle database input verification attack, the infection probability is 0.89, and the supply cost is 0.15; a3 is a Word plug-in path traversal attack, the infection probability is 0.93, and the attack cost is 0.10; a4 is a Microsoft Edge cross-site scripting attack, the infection probability is 0.73, and the attack cost is 0.25.
[0112] The corresponding defense strategy is the defense strategy S of the network D1 , and the defense strategy S of the user terminal D2 . Each type of defense strategy contains two atomic defense strategies S D1 = (b1, b2), S D2 = (b3, b4), as shown in the following Table 2:
[0113] Table 2: Atomic defense strategy
[0114]
[0115] As can be seen from Table 2, b1 is to set a black hole path, the operation cost is 0.30, and the defense effect is 0.59; b2 is to discard suspicious data packets, the operation cost is 0.10, and the defense effect is 0.25; b3 is to limit user activity, the operation cost is 0.50, and the defense effect is 0.83; b4 is to format the hard disk, the operation cost is 0.80, and the defense effect is 0.99.
[0116] When calculating the strategy benefit, it is considered that the strategy benefit is equal to the average benefit of the atomic attack and defense actions contained in the strategy. Combining formulas (1) and (2), the benefit quantization matrix of the attack and defense sides is given:
[0117]
[0118]
[0119] The disclosure will take a 2X2 symmetric game as an example. Through simulation experiments, first, the evolution solving process of the ideal Nash equilibrium solution set Q * is given, then the convergence stability of the actual solution set P(t0) when the observation error e(t) exists is analyzed, and finally the advantages of the algorithm and model of the disclosure in time complexity are compared with the classical model.
[0120] In a 2x2 symmetric game, the attacking side consists of two subgroups: N A1 and N A2 The defending side also contains two subgroups N. D1 and N D2 The attacker's corresponding pure strategy is: S A1 and S A2 The defending side's corresponding pure strategy is: S D1 and S D2 Taking the defensive side in a game as an example, the standardized payoff matrix can be expressed as:
[0121]
[0122] Where u1 is the attacker's pure strategy S A1 At that time, the defending side adopts a pure strategy S D1 The relative gains obtained; u2 is the attacker's pure strategy S A2 At that time, the defending side adopts a pure strategy S D2 The relative gains obtained. Substituting this into formula (3) yields the corresponding improved replication dynamic equations for the defender and attacker:
[0123]
[0124]
[0125]
[0126]
[0127] Experiment 1: Verify that the game model stably converges to Q when e(t)≡0. *
[0128] First, simulations verify that the model stably converges to Q when e(t)≡0. * , refer to Figure 5 As shown, the evolution trend of the game was analyzed using MATLAB experimental tools. In the simulation experiment, the values of u1 and u2 were adjusted multiple times, and it was found that this did not affect the convergence result of the stable solution. We set |u1|=0.4, |u2|=0.6, e(t)≡0, and the initial game belief P... A1 ,P D1 A random number within (0,1) Figure 5 The results correspond to 300 Monte Carlo simulation experiments. Figure 5 a = u1 = 0.4, u2 = 0.6; Figure 5 b is u1 = -0.4, u2 = 0.6; Figure 5 c = u1 = -0.4, u2 = -0.6; Figure 5d is u1=0.4, u2=-0.6.
[0129] From Figure 5 b and Figure 5 d, when u1*u2<0, the game belief does not change the sign in the state space, and from any initial position inside the state space, the overall state of the game will converge to a strictly dominant pure strategy. From 5a and Figure 5 c, when u1*u2>0, the game has two strictly pure strategy Nash equilibria and ideal mixed strategy Nash equilibria. According to formula (8), when the game converges to the ideal mixed strategy Nash equilibrium, P A1 (t)=u2 / (u1+u2), P D1 (t)=u2 / (u1+u2). The mixed strategy Nash equilibrium point of the game is unstable and will change with the change of u1 and u2 values, and since the mixed Nash equilibrium point is a saddle point, there is no actual solution trajectory, so in actual calculation the game result will not converge to the mixed strategy Nash equilibrium point. When u1*u2>0, the game converges to two stable strictly pure strategy Nash equilibria. The simulation results show that when e(t)=0, the game model stably converges to Q * , and all solutions in Q * are deterministic solutions, that is, in the ideal state without observation error, the game model can output stable and reliable pure strategies to support network defense decision-making.
[0130] Experiment 2: Verify that when e(t)≠0, the game model stably converges to the ε-neighborhood of Q * .
[0131] According to the attack and defense side revenue quantification results, u1=0.22, u2=0.26, set P A1 =0.6, P D1 =0.4, respectively, set δ1=1, δ2=0.1, δ3=0.01 three groups of control experiment group. |e(t)| in each group is a random number in (0, δ], and to observe the experimental results more clearly, the control curve of e(t)=0 in the group, Figure 6 the simulation results are given.
[0132] By comparing Figure 6 a, 6b and 6c, it can be seen that as the order of magnitude of δ decreases, the shock phenomenon in the early stage of model evolution gradually weakens, but when there is observation error, the model needs more evolution iterations to reach a stable convergence state. Within the range of δ≤1, no matter how large or small the value of δ is, the model can stably converge to the ε-neighborhood of the evolutionary equilibrium solution (0 * , 0 * ).
[0133] It should be particularly pointed out that the simulation experiment results show that when the evolution algebra ite→∞, ε→0. But due to the time sensitivity characteristics of network defense, it is necessary to output effective defense strategy in a limited time window, therefore, in the simulation experiment of the present disclosure, ite=300 is selected, at this time, ε<0.01, and the game model is stably converged to the ε-neighborhood of Q * .
[0134] Through theoretical analysis and simulation experiment, it is shown that under the condition that observation error exists, the model disclosed in the present disclosure can still stably output evolution results of approximate pure strategy. In a reasonable time window, the error order of magnitude can be controlled within 0.01%, and under the fierce confrontation situation of network attack and defense, it can be considered that the results of the model disclosed in the present disclosure are feasible and reliable.
[0135] Experiment 3: Performance comparison and analysis of different defense decision models in interference environment
[0136] The model disclosed in the present disclosure and some classical network defense decision models are all based on the simulation of decision process by using replication dynamics, and only differ in decision criteria and target application scenarios. Therefore, in the interference environment, three classical defense decision models are selected for comparison with the model disclosed in the present disclosure, to analyze the difference in time performance of algorithms used by different models.
[0137] The improved evolutionary game model IEGM of the present disclosure: based on the classical replication dynamics decision model, aiming at the interference environment and the time efficiency requirement of decision, the evolutionary game model is improved in complexity adaptability. The observation error e(t) and the short-term prediction of game belief are quantitatively defined. The optimal defense pure strategy selection algorithm is adopted to output the optimal defense pure strategy set S D (Best).
[0138] The network attack and defense evolutionary game model ADEGM: based on the limited rationality constraint, the non-cooperative evolutionary game process of attack and defense parties is simulated, and the optimal defense strategy selection algorithm of evolutionary stability is proposed. The advantage of the algorithm lies in the rigorous stability analysis of defense strategy, so that the evolutionary equilibrium strategy ESS output by the algorithm has strong stability and prediction ability.
[0139] The network attack and defense game model NADG: based on the incomplete information condition, aiming at the characteristics of military information network, an active defense strategy selection algorithm based on attack and defense game is proposed to output the optimal defense pure strategy. The advantage of the algorithm lies in that the active defense strategy selection is in the form of pure strategy, which effectively solves the problem that the defense strategy selection in the form of probability is not convenient to understand and operate, and also has certain advantages in efficiency because it does not need to solve all Nash equilibrium solutions.
[0140] The Improved Replicating Dynamic Attack-Defense Evolutionary Game Model (IADEGM) quantifies the strategy dependency effect among defensive groups based on the classic replicating dynamic equations and introduces an incentive coefficient λ. ij An improved copy dynamic attack-defense evolutionary game model is proposed. The advantage of the algorithm lies in its ability to accelerate strategy selection by increasing the incentive coefficient of the defending group.
[0141] In this experiment, three sets of interference experimental environments were set: δ1 = 1, δ2 = 0.1, and δ3 = 0.01. The experimental parameters for each environment are shown in Table 3 below.
[0142] Table 3: Experimental parameter settings in Experiment 3
[0143]
[0144] Due to the two defense strategies S in this simulation experiment D1 and S D2 There is no dependency relationship, therefore the excitation coefficient λ is set. 21 =1. The simulation results are shown in Table 4 below:
[0145] Table 4: Number of evolutionary algebras required to achieve the optimal solution for different models in Experiment 3
[0146]
[0147] Reference Figure 5 As shown, Figure 5 a is |e(t)|≤1; Figure 5 b is |e(t)|≤0.1; Figure 5 c is |e(t)|≤0.01. (From...) Figure 5 As shown, regardless of the order of magnitude of δ, the number of evolutionary generations required for the IEGM model to achieve stable output results is less than that of other comparative models. When δ = 0.01, the convergence speed of IEGM is improved by 3.29%-28.57% compared to other models; when δ = 0.1, the timeliness of IEGM can be improved by 4.26%-41.49%; when δ = 1, the improvement in the evolution speed of IEGM is very close to that of models such as NADG and IADEGM. It should be noted that, since there is no dependency between the available defense strategies of this disclosure, the incentive effect in the IADEGM model cannot be well reflected.
[0148] Simulation results show that, under strong interference, the model proposed in this disclosure can output an evolutionarily stable defense pure strategy faster than classic models such as ADEGM, NADG, and IADEGM, and better meet the time-sensitivity requirements of network attack and defense games.
[0149] In summary, the present disclosure aims at the problem that the traditional evolutionary game model is difficult to balance emergency response and error tolerance. Combining the active nerve-mediated segment servo system hypothesis in the biological motion control theory, an improved network security evolutionary game model is proposed, an optimal defense pure strategy selection algorithm considering factual bias is designed, and simulation verification is carried out combined with the simulation network attack and defense experimental environment.
[0150] When the observation error is 0, the improved evolutionary game model proposed by the present disclosure can accurately converge to the pure strategy Nash equilibrium, which improves the operability of network defense to a certain extent.
[0151] The improved network security evolutionary game model considering observation error cannot converge to the pure strategy Nash equilibrium, but will fall within the ε-neighborhood of the Nash equilibrium solution. And in the selected time window, ε<0.01, the feasibility and reliability of the obtained results are relatively high.
[0152] Within the given error range, the model shows good adaptability to errors of different orders of magnitude, and has certain advantages in response speed compared with other classical models.
[0153] It should be noted that: in the lock body model solving and example analysis of the present disclosure, it is assumed that the number of selectable strategies is 2, and the stability and applicability of the cover game model in the case of multi-dimensional game strategy space can be considered in the later research. In addition, in the present disclosure, the step length of short-term prediction is a fixed value of 1, and the time-varying step length adjustment method can be studied in the future, which can increase the search range of the optimal solution in the early stage and improve the stable accuracy of the optimal solution in the later stage.
[0154] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon considering the specification and practice of the subject disclosure. The present application is intended to cover any variations, uses or adaptive changes of the present disclosure that follow the general principles of the present disclosure and include known or customary technical methods in the art that are not disclosed by the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are indicated by the appended claims.
Claims
1. A multi-stage network defense decision-making method based on an improved evolutionary game model, characterized in that, The method comprises: An improved evolutionary game model is constructed, and parameters and steps in the improved evolutionary game model are initialized; According to an attack-defense dynamic game process in the improved evolutionary game model, a replicator dynamic equation is constructed and solved to obtain an evolutionary stable defense strategy; All the above steps are repeated to traverse and solve all the evolutionary stable defense strategies to obtain an optimal defense pure strategy set; The improved evolutionary game model includes a 7-element ordered group model IEGM wherein, , N denotes a set of attack-defense game players; , denotes a set of attack game players, denotes an attack game player individual; , denotes a set of defense game players, denotes a defense game player individual; , S denotes the set of mixed strategies of the attacker player, , denotes the set of pure strategies of the attacker player, denotes an individual pure strategy of the attacker player, , denotes the set of pure strategies of the defender player, denotes an individual pure strategy of the defender player, , P denotes the set of actual game beliefs of the attacker, , denotes the set of actual game beliefs of the defender, denotes the probability of the pure strategy being actually chosen by the attacker; , denotes the set of actual game beliefs of the defender, denotes the probability of the pure strategy being actually chosen by the defender; , denotes the set of beliefs of the attacker about the experience game, denotes the probability that the defender judges the choice of the attacker to be a pure strategy . , U denotes the set of payoffs of the game; , denotes the set of payoffs of the attacker, denotes the individual payoffs of the attacker's game participants adopting a pure strategy the expected payoff obtained in a one-stage game; , denotes the set of payoffs of the defender, denotes the individual payoffs of the defender's game participants adopting a pure strategy the expected payoff obtained in a one-stage game; represents the observation error; represents a set of short-term predictions; , represents a set of short-term predictions of the attacker's game belief, represents a short-term prediction for the game belief ; , represents a set of short-term predictions for the game belief , represents a set of short-term predictions of the attacker's game belief; The calculation formula of the replicator dynamic equation comprises: (3) wherein, = , denotes the game belief over time ; denotes the experienced game belief of the attacker at time ; denotes the game belief prediction over time ; denotes the response function, denotes the joint response of the defender to the game belief of the attacker and its derivative at time ; denotes the average payoff of the set of game participants of the defender at time ; i and j are in one-to-one correspondence; The attacker is in the experience game belief at the formula for calculating the formula for calculating (4) wherein, represents the actual game belief of the attacker at the time instant, represents the observation error at the time instant; The The average payoff of the set of defense players at the moment The formula for calculating the average payoff of the set of defense players at the moment includes: (5).
2. The multi-stage network defense decision method according to claim 1, characterized in that, When , the game has no influence on the pure strategy selection of both sides of the game. When , the game has a positive impact on the choice of pure strategies of both sides of attack and defense. When , the game has a negative impact on the pure strategy selection of both the attacker and the defender.
3. The multi-stage network defense decision method according to claim 1, characterized in that, The defensive game payoff is calculated by the formula comprises: (1) wherein, represents the resource benefit; represents the defense cost; represents the defense effect.
4. The multi-stage network defense decision method according to claim 1, characterized in that, The attacker game payoff is calculated by the formula comprises: (2) wherein, represents the resource benefit; represents the cost of the attacker; represents the infection probability.
5. The method of claim 1, wherein, In the step of constructing an improved evolutionary game model and initializing parameters and steps in the improved evolutionary game model, the steps are time delays between inputs and outputs in the improved evolutionary game model T .
6. The multi-stage cyber defense decision method of claim 1, wherein, The Observation error of time The value range includes Any random number in 7. The multi-stage cyber defense decision method of claim 1, wherein, the set of attack-defense game participants N There exists a strict Nash equilibrium, the response function is a monotone function.