Safety protection method and device based on multi-level multi-player safety master-slave game

By building a multi-level multi-player security master-slave game decision model and analyzing Stackelberg and Nash equilibrium, the problem of poor security protection effects in complex offensive and defense scenarios in the existing technology is solved, and stronger interpretability and practicality are achieved, and more reliable security protection strategies are provided.

CN120337483APending Publication Date: 2025-07-18TONGJI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510195080.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing security master-slave game decision-making model cannot fully simulate multi-level and multi-player participants and their operating modes in complex offensive and defense scenarios, resulting in poor security protection effects.

Method used

Build a multi-level multi-player security master-slave game decision model, analyze Stackelberg equilibrium and Nash equilibrium, determine consistency conditions, apply it to advanced continuous threat confrontation scenarios for security protection, introduce deep neural networks and reinforcement learning to predict attack behavior.

Benefits of technology

It improves the hierarchy and strategic model of the security protection system, enhances the explanatory and practicality of complex offensive and defense scenarios, provides a more reliable security decision theory, and realizes the dual functions of intrusion detection and fault repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337483A_ABST
    Figure CN120337483A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a security protection method and device based on a multi-level multi-player security master-slave game, and the method comprises the steps: constructing a multi-level multi-player security master-slave game decision model corresponding to a security protection system according to a large complex attack and defense scene facing the security protection system; analyzing Stackelberg equilibrium and Nash equilibrium of the model so as to solve a Stackelberg equilibrium strategy and a Nash equilibrium strategy; determining the consistency condition of the Stackelberg equilibrium strategy and the Nash equilibrium strategy of the model according to the disturbance of the model in practical application and the threat of uncertainty factors to the Stackelberg equilibrium reliability of the model; and setting the model according to the consistency condition, and applying the set model to an advanced continuous threat attack and defense scene which is actually faced by a security protection system for security protection. In this way, the security protection effect can be improved based on the multi-level multi-player security master-slave game decision model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of security protection, and particularly relates to a security protection method and device based on a multi-level multi-player security principal-agent game. Background Art

[0002] Security issues involve multiple fields such as engineering systems, networks, and privacy protection. Constructing a reasonable security decision-making model and performing security protection based on it can effectively defend against the risks and damages caused by attacks. Among them, the security principal-agent game decision-making model has attracted the attention of scholars in various fields. The principal-agent game was proposed by Von Stackelberg. This game consists of a leader player and a follower player. Among them, the leader has a higher dominant position and a first-mover advantage, while the follower's dominant position is relatively weak and the action is delayed after the leader. The security principal-agent game can simulate the differences in the internal action order, cognition, dominance, etc. between the defender and the attacker in the attack and defense scenario, can more vividly depict the hierarchical structure of the attack and defense scenario, and has significant practicality and interpretability.

[0003] In recent years, although existing research has expanded the scale of the security principal-agent game to make the model more suitable for real-world scenarios, the existing security principal-agent game decision-making model still has limitations in the application to actual attack and defense scenarios. On the one hand, real attack and defense scenarios are often intricate, and two-level and three-level principal-agent games cannot fully classify participants with different identities and different operation modes; on the other hand, the actions of players at the same hierarchical level may also have a sequential order, forming multiple sub-levels. In view of this, the current security protection based on the security principal-agent game decision-making model has the defect of poor effect. Summary of the Invention

[0004] In a first aspect, an embodiment of the present invention provides a security protection method based on a multi-level multi-player security principal-agent game, and the method includes:

[0005] Construct a multi-level multi-player security principal-agent game decision-making model corresponding to the security protection system according to the large and complex attack and defense scenario faced by the security protection system;

[0006] Analyze the Stackelberg equilibrium and Nash equilibrium of the constructed multi-level multi-player security principal-agent game decision-making model to solve the Stackelberg equilibrium strategy and Nash equilibrium strategy;

[0007] Determine the consistency condition of the Stackelberg equilibrium strategy and Nash equilibrium strategy of the multi-level multi-player security principal-agent game decision-making model according to the threat of the disturbance and uncertainty factors faced by the multi-level multi-player security principal-agent game decision-making model in actual application to the reliability of the Stackelberg equilibrium of the constructed multi-level multi-player security principal-agent game decision-making model;

[0008] Set up a multi - level multi - player secure master - slave game decision model according to the consistency condition, and apply the set multi - level multi - player secure master - slave game decision model to the advanced persistent threat confrontation and defense scenarios actually faced by the security protection system for security protection.

[0009] In some realizable ways of the first aspect, according to the large and complex attack - defense scenarios faced by the security protection system, construct a multi - level multi - player secure master - slave game decision model corresponding to the security protection system, including:

[0010] According to the large and complex attack - defense scenarios faced by the security protection system, combined with the multi - level multi - player secure master - slave game architecture, construct a multi - level multi - player secure master - slave game decision model corresponding to the security protection system.

[0011] In some realizable ways of the first aspect, the multi - level multi - player secure master - slave game architecture is described as:

[0012] The multi - level multi - player secure master - slave game generally presents a chain - series structure. Each player has a hierarchical level. The dominant position of the player increases sequentially with the increase of the hierarchical level, and the players act sequentially in the series direction. The highest hierarchical level of the player is called the senior leader, belonging to the defense side, which plays a key role in the operation of the security protection system. Its specific actions include prevention, detection, and response. The lowest hierarchical level of the player is the junior follower, belonging to the attack side, which decides the attack strategy after observing the situation of the security protection system and cannot obtain the internal information of the security protection system. A number of intermediate followers are introduced between the senior leader and the junior follower, belonging to the defense side, which assist the senior leader to consolidate and improve the security protection system. Its specific actions include fault repair and internal threat investigation.

[0013] In some realizable ways of the first aspect, in the multi - level multi - player secure master - slave game decision model, the intermediate and junior followers make decisions based on known information, that is, make the best response according to the determined strategies of the players with hierarchical levels higher than their own. In addition, the defense side needs to establish a dominant advantage over the attack side to anticipate the behavior of the attack side in advance. The process of establishing the defense side's dominant advantage includes: collecting attack behavior logs, system responses and states, historical attack - defense game data, and extracting them to obtain the characteristics related to the behavior of the attack side. Then, based on this, abstract the attack patterns and the behavior sequence of the attack side through a deep neural network and a long - short - term memory network, and on this basis, use the policy iteration process of reinforcement learning to model the best response strategy of the attack side.

[0014] In some realizable ways of the first aspect, the operation process of the multi - level multi - player secure master - slave game decision model is as follows:

[0015] Senior leaders can obtain the best responses of all intermediate and junior followers and incorporate them into their own utility functions for optimization to obtain the best strategies;

[0016] Each intermediate follower can observe the definite strategy of the senior leader and, relying on the dominant relative advantage, obtain the best response of the junior follower, and incorporate the best response of the junior follower into its own utility function for optimization to obtain the best strategy;

[0017] After obtaining the best strategies of the senior leader and the intermediate follower, the junior follower optimizes its own utility function based on this for decision-making.

[0018] In some realizable ways of the first aspect, analyze the Stackelberg equilibrium and Nash equilibrium of the constructed multi-level multi-player secure master-slave game decision model to solve the Stackelberg equilibrium strategy and Nash equilibrium strategy, including:

[0019] Calculate the best response of the junior follower, and based on the best response of the junior follower, calculate the best response of the intermediate follower. Then, based on the best responses of the junior and intermediate followers, calculate the best response of the senior leader to obtain the Stackelberg equilibrium strategy of the senior leader. After that, substitute the Stackelberg equilibrium strategy of the senior leader into the best responses of the junior and intermediate followers to obtain the Stackelberg equilibrium strategies of the junior and intermediate followers;

[0020] Regard the junior and intermediate followers and the senior leader as having equal dominant status and acting synchronously, so as to calculate the Nash equilibrium strategies of the junior and intermediate followers and the senior leader.

[0021] In some realizable ways of the first aspect, according to the threats to the reliability of the Stackelberg equilibrium of the constructed multi-level multi-player secure master-slave game decision model caused by perturbations and uncertainty factors faced in the actual application of the multi-level multi-player secure master-slave game decision model, determine the consistency conditions of the Stackelberg equilibrium strategy and Nash equilibrium strategy of the multi-level multi-player secure master-slave game decision model, including:

[0022] Calculate the utility function gradients of the senior leader and the intermediate follower in pursuing the Stackelberg equilibrium strategy;

[0023] Calculate the utility function gradients of the senior leader and the intermediate follower in pursuing the Nash equilibrium strategy;

[0024] Based on the concavity and monotonic interval of the utility function, analyze the positivity and negativity of the calculated utility function gradients within the strategy set constraints to obtain the necessary and sufficient conditions for the consistency of the Stackelberg equilibrium strategy and Nash equilibrium strategy of the senior leader and the intermediate follower.

[0025] In some realizable ways of the first aspect, setting a multi-level multi-player secure leader-follower game decision model according to consistency conditions includes:

[0026] Substitute the utility functions of the high-level leader and the mid-level follower into the utility function gradients and necessary and sufficient conditions of the high-level leader and the mid-level follower in pursuing the Stackelberg equilibrium strategy and the Nash equilibrium strategy, and deduct through numerical calculation to obtain the adjustable parameters in the utility functions of the high-level leader and the mid-level follower that satisfy the necessary and sufficient conditions, and set the adjustable parameters for the high-level leader and the mid-level follower so that the Stackelberg equilibrium strategy of the high-level leader and the mid-level follower is consistent with the Nash equilibrium strategy.

[0027] In a second aspect, an embodiment of the present invention provides a security protection device based on a multi-level multi-player secure leader-follower game, and the device includes:

[0028] A construction module, configured to construct a multi-level multi-player secure leader-follower game decision model corresponding to the security protection system according to the large and complex attack and defense scenarios faced by the security protection system;

[0029] An analysis module, configured to analyze the Stackelberg equilibrium and the Nash equilibrium of the constructed multi-level multi-player secure leader-follower game decision model to solve the Stackelberg equilibrium strategy and the Nash equilibrium strategy;

[0030] A determination module, configured to determine the consistency conditions of the Stackelberg equilibrium strategy and the Nash equilibrium strategy of the multi-level multi-player secure leader-follower game decision model according to the threats to the reliability of the Stackelberg equilibrium of the constructed multi-level multi-player secure leader-follower game decision model caused by perturbations and uncertain factors faced during the actual application of the multi-level multi-player secure leader-follower game decision model;

[0031] An application module, configured to set the multi-level multi-player secure leader-follower game decision model according to the consistency conditions, and apply the set multi-level multi-player secure leader-follower game decision model to the advanced persistent threat confrontation attack and defense scenario actually faced by the security protection system for security protection.

[0032] In a third aspect, an embodiment of the present invention provides an electronic device, and the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0033] Fourthly, an embodiment of the present invention provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the method as described above.

[0034] Compared with the prior art, the present invention has at least the following technical effects:

[0035] 1. The construction and proposal of the multi-level multi-player security master-slave game decision model make the security game have stronger interpretability and practicability for larger-scale and more complex attack and defense scenarios. It refines the single defender level in the existing security game, making the modeling of the security protection system more hierarchical, including more strategy modes, and can simulate more operation details of the actual engineering system.

[0036] 2. Studying the Stackelberg equilibrium and Nash equilibrium solutions of the multi-level multi-player security master-slave game provides a theoretical basis and guidance for the security decision-making of the security protection system, and greatly improves the security protection effect.

[0037] 3. Studying the sufficient and necessary conditions for the consistency of the Stackelberg equilibrium strategy and Nash equilibrium strategy of the multi-level multi-player security master-slave game makes the equilibrium strategy adopted by the security protection system in the face of disturbances and uncertain factors more robust and reliable.

[0038] 4. Studying the advanced persistent threat countermeasure security protection system based on the multi-level multi-player security master-slave game decision model, applying the model constructed by the present invention to the actual attack and defense scenario, enabling the security protection system to have the dual functions of intrusion detection and resistance and fault repair, and when the relevant parameters set by the security protection system meet the sufficient and necessary conditions for the consistency of the two equilibrium strategies, the abnormal behavior of the attacker taking synchronous actions with the defender will not cause large fluctuations and impacts on the strategy selection and utility of the senior leader.

[0039] It should be understood that the content described in the summary of the invention is not intended to limit the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Combined with the drawings and referring to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present invention will become more obvious. The drawings are used to better understand the present invention and do not constitute a limitation to the present invention. In the drawings, the same or similar reference numerals represent the same or similar elements, where:

[0041] Figure 1 is a flowchart of a security protection method based on a multi-level multi-player security master-slave game provided by an embodiment of the present invention;

[0042] Figure 2Schematic diagram of a multi-level multi-player secure master-slave game architecture provided by an embodiment of the present invention;

[0043] Figure 3 Schematic diagram of the Flipit game with a fault fixer provided by an embodiment of the present invention;

[0044] Figure 4 For different L provided by an embodiment of the present invention X Schematic diagram of the Stackelberg equilibrium strategy and Nash equilibrium strategy of leader X under different L coefficients;

[0045] Figure 5 For different L when the attacker adopts the Nash equilibrium strategy provided by an embodiment of the present invention X Schematic diagram of the utility of the senior leader under different L coefficients;

[0046] Figure 6 Structural diagram of a security protection device based on a multi-level multi-player secure master-slave game provided by an embodiment of the present invention;

[0047] Figure 7 Structural diagram of an exemplary electronic device capable of implementing the embodiments of the present invention. Detailed implementation manners

[0048] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] In addition, the term "and / or" in the present invention is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the associated objects before and after.

[0050] To solve the technical problems in the background art, the embodiments of the present invention provide a security protection method, device, equipment and storage medium based on a multi-level multi-player secure master-slave game. The following will describe in detail a security protection method, device, equipment and storage medium based on a multi-level multi-player secure master-slave game provided by the embodiments of the present invention with reference to the accompanying drawings through specific embodiments.

[0051] Figure 1 Flowchart of a security protection method based on a multi-level multi-player secure master-slave game provided by an embodiment of the present invention, asFigure 1 As shown in Figure 1 , the security protection method 100 may include:

[0052] S110. According to the large and complex attack and defense scenarios faced by the security protection system, construct a multi-level and multi-player security master-slave game decision-making model corresponding to the security protection system.

[0053] S120. Analyze the Stackelberg equilibrium and Nash equilibrium of the constructed multi-level and multi-player security master-slave game decision-making model to solve the Stackelberg equilibrium strategy and Nash equilibrium strategy.

[0054] S130. According to the threats to the reliability of the Stackelberg equilibrium of the constructed multi-level and multi-player security master-slave game decision-making model caused by perturbations and uncertainty factors faced by the multi-level and multi-player security master-slave game decision-making model in actual applications, determine the consistency conditions for the Stackelberg equilibrium strategy and Nash equilibrium strategy of the multi-level and multi-player security master-slave game decision-making model.

[0055] S140. Set the multi-level and multi-player security master-slave game decision-making model according to the consistency conditions, and apply the set multi-level and multi-player security master-slave game decision-making model to the advanced persistent threat confrontation attack and defense scenario actually faced by the security protection system for security protection.

[0056] For the convenience of further understanding, the above steps will be described in detail below with specific embodiments:

[0057] a) Construction of the multi-level and multi-player security master-slave game decision-making model

[0058] Specifically, according to the large and complex attack and defense scenarios faced by the security protection system, combined with the multi-level and multi-player security master-slave game architecture, construct a multi-level and multi-player security master-slave game decision-making model corresponding to the security protection system.

[0059] Among them, the multi-level and multi-player security master-slave game architecture can be as Figure 2 shown and expressed as:

[0060] The multi-level multi-player security master-slave game generally presents a chain-tandem structure. Each player has a hierarchical level, and the dominant position of the player increases sequentially with the increase of the hierarchical level. Moreover, the players act sequentially in the tandem direction. The highest hierarchical level of the player is called the senior leader, who belongs to the defense side and plays a key role in the operation of the security protection system. Its specific actions include prevention, detection, and response. The lowest hierarchical level of the player is the junior follower, who belongs to the attack side. After observing the situation of the security protection system, it decides on the attack strategy and cannot obtain the internal information of the security protection system. Multiple intermediate followers are introduced between the senior leader and the junior follower. They belong to the defense side and assist the senior leader in consolidating and improving the security protection system. Their specific actions include fault repair and internal threat investigation.

[0061] When constructing the decision-making model of the multi-level multi-player security master-slave game, the senior leader is denoted as X, and the n intermediate followers are denoted as Y = {Y1, Y2, …, Y n}, and the junior follower is denoted as Z. The strategies of the senior leader, intermediate followers, and junior follower are represented by respectively, where, Ω x , Ω yi (i = 1, …, n), Ω z represent the strategy set constraints of the three types of players respectively.

[0062] Combining the success rates, costs, and losses of the defense side and the attack side in the actual attack and defense scenarios, the utility functions of the senior leader, intermediate followers, and junior follower can be expressed as:

[0063] U X (x, y1, …, y n , z) = B(x) + f x (y1, …, y n , z)x(1)

[0064]

[0065] U Z (x, y1, …, y n , z) = f z (x, y1, …, y n , z)(3)

[0066] Among them, U X (·), U Z (·) represent the utility functions of the senior leader, intermediate followers, and junior follower respectively; B(·), f x (·), f z(·) are all real-valued functions. It is worth noting that the utility function is usually used to measure the degree of goal achievement of participants (attackers and defenders), expressing the impact of each decision on the interests or goals of the participants.

[0067] In the multi-level multi-player security master-slave game decision-making model, the intermediate and low-level followers make decisions based on known information, that is, they make the best response according to the determined strategies of the players at a higher level than their own. In addition, in order to obtain a more reasonable strategy, the defender needs to establish a dominant advantage over the attacker in order to predict the attacker's behavior in advance. The process of establishing the dominant advantage of the defender includes: collecting attack behavior logs, system responses and status, historical attack and defense game data, and extracting them to obtain features related to the attacker's behavior (such as attack frequency, attack duration, attack intensity, attacker's changing trend, historical behavior pattern, etc.), and then abstracting the attack mode and attacker behavior sequence through deep neural networks and long short-term memory networks. On this basis, the strategy iteration process of reinforcement learning is used to model the attacker's best response strategy. Relying on the above process, the defender can learn the attributes of the attacker's action mode, motivation, utility, purpose, etc., so as to gain the upper hand in the confrontation game.

[0068] Therefore, the operation process of the multi-level multi-player security master-slave game decision model can be summarized as follows:

[0069] 1. Senior leader X can obtain the best responses of all mid-level and low-level followers and incorporate them into its own utility function to optimize and obtain the best strategy;

[0070] 2. Each intermediate follower Y i (i=1,…,n) The determined strategy of the senior leader can be observed, and the best response of the low-level followers can be known by virtue of the dominant relative advantage. The best response of the low-level followers can be incorporated into the utility function to optimize and obtain the best strategy;

[0071] 3. After obtaining the best strategies from senior leaders and mid-level followers, low-level follower Z optimizes its own utility function based on this to make decisions.

[0072] b) Stackelberg equilibrium strategy and Nash equilibrium strategy solution

[0073] Analyze the stable state of the multi-level multi-player security master-slave game decision model, that is, the equilibrium solution, which is the set of security strategies of all players. First, analyze the Stackelberg equilibrium strategy of the model.

[0074] For a low-level follower Z, for any {x,y1,…,y n The best response for} is:

[0075]

[0076] For intermediate follower Y n , based on the best response of lower-level follower Z, its best response to any {x, y1, …, y n―1} is:

[0077]

[0078] Similarly for intermediate follower Y i (i = 1, …, n−1), it can also obtain the best responses of lower-level players by virtue of its dominant relative advantage, and consider the influence of these best responses on its own utility function to obtain a more rational strategy. Therefore, the best response of intermediate follower Y i only depends on {x, y1, …, y i―1}(i > 1), and the best response of intermediate follower Y1 only depends on the strategy x of leader X. Denote as the best response vector of players {Y i deduced by the higher-level player according to its own strategy and the strategies of players {Y i+1 , …, Y n , Z}, that is, only depends on {x, y1, …, y i}. Thus, the best response of intermediate follower Y i (i = 1, …, n) to any {x, y1, …, y i―1} is:

[0079]

[0080] where the best response of Y1 is directly expressed as

[0081] The highest-ranking senior leader X can consider the best responses of all followers when making decisions. Denote as the best responses of all intermediate and lower-level followers deduced by X according to its own strategy x. Thus, the Stackelberg strategy of senior leader X (denoted as x SE ) can be calculated by the following formula:

[0082]

[0083] After obtaining the Stackelberg strategy of senior leader X, intermediate and lower-level followers substitute the Stackelberg strategy of senior leader X into their own best response expressions (4) and (6) successively according to the multi-level multi-player secure master-slave game series structure to obtain the corresponding Stackelberg strategies, and finally obtain the strategy set This is the Stackelberg equilibrium strategy of the constructed model.

[0084] When the dominant position of the senior leader X is threatened, the player may abandon the Stackelberg equilibrium and pursue the Nash equilibrium commonly used in the game. The Nash equilibrium strategy of the constructed model is obtained from the following formula:

[0085]

[0086] Each player has an equal dominant position and acts synchronously when pursuing the Nash equilibrium. When the players are in the Nash equilibrium, no player can obtain more benefits by unilaterally changing their own strategy.

[0087] To ensure the existence of the equilibrium of the constructed multi-level multi-player secure leader-follower game decision-making model, the strategy sets, utility functions, and best responses of the players need to satisfy the following conditions:

[0088] 1. Ω x , Ω z are both non-empty convex compact sets.

[0089] 2. B(x) is continuously differentiable with respect to x and is a concave function; f x (y1, y2, …, y n , z) is continuously differentiable with respect to y1, …, y n , z.

[0090] 3. For any are all continuously differentiable with respect to x, y1, …, y n , z, and is a concave function with respect to y i .

[0091] 4. f z (x, y1, …, y n , z) are all continuously differentiable with respect to x, y1, …, y n , z, and is a concave function with respect to z.

[0092] 5. The best responses are all continuously differentiable of the first order.

[0093] The above two equilibrium solutions will be obtained through optimization methods and computational tools. For small-scale game problems, equations (4)-(7) and equation (8) are arranged into a system of equations for the corresponding equilibrium, and combined with the constraint conditions corresponding to the strategy variables, and solved through tool functions such as fsolve, vpasolve, and fmincon in Matlab. Among them, fsolve and fmincon can handle nonlinear functions and systems of equations; for large-scale and more complex game problems, iterative methods such as the gradient descent algorithm are used to find the equilibrium solution, calculate the partial derivative of the utility function of each player or the utility function after substituting the best response of the lower-level players with respect to its own strategy variable, and subtract the product of the step size and the partial derivative from the previous-round strategy value in each iteration to obtain the latest strategy until convergence.

[0094] c) Determination of the consistency conditions for Stackelberg equilibrium and Nash equilibrium strategies

[0095] When the multi-level multi-player security leader-follower game decision model is applied to actual attack and defense scenarios, the reliability of Stackelberg equilibrium may be threatened by many internal and external disturbance factors. Especially in attack and defense scenarios, as the attacking party of the lower-level follower, it is often unwilling to be in a dominant and disadvantaged position, and may act synchronously with the defense party, or master the internal information of the defense party through various channels to improve its own cognitive ability. At this time, it is another alternative for the defense party to consider Nash equilibrium. However, the defense party's adoption of Nash equilibrium strategy (that is, giving up the leading initiative) may reduce its utility, and even in the scenario where the attacking party does not break the leader-follower structure, the defense party's direct adoption of Nash equilibrium strategy may lead to serious consequences; while adhering to the Stackelberg equilibrium strategy will also make the defense party may fall into the dilemma of the above disturbance factors. Facing the dual choices, if the two equilibrium strategies of the defense party are consistent, the defense party does not need to consider the influence of disturbance factors, and the stability and robustness of the constructed multi-level multi-player security leader-follower game are also improved. Therefore, it is crucial to analyze the consistency conditions of the two equilibrium strategies of the defense party.

[0096] To establish the connection between different equilibrium strategies, first calculate the gradients of the utility functions of the high-level leader and the middle-level follower. During the calculation of the Stackelberg equilibrium, when the lower-level follower takes the best response, the utilities of the high-level leader and the middle-level follower are respectively expressed as Their gradients are respectively denoted as:

[0097]

[0098] During the process of the defense party pursuing the Nash equilibrium strategy, when other players adopt the Nash equilibrium, find the gradient of its original utility function, denoted as:

[0099]

[0100] The Stackelberg equilibrium strategies for the senior leader and the intermediate follower are extreme points or the boundaries of the strategy set constraints of SE ), there exists a neighborhood δ(x satisfying: is monotonically increasing in the left neighborhood interval δ SE (x ― ), and is monotonically decreasing in the right neighborhood interval δ SE ; SE (x + ) is monotonically increasing in the left neighborhood interval SE ) and monotonically decreasing in the right neighborhood interval In left neighborhood interval it is monotonically increasing, and in right neighborhood interval it is monotonically decreasing.

[0101] Based on the concavity of the utility function and the monotonic intervals, analyzing the positivity and negativity of the gradient within the strategy set constraints, it can be obtained that the necessary and sufficient conditions for the Stackelberg equilibrium strategies of the defender's senior leader X and intermediate follower Y i to be consistent with the Nash equilibrium strategies are:

[0102] 1. When any y in the interval i satisfies , the Stackelberg equilibrium strategy of the intermediate follower Y i is also the Nash equilibrium strategy.

[0103] 2. When any x in the interval δ(x SE ) ∩ rint(Ω x ) satisfies D x (x SE ) = 0 or T x (x) · D x (x) > 0, the Stackelberg equilibrium strategy of the leader X is also the Nash equilibrium strategy, where rint(·) represents the relative interior point.

[0104] d) Application of the multi-level multi-player secure leader-follower game decision model

[0105] Substitute the utility functions of the senior leader and the intermediate follower into equations (9)-(12) and the necessary and sufficient conditions, and deduce the adjustable parameters in the utility functions of the senior leader and the intermediate follower that satisfy the necessary and sufficient conditions through numerical calculation. Set the adjustable parameters for the senior leader and the intermediate follower so that the Stackelberg equilibrium strategy of the senior leader and the intermediate follower is consistent with the Nash equilibrium strategy, improving the robustness of the model and reducing the risk of decision-making errors. After that, apply the established multi-level multi-player secure master-slave game decision model to the actual advanced persistent threat confrontation and defense scenario in the security protection system for security protection.

[0106] As an example, applying the multi-level multi-player secure master-slave game decision model to the actual advanced persistent threat confrontation and defense scenario for security protection can be shown as follows:

[0107] In the face of advanced persistent threats, existing security games are often modeled as the Flipit game:

[0108]

[0109] Among them, X is the defender, acting as the senior leader in the master-slave game, and Z is the attacker who launches the advanced persistent threat, acting as the junior follower in the master-slave game. x represents the defense frequency, z represents the attack frequency, and L X ,L Z respectively represent the cost coefficients for the defender and the attacker to initiate actions. Both the defender and the attacker are committed to maximizing their occupied time and control rights during the persistent threat period. The present invention will introduce multiple fault repairers {Y1,…,Y n} as intermediate followers. These fault repairers can promptly repair damaged and infected defense targets to block the secondary lateral spread of infections and improve the security and efficiency of the defense system. Thus, the utility functions of all players in the multi-level multi-player secure master-slave Flipit game in the face of advanced persistent threats are expressed as:

[0110]

[0111] Among them, ρ represents the occupancy rate of the fault repairer during the attack period, and y i represents the repair progress of the i-th fault repairer. In the actual scenario, due to the limited capabilities of a single fault repairer, multiple fault repairers observe the real-time repair progress and carry out repair work in sequence until the faulty equipment starts to operate. Each player is committed to maximizing their own utility function. Figure 3 is the action process of the players in the extended Flipit game when facing advanced persistent threat attacks.

[0112] Set L Z = 1.85, ρ = 0.3, Ω x= [0.1, 1], Ω z = [0.1, 2], the repair target of the faulty equipment is 70%. For this attack and defense scenario, a simulation experiment is conducted. It is set that the repair target is reached when the 4th faulty repairer is introduced. Randomize the repair capabilities of the 4 faulty repairers, where are all subsets of [0, 0.7].

[0113] When given x and y n , U can be obtained Z The extreme point z of ◇ is:

[0114]

[0115] According to the value ranges of x and y n , the value range of z ◇ can be calculated as [0.16, 1.73], which is within Ω z range. Therefore, in this experimental scenario, BR Z (x, y n ) = z ◇ to ensure the continuous differentiability of the best response function. Thus, T x (x) and D x (x) can be calculated as:

[0116]

[0117] According to the necessary and sufficient conditions for the consistency of the Stackelberg equilibrium strategy and the Nash equilibrium strategy derived, when L X ≤ 0.14 or L X ≥ 0.91, the consistency can be satisfied. Adjust the value of the adjustment coefficient L X . In this scenario, the changes in the Stackelberg equilibrium strategy (x SE ) and the Nash equilibrium strategy (x NE ) of the senior leader X are as shown in Figure 4 .

[0118] It can be found from the figure that when L X < 0.15 or L X ≥ 1, the consistency of the two equilibrium strategies can be satisfied. Thus, the senior leader can improve the robustness of the strategy against several perturbations and uncertainties by adjusting its own cost coefficient, and solve the dilemma of choosing different equilibrium schemes.

[0119] Further simulate the abnormal scenario where the attacker does not follow the master - slave game framework and acts synchronously with the defender, that is, the attacker adopts the Nash equilibrium strategy z NE . The utilities of the senior leader under different cost coefficients in this scenario are as shown in Figure 5 .

[0120] Depend on Figure 5 It can be seen that when L X =0.2, that is, when the necessary and sufficient conditions for consistency are not met, if the senior leader insists on using the Stackelberg equilibrium strategy when the attacker destroys the master-slave game architecture, he will gain less benefits. If he switches to the Nash equilibrium strategy, the utility will increase but still be at a low level. X =0.1, no matter which equilibrium strategy the senior leader adopts, its utility is consistent and significantly higher than L X =0.2. This shows that making the constructed multi-level multi-player security master-slave game decision model meet the necessary and sufficient conditions for the consistency of the two equilibrium strategies can improve the decision-making stability and robustness of the defending players when facing abnormal actions of attackers.

[0121] In summary, the present invention achieves at least the following technical effects:

[0122] 1. The construction and proposal of a multi-level multi-player security master-slave game decision model makes security games more interpretable and practical for larger-scale and more complex attack and defense scenarios. It refines the single defender level in the existing security game, making the modeling of the security protection system more hierarchical, including more strategy modes, and can simulate more operational details of the actual engineering system.

[0123] 2. Study the Stackelberg equilibrium and Nash equilibrium solutions of multi-level and multi-player security master-slave games, provide theoretical basis and guidance for security decision-making of security protection systems, and greatly improve the security protection effect.

[0124] 3. Study the necessary and sufficient conditions for the consistency of the Stackelberg equilibrium strategy and the Nash equilibrium strategy in the multi-level and multi-player security master-slave game, so as to make the equilibrium strategy adopted by the security protection system more robust and reliable when facing disturbances and uncertain factors.

[0125] 4. Research on an advanced persistent threat countermeasure security protection system based on a multi-level multi-player security master-slave game decision model, apply the model constructed by the present invention to actual attack and defense scenarios, so that the security protection system has the dual functions of intrusion detection and resistance, and fault repair. When the relevant parameters set by the security protection system meet the necessary and sufficient conditions for the consistency of the two equilibrium strategies, the abnormal behavior of the attacker taking synchronous actions with the defender will not cause significant fluctuations and impacts on the strategic choices and effectiveness of senior leaders.

[0126] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0127] The above is the introduction of the method embodiments. The following further illustrates the solution of the present invention through device embodiments.

[0128] Figure 6 The following is a structural diagram of a security protection device provided by an embodiment of the present invention based on a multi-level multi-player secure Stackelberg game, as Figure 6 shown. The security protection device 600 may include:

[0129] A construction module 610, configured to construct a multi-level multi-player secure Stackelberg game decision model corresponding to the security protection system according to the large and complex attack and defense scenarios faced by the security protection system;

[0130] An analysis module 620, configured to analyze the Stackelberg equilibrium and Nash equilibrium of the constructed multi-level multi-player secure Stackelberg game decision model to solve the Stackelberg equilibrium strategy and Nash equilibrium strategy;

[0131] A determination module 630, configured to determine the consistency condition between the Stackelberg equilibrium strategy and the Nash equilibrium strategy of the multi-level multi-player secure Stackelberg game decision model according to the threats to the reliability of the Stackelberg equilibrium of the constructed multi-level multi-player secure Stackelberg game decision model caused by perturbations and uncertainty factors faced in actual applications of the multi-level multi-player secure Stackelberg game decision model;

[0132] An application module 640, configured to set the multi-level multi-player secure Stackelberg game decision model according to the consistency condition, and apply the set multi-level multi-player secure Stackelberg game decision model to the advanced persistent threat confrontation attack and defense scenario actually faced by the security protection system for security protection.

[0133] It can be understood that Figure 6 each module / unit in the security protection device 600 shown has the functions of implementing Figure 1 each step in the security protection method 100 shown, and can achieve its corresponding technical effects. For the sake of brevity, it will not be elaborated here.

[0134] Figure 7It is a structural diagram of an exemplary electronic device capable of implementing the embodiments of the present invention. The electronic device 700 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 700 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown in the present invention, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed in the present invention.

[0135] As Figure 7 shown, the electronic device 700 may include a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0136] A plurality of components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0137] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 may be implemented as a computer program product, including a computer program tangibly embodied in a computer-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of method 100 described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute method 100 by any other suitable means (e.g., by means of firmware).

[0138] The various embodiments described above in the present invention can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0139] The program code for implementing the methods of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0140] In the context of the present invention, a computer-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0141] It should be noted that the present invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute method 100 and achieve the corresponding technical effects achieved by the method of the embodiments of the present invention. For the sake of concise description, details are not repeated herein.

[0142] In addition, the present invention also provides a computer program product, which includes a computer program that implements method 100 when executed by a processor.

[0143] It should be understood that various forms of the processes shown above may be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. The present invention places no restrictions herein.

[0144] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A security protection method based on multi-level multi-player secure master-slave game, characterized in that The method includes: Constructing a multi-level multi-player security leader-follower game decision model corresponding to the security protection system according to the large-scale complex attack and defense scenarios faced by the security protection system; Analyzing the Stackelberg equilibrium and Nash equilibrium of the constructed multi-level multi-player security leader-follower game decision model to solve the Stackelberg equilibrium strategy and Nash equilibrium strategy; Determining the consistency conditions of the Stackelberg equilibrium strategy and Nash equilibrium strategy of the multi-level multi-player security leader-follower game decision model according to the threats to the reliability of the Stackelberg equilibrium of the constructed multi-level multi-player security leader-follower game decision model caused by perturbations and uncertainty factors faced in actual applications; Setting the multi-level multi-player security leader-follower game decision model according to the consistency conditions, and applying the set multi-level multi-player security leader-follower game decision model to the advanced persistent threat confrontation attack and defense scenarios actually faced by the security protection system for security protection.

2. The method according to claim 1, wherein The constructing a multi-level multi-player security leader-follower game decision model corresponding to the security protection system according to the large-scale complex attack and defense scenarios faced by the security protection system includes: Constructing a multi-level multi-player security leader-follower game decision model corresponding to the security protection system according to the large-scale complex attack and defense scenarios faced by the security protection system and combining with the multi-level multi-player security leader-follower game architecture.

3. The method according to claim 2, wherein The multi-level multi-player security leader-follower game architecture is described as: The multi-level multi-player security leader-follower game generally presents a chain series structure. Each player has a hierarchical level, and the dominant position of the player increases sequentially with the increase of the hierarchical level, and the players act sequentially in the series direction; the highest hierarchical level of the player is called the senior leader, belonging to the defense side, which plays a key role in the operation of the security protection system, and its specific actions include prevention, detection and response; the lowest hierarchical level of the player is the junior follower, belonging to the attack side, which decides the attack strategy after observing the situation of the security protection system and cannot obtain the internal information of the security protection system; several intermediate followers are introduced between the senior leader and the junior follower, belonging to the defense side, which assist the senior leader to consolidate and improve the security protection system, and its specific actions include fault repair and internal threat investigation.

4. The method according to claim 3, wherein In the multi-level multi-player security leader-follower game decision model, the intermediate and junior followers make decisions based on known information, that is, make the best response according to the determined strategies of the players with hierarchical levels higher than their hierarchical levels. In addition, the defense side needs to establish a dominant advantage over the attack side to anticipate the behavior of the attack side in advance. The process of establishing the dominant advantage of the defense side includes: collecting attack behavior logs, system responses and states, historical attack and defense game data, and extracting them to obtain the characteristics related to the behavior of the attack side, and then abstracting the attack patterns and attack side behavior sequences based on this through deep neural networks and long short-term memory networks, and on this basis, using the policy iteration process of reinforcement learning to model the best response strategy of the attack side.

5. The method according to claim 4, wherein The operation process of the multi-level multi-player security leader-follower game decision model is: The senior leader can obtain the best responses of all intermediate and junior followers and incorporate them into its own utility function for optimization to obtain the best strategy; Each intermediate follower can observe the definite strategy of the senior leader and, relying on the dominant relative advantage, obtain the best response of the junior follower, and incorporate the best response of the junior follower into its own utility function for optimization to obtain the best strategy; After obtaining the best strategies of the senior leader and the intermediate follower, the junior follower optimizes its own utility function based on this for decision-making.

6. The method according to claim 5, characterized in that, The Stackelberg equilibrium and Nash equilibrium of the multi-level multi-player secure master-slave game decision model constructed by the above analysis are used to solve the Stackelberg equilibrium strategy and Nash equilibrium strategy, including: Calculating the best response of the junior follower, and calculating the best response of the intermediate follower according to the best response of the junior follower, and then calculating the best response of the senior leader according to the best responses of the junior and intermediate followers to obtain the Stackelberg equilibrium strategy of the senior leader. Then, substituting the Stackelberg equilibrium strategy of the senior leader into the best responses of the junior and intermediate followers to obtain the Stackelberg equilibrium strategies of the junior and intermediate followers; Regarding the junior and intermediate followers and the senior leader as having equal dominant status and acting synchronously, so as to calculate the Nash equilibrium strategies of the junior and intermediate followers and the senior leader.

7. The method according to claim 6, wherein Regarding the threats to the reliability of the Stackelberg equilibrium of the multi-level multi-player secure master-slave game decision model constructed by the perturbations and uncertainty factors faced by the multi-level multi-player secure master-slave game decision model in practical applications, determining the consistency conditions of the Stackelberg equilibrium strategy and Nash equilibrium strategy of the multi-level multi-player secure master-slave game decision model, including: Calculating the utility function gradients of the senior leader and intermediate follower in pursuing the Stackelberg equilibrium strategy; Calculating the utility function gradients of the senior leader and intermediate follower in pursuing the Nash equilibrium strategy; Based on the concavity and monotonic interval of the utility function, analyzing the positive and negative of the calculated utility function gradients within the strategy set constraints to obtain the necessary and sufficient conditions for the consistency of the Stackelberg equilibrium strategy and Nash equilibrium strategy of the senior leader and intermediate follower.

8. The method according to claim 7, characterized in that Setting the multi-level multi-player secure master-slave game decision model according to the consistency conditions, including: Substituting the utility functions of the senior leader and intermediate follower into the utility function gradients and necessary and sufficient conditions of the senior leader and intermediate follower in pursuing the Stackelberg equilibrium strategy and Nash equilibrium strategy, and numerically deducing the adjustable parameters in the utility functions of the senior leader and intermediate follower that satisfy the necessary and sufficient conditions, and setting the adjustable parameters for the senior leader and intermediate follower so that the Stackelberg equilibrium strategy and Nash equilibrium strategy of the senior leader and intermediate follower are consistent.

9. A security protection device based on a multi-level multi-player secure master-slave game, characterized in that, The device includes: A construction module, configured to construct a multi-level and multi-player security leader-follower game decision model corresponding to a security protection system according to large-scale and complex attack and defense scenarios faced by the security protection system; An analysis module, configured to analyze the Stackelberg equilibrium and Nash equilibrium of the constructed multi-level and multi-player security leader-follower game decision model to solve the Stackelberg equilibrium strategy and Nash equilibrium strategy; A determination module, configured to determine the consistency condition between the Stackelberg equilibrium strategy and the Nash equilibrium strategy of the multi-level and multi-player security leader-follower game decision model according to the threats to the reliability of the Stackelberg equilibrium of the constructed multi-level and multi-player security leader-follower game decision model caused by disturbances and uncertainty factors faced during actual application; An application module, configured to set the multi-level and multi-player security leader-follower game decision model according to the consistency condition, and apply the set multi-level and multi-player security leader-follower game decision model to the advanced persistent threat confrontation attack and defense scenario actually faced by the security protection system for security protection.

Citation Information

Cited By

  • Robot decision planning method and device based on three-level player master-slave game

    CN120287283A

  • A robot decision planning method and device based on a three-level player master-slave game

    CN120287283B