A Network Security Defense Method Based on Game Reinforcement Learning
Through game reinforcement learning methods, key assets and action space are determined, PPO agents are created, and defense strategies are optimized, which solves the problem of difficulty in selecting network security defense strategies in the existing technology, and realizes automated and effective network security defense.
Patent Information
- Application Number
- CN202411787349.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Existing cybersecurity defense measures are difficult to effectively and automatically choose the best defense strategy when facing complex and dynamic cyber threats, resulting in high defense costs and inefficiency.
The game reinforcement learning method is adopted to establish offensive and defensive game goals by determining key assets and importance weights, attacker and defender action space and probability, create PPO agents, alternately train attacker and defender strategies, and optimize defense strategies to deal with network security challenges.
Simulate real network environments, improve defense capabilities, automatically detect and respond to attacks, optimize defense strategies, identify vulnerabilities, and improve network security defense levels.
Smart Images

Figure CN119675938B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security. More specifically, it relates to a network security defense method based on game reinforcement learning. Background Art
[0002] Network information systems have now been deeply integrated into people's daily lives, government operations, industrial development, and academic research. Therefore, the importance of network security has become increasingly prominent. With the rapid development of technology, various network security threats are also growing continuously. In particular, Advanced Persistent Threats (APTs) have brought severe challenges to various fields. At the same time, network information systems have become more and more complex and interconnected, which has exacerbated the severity and destructiveness of network security problems. Therefore, ensuring network security has become a major issue that urgently needs to be solved, and automated network security defense is an important measure to ensure network security.
[0003] Currently, automated network security defense faces many challenges. To effectively resist malicious attacks, continuous monitoring, traffic analysis, and vulnerability patching are required. This is not only costly but also requires a large amount of human input. In addition, there is an obvious asymmetry in the field of network security. Attackers can often easily launch attacks, while defenders face great difficulties, which increases the difficulty of network security defense.
[0004] Traditional network security defense measures, such as firewalls, Intrusion Detection Systems (IDS), and Intrusion Prevention Systems (IPS), have shown certain limitations in the process of the rapid development of the Internet. With the continuous complexity of the network environment and the continuous increase of security risks, developing efficient and intelligent automated defense methods has become an important research direction in the field of network security. Automated defense mechanisms can effectively monitor network activities and timely detect and prevent various types of network attacks, providing strong guarantees for the safe operation of the network environment. However, how to automatically, accurately, and effectively select the best defense strategy in this complex situation has become a major challenge that urgently needs to be solved. Summary of the Invention
[0005] The object of the present invention is to overcome the deficiencies of the prior art and provide a network security defense method based on game reinforcement learning. In the dynamic attack and defense scenario, through the mutual cooperation of reinforcement learning and game theory, the optimal network security defense strategy is sought to better cope with various challenges in the field of network security and improve the overall network security defense ability.
[0006] In order to achieve the above-mentioned object of the invention, the network security defense method based on game reinforcement learning of the present invention is characterized by comprising the following steps:
[0007] (1) Establishing offensive and defensive game objectives
[0008] Determine key assets and importance weights: S = {s1, s2, ...s J}, W={w1,w2,…s J}, where S is the key asset set, s j ,j=1,2,…,J is the jth key asset, W is the importance weight set, w j ,j=1,2,…,J is the weight of the jth key asset, and J is the number of key assets;
[0009] Determine the attacker's action space, probability of occurrence, and impact: A attacker =
[0010] {a attacker1 ,a attacker ,…a attackerM}, P={p1,p2,…p M}, I={i1,i2,…i M}, where A attacker is the attacker's action space, a attackerm ,m=1,2,…,M is the mth attacker action, P is the occurrence probability set, p m ,m=1,2,…,M is the probability of the mth attacker’s action, I is the set of influence levels, i m ,m=1,2,…,M is the impact degree of the mth attacker action on the key assets, and M is the number of attacker actions;
[0011] Determine the defender's action space, action probability, and cost: A defender =
[0012] {a defender1 ,a defender2 ,…a defenderN}, Q={q1,q2,…q N}, C={c1,c2,…d N}, where A defender is the defender's action space, a defendern ,n=1,2,…,N is the nth defender action, Q is the occurrence probability set, q n ,n=1,2,…,N is the probability of the nth defender’s action, C is the cost set, c n ,n=1,2,…,N is the cost of the nth defender action, and N is the number of defender actions;
[0013] Determine the offense-defense game objective: Among them, F is the defense objective function, T is the time step of a complete offense-defense interaction round, and w(t) is the weight of the critical asset targeted by the attacker's action at the t-th time step, that is, at the t-th time step for the j-th critical asset s j , and its corresponding importance weight is w(t), then w(t) = w j , p(t) and i(t) are respectively the occurrence probability and the impact degree on the critical asset of the attacker's action at the t-th time step, that is, at the t-th time step for the m-th attacker's action a attackerm , then p(t) = p m , i(t) = i m , q(t) and c(t) are respectively the occurrence probability and the cost of the defender's action at the t-th time step, that is, at the t-th time step for the n-th attacker's action a defendern , then q(t) = q n , c(t) = c n .
[0014] (2) Set the game environment
[0015] Under the framework of game reinforcement learning, the reward r attacker (t) of the attacker is:
[0016] r attacker (t) = RA suc (attacker state, attacker action) - RA cost (attacker state, attacker action), where RA suc (attacker state, attacker action) = p(t) × i(t), RA cost (attacker state, attacker action) is the risk perceived by the system after taking the attacker's action a attacker (t), a attacker (t) is the attacker's action at the t-th time step, that is, at the t-th time step for the m-th attacker's action a attackerm , then a attacker (t) = a attackerm ;
[0017] The reward r defende (t) of the defender is:
[0018] r defender (t) = RD suc (defender state, defender action) - RD cost (defender state, defender action), where RD suc (defender state, defender action) is the risk perceived by the system after taking the defense action a defender(t), Reward for successful defense, RD cost (Defender state, Defender action) is the cost of taking the defender action, a defender (t) is the defender action at the t-th time step, i.e., the n-th defender action a at the t-th time step defendern , then a defender (t) = a defendern , RD cost (Defender state, Defender action) = q(t) × c(t);
[0019] (3), Attack-defense game
[0020] Create two PPO agents, namely the attacker PPO agent and the defender PPO agent. The attacker PPO agent is used to obtain the attacker action a attacker (t), the attacker's reward r attacker (t), and the attacker PPO agent's network parameters are the attacker's policy π attacker , and the defender PPO agent is used to obtain the defender action a attacker (t) based on the defender's environmental state s defender (t) at the t-th time step. The network parameters of the defender PPO agent are the defender's policy π defender ; defender ;
[0021] Build a network simulation environment E. The input of the simulation environment E is the attacker action a attacker (t) and the defender action a defender (t) at the t-th time step. The output of the simulation environment E is the attacker's environmental state s attacker (t + 1) and the reward r attacker (t + 1), as well as the defender's environmental state s defender (t + 1) and the reward r attacker (t + 1);
[0022] Alternately train the attacker's policy and the defender's policy:
[0023] 3.1), Initialization
[0024] Initialize the attacker's policy pool Pool attacker and the defender's policy pool Pool defender , set the maximum number of iterations to K, and initialize the iteration number k = 0;
[0025] 3.2), Defender policy optimization
[0026] Clear the experience buffer, initialize the simulation environment E, and initialize the network parameters of the defender PPO agent, i.e., the defender policy π defender , and sample and select an attacker policy π from the attacker policy pool Pool attacker for the attacker PPO agent. Among them, there is an 80% probability of sampling and selecting an attacker policy π randomly from the first L attacker policies in the attacker policy pool Pool attAckEr , and a 20% probability of sampling and selecting an attacker policy π randomly from the remaining attacker policies in the attacker policy pool Pool attacker ; attacker attacker According to the simulation environment E, obtain the environmental state s(1) of the attacker at the first time step and the attacker's reward r(1), as well as the environmental state s(1) of the defender and the defender's reward r(1). Then, send the environmental state s(1) into the attacker PPO agent to obtain the attacker's action a(1), send the environmental state s(1) of the defender into the defender PPO agent to obtain the defender's action a(1), and input the attacker's action a(1) and the defender's action a(1) into the simulation environment E to obtain the environmental state s(2) of the attacker, the attacker's reward r(2), the environmental state s(2) of the defender, and the defender's reward r(2) at the second time step; attacker ;
[0027] Obtain the environmental state s(1) of the attacker at the first time step according to the simulation environment E attacker (1) and the attacker's reward r attacker (1), as well as the environmental state s defender (1) of the defender and the defender's reward r defender (1). Then, send the environmental state s attacker (1) into the attacker PPO agent to obtain the attacker's action a attacker (1), send the environmental state s defendee (1) of the defender into the defender PPO agent to obtain the defender's action a defender (1), and input the attacker's action a attacker (1) and the defender's action a defende (1) into the simulation environment E to obtain the environmental state s attacker (2) of the attacker at the second time step and the attacker's reward r attacker (2), as well as the environmental state s defender (2) of the defender and the defender's reward r defender (2);
[0028] Sample and select an attacker policy π from the attacker policy pool Pool attacker for the attacker PPO agent. Send the environmental state s attacker (2) into the attacker PPO agent to obtain the attacker's action a attacker (2), send the environmental state s attacker (2) of the defender into the defender PPO agent to obtain the defender's action a defender (2), and input the attacker's action a defender (2) and the defender's action a attacker (2) into the simulation environment E to obtain the environmental state s defender (3) of the attacker at the third time step and the attacker's reward r attacker (3) and the defender's reward r attacker(3) and the environmental state s of the defender defender (3) and the reward r of the defender defender (3);
[0029] Perform adversarial training in this way to obtain a series of adversarial samples:
[0030] (s defender (t), a defender (t), r defender (t), s defender (t + 1))
[0031] Put the adversarial samples into the experience buffer. When it is full, take out all the adversarial samples from the experience buffer and use the PPO algorithm to update the network parameters of the defender's PPO agent, that is, the defender's policy π defender ;
[0032] Judge the updated defender's policy π defebder compared with the current defender's policy π defender to see if it has improved. If so, clear the experience buffer, continue to put in adversarial samples. When it is full, continue to take them out and update the defender's policy π defender , continue to judge if it has improved. If so, continue to clear the experience buffer, continue to put in adversarial samples until there is no improvement, and put the new defender's policy π defender at the front of the defender's policy pool Pool defender ;
[0033] 3.3) Optimization of the attacker's policy
[0034] Clear the experience buffer, initialize the simulation environment E, and initialize the network parameters of the attacker's PPO agent, that is, the attacker's policy π attacker , and sample and select a defender's policy π defender from the defender's policy pool Pool attAckEr to give to the defender's PPO agent. Among them, there is an 80% probability of sampling and selecting a defender's policy π defender randomly from the first L in the defender's policy pool Pool defender , and a 20% probability of randomly selecting a defender's policy π defender from the remaining defender's policies in the defender's policy pool Pool defender ;
[0035] According to the simulation environment E, obtain the environmental state s attacker (1) of the attacker at the first time step and the reward r attacker (1) of the attacker and the environmental state s defender (1) of the defender and the reward r defender (1) of the defender, and put the environmental state sattacker (1) Feed the attacker's PPO agent to obtain the attacker's action a attacker (1), and the environmental state s of the defender defendee (1) Feed the defender's PPO agent to obtain the defender's action a defender (1), and the attacker's action a attacker (1), the defender's action a defender (1) Input them into the simulation environment E to obtain the attacker's environmental state s at the second time step attacker (2) and the attacker's reward r attacker (2) as well as the defender's environmental state s defender (2) and the defender's reward r defender (2);
[0036] Sample and select a defender's policy π from the defender's policy pool Pool defender and give it to the defender's PPO agent. Feed the environmental state s defender (2) to the attacker's PPO agent to obtain the attacker's action a attacker (2), and feed the defender's environmental state s attacker (2) to the defender's PPO agent to obtain the defender's action a defender (2), and the attacker's action a defender (2), the defender's action a attacker (2) Input them into the simulation environment E to obtain the attacker's environmental state s at the third time step defender (3) and the attacker's reward r attacker (3) as well as the defender's environmental state s attacker (3) and the defender's reward r defender (3); defender (3);
[0037] Conduct adversarial training in this way to obtain a series of adversarial samples:
[0038] (s attacker (t), a attacker (t), r attacker (t), s attacker (t + 1))
[0039] Put the adversarial samples into the experience buffer. When it is full, take out all the adversarial samples from the experience buffer and use the PPO algorithm to update the network parameters of the attacker's PPO agent, that is, the attacker's policy π attacker ;
[0040] Judge the updated attacker's policy π attacker and the current attacker's policy π attackerCompare whether there is an improvement. If so, clear the experience buffer, continue to put in adversarial samples, and when it is full, continue to take out and update the attacker's strategy π attacker , continue to judge whether there is an improvement. If so, continue to clear the experience buffer, continue to put in adversarial samples, and continue to judge until there is no improvement. Then, put the current attacker's strategy π attacker into the attacker's strategy pool Pool attacker at the front;
[0041] 3.3), k = k + 1, judge whether the iteration number k is less than the maximum iteration number K. If it is not less, the iteration ends, and the obtained defender's strategy π defender is the optimal network security defense strategy. Otherwise, return to step 3.2).
[0042] The invention object of the present invention is realized as follows:
[0043] The network security defense method based on game reinforcement learning of the present invention first determines key assets and importance weights, determines the attacker's action space, action occurrence probability and influence degree, determines the defender's action space, action occurrence probability and cost, and then establishes the attack - defense game goal. Then, a game environment is set. On this basis, two PPO agents and a network construction simulation environment E are created. In the scenario of dynamic changes in attack and defense, through the mutual cooperation of reinforcement learning and game theory, the attacker's strategy and the defender's strategy are alternately trained to obtain the optimal network security defense strategy to better cope with various challenges in the field of network security and improve the overall network security defense ability.
[0044] The present invention has the following beneficial effects:
[0045] 1. Simulate a real network attack - defense environment. By establishing a capture - the - flag environment and virtual games, the interaction between the attacker and the defender can be made closer to the interaction in the real scenario. At the same time, the opponent in the game is also evolving continuously, increasing the exploratory nature of the attacker's and defender's strategies;
[0046] 2. Improve the network security defense ability. The trained defender will have the ability to automatically detect and respond to network attacks, and it continuously learns and optimizes the defender's strategy, thereby enhancing the overall network protection level;
[0047] 3. Explore new network security strategies. Through the game interaction between the attacker and the defender, we can discover the shortcomings of the current defense mechanism, and then develop better defender strategies to counter new types of attacks;
[0048] 4. Identify network vulnerabilities. By simulating various attack behaviors, the attacker can help identify the security vulnerabilities of the actual system, providing a reference for subsequent repair and defense. Description of the Drawings
[0049] Figure 1 This is a flowchart of a specific implementation manner of the network security defense method based on game reinforcement learning of the present invention;
[0050] Figure 2 This is a schematic diagram of the simulation environment;
[0051] Figure 3 This is a schematic diagram of the relationship between the agent and the environment;
[0052] Figure 4 This is a schematic diagram of the attack - defense iterative training;
[0053] Figure 5 This is a schematic diagram of the defender's strategy training;
[0054] Figure 6 This is a schematic diagram of the attacker's strategy training. Specific implementation manner
[0055] The following describes the specific implementation manner of the present invention in conjunction with the accompanying drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0056] Reinforcement learning (RL) in artificial intelligence technology can help develop autonomous, automatic, and accurate network security decision - making algorithms, which is crucial for solving complex attack - defense interaction problems in the field of network security. In addition, game theory, as a mathematical tool, can also be used for strategy selection in a highly competitive network security environment, thereby improving the accuracy and efficiency of network security decision - making. The present invention focuses on the urgent needs in the field of network security and related automated defense mechanisms, and through the application of these technologies, it better addresses various challenges in the field of network security, finds the optimal defender strategy in a dynamically changing network through game theory and reinforcement learning, and improves the overall network security defense ability.
[0057] Figure 1 This is a flowchart of a specific implementation manner of the network security defense method based on game reinforcement learning of the present invention.
[0058] In this embodiment, as Figure 1 shown, the network security defense method based on game reinforcement learning of the present invention includes the following steps:
[0059] Step S1: Establish the attack - defense game goal
[0060] Step S1.1: Determine critical assets and importance weights: Sort out the organization's key business processes, information systems, and data to identify critical assets, and assign corresponding importance weights to each critical asset. The weights can be determined based on characteristics such as the confidentiality, integrity, and availability of the assets. If a critical asset involves private data, the importance weight of the corresponding asset is higher. Specifically, it can be expressed as:
[0061] S = {s1, s2, … s J}, W = {w1, w2, … s J}, where S is the set of critical assets, s j , j = 1, 2, …, J is the jth critical asset, W is the set of importance weights, w j , j = 1, 2, …, J is the weight of the jth critical asset, and J is the number of critical assets.
[0062] In this embodiment, as Figure 2 shown, there are 3 subnets and 13 machines. In subnet 1, there are 5 user hosts; in subnet 2, there are 3 enterprise servers and 1 defense host; in subnet 3, there are 3 operation hosts and 1 operation server.
[0063] Each machine exposes some network services that other machines can connect to, and there may also be exploitable vulnerabilities. However, there is a firewall in this network, and the machines in subnet 1 cannot directly connect to the machines in subnet 3. Only the operation hosts have permission to access the operation server, and the security of the operation server is crucial to the entire manufacturing process.
[0064] The critical assets S = {subnet 1 user machines, subnet 2 enterprise servers, subnet 3 operation servers}, that is, J = 3, s1 = subnet 1 user machines, s2 = subnet 2 enterprise servers, s3 = subnet 3 operation servers. The importance weights W = {0.05, 0.15, 0.8}, that is, w1 = 0.05, w2 = 0.15, w3 = 0.8. The weight of the subnet 1 user machines s1 is relatively low because it is not directly involved in critical business; the weight of the subnet 2 enterprise servers s2 is relatively high because they store important data; the weight of the subnet 3 operation servers s3 is the highest because it is crucial to the entire manufacturing process.
[0065] Step S1.2: Determine the attacker's action space, occurrence probability, and impact degree: Based on historical events, intelligence information, security vulnerability reports, etc., determine the possible network attack actions that the organization may face, analyze each attack action, and determine its occurrence probability and impact degree on critical assets. Specifically, it can be expressed as:
[0066] A attacker = {a attacker1 , a attacker2 , … aattackerM}, P={p1,p2,…p M}, I={i1,i2,…i M}, where A attacker is the attacker's action space, a attackerm ,m=1,2,…,M is the mth attacker action, P is the occurrence probability set, p m ,m=1,2,…,M is the probability of the mth attacker’s action, I is the set of influence levels, i m ,m=1,2,…,M is the impact degree of the mth attacker action on the critical assets, and M is the number of attacker actions.
[0067] In this embodiment, the attacker's actions include port scanning, vulnerability exploitation, privilege escalation, and data theft, that is, the attacker's action data M=4, a attacker1 =Port scan, a attacker2 =Vulnerability exploitation,
[0068] a attacker3 =Privilege escalation, a attacker = data theft, with probabilities of 0.5, 0.25, 0.15, and 0.1, respectively: p1 = 0.5, p2 = 0.25, p3 = 0.15, and p4 = 0.1. Port scanning has a higher probability because it requires the discovery of system vulnerabilities; vulnerability exploitation has the next highest probability, as the system may have unpatched vulnerabilities; privilege escalation has the next highest probability, as elevated privileges allow access to critical assets; and data theft has the lowest probability. The impact levels are set to i1 = 0.1, i2 = 0.15, i3 = 0.25, and i4 = 0.5. Port scanning has the lowest impact, exposing only system information; vulnerability exploitation has a higher impact, potentially leading to system compromise; privilege escalation has an even higher impact, allowing access to critical assets; and data theft has the highest impact, potentially causing significant losses.
[0069] Step S1.3: Determine the defender's action space, probability of occurrence, and cost: A defender =
[0070] {a defender1 ,a defender2 ,…a defenderN}, Q={q1,q2,…q N}, C={c1,c2,…d N}, where D is the defender's action space, a defendern ,n=1,2,…,N is the nth defender action, Q is the occurrence probability set, q n ,n=1,2,…,N is the probability of the nth defender’s action, C is the cost set, c n, where \(n = 1, 2, \ldots, N\) is the cost of the \(n\)th defender action, and \(N\) is the number of defender actions.
[0071] In this embodiment, there are four defender actions, i.e., \(N = 4\), including: defender a = Analyze the machine, by collecting machine information, identify malicious behaviors and permission status, and master the current system security situation;
[0072] a defendrt2 a = Terminate malicious processes, directly prevent the attacker's malicious activities, but will affect the system availability in the short term; defender3 a = Restore the system; defender4 a = Deploy a honeypot service. A honeypot is a defensive measure with a trapping nature that can attract and observe attackers, but it needs to be deployed carefully to avoid affecting normal business.
[0073] The costs of the defender actions are \(c1 = 0\), \(c2 = 0.1\), \(c3 = 0.7\), \(c4 = 0.2\). The occurrence probabilities of the defender actions are \(q1 = 0.4\), \(q2 = 0.2\), \(q3 = 0.3\), \(q4 = 0.1\).
[0074] Step S1.4: Determine the attack - defense game objective: According to the critical assets and the attacker's actions, determine the organization's attack - defense game objective. The attack - defense game objective can be defined as: minimizing the potential loss and recovery cost of critical assets, which can be expressed by the following formula:
[0075] where \(F\) is the defense objective function, \(T\) is the time step of a complete round of attack - defense interaction, \(w(t)\) is the weight of the critical asset targeted by the attacker's action at the \(t\)th time step, i.e., at the \(t\)th time step for the \(j\)th critical asset \(s\) j , then \(w(t)=s\) j , \(p(t)\) and \(i(t)\) are respectively the occurrence probability and the impact degree on the critical asset of the attacker's action at the \(t\)th time step, i.e., at the \(t\)th time step for the \(m\)th attacker action \(a\) attackerm , then \(p(t)=p\) m , \(i(t)=i\) m , \(q(t)\) and \(c(t)\) are respectively the occurrence probability and the cost of the defender's action at the \(t\)th time step, i.e., at the \(t\)th time step for the \(n\)th attacker action \(a\) defendern , then \(q(t)=q\) n , \(c(t)=c\) n .
[0076] Step S2: Set the game environment
[0077] In the framework of game - based reinforcement learning, the attacker's reward \(r\) attacker (t) is:
[0078] rattacker R_A(t) = R_A suc (Attacker's state, Attacker's action) - R_A cost (Attacker's state, Attacker's action), where R_A suc (Attacker's state, Attacker's action) = p(t) × i(t), R_A cost (Attacker's state, Attacker's action) is the risk that the attacker is detected by the system after taking the attacker's action a attacker (t), and a attacker (t) is the attacker's action at the t-th time step, i.e., the t-th time step is the m-th attacker's action a attackerm , then a attacker (t) = a attackerm .
[0079] The defender's reward r defender (t) is as follows:
[0080] r defender (t) = R_D suc (Defender's state, Defender's action) - R_D cost (Defender's state, Defender's action), where R_D suc (Defender's state, Defender's action) is the reward for successful defense after taking the defense action a defender (t). In a specific real-time process, it can be defined as a large constant representing the reward for successful defense. R_D cost (Defender's state, Defender's action) represents the cost of taking the defender's action, a defender (t) is the defender's action at the t-th time step, i.e., the t-th time step is the n-th defender's action a defendern , then a defender (t) = a defendern , R_D cost (Defender's state, Defender's action) = q(t) × c(t).
[0081] In this embodiment, the state space and action space of the attacker and the defender in the network security attack and defense game can be described as follows:
[0082] The attacker's state space is divided into four parts. Network topology information, including the number of machines and the number of connections mastered by the attacker; system vulnerability information, including the types and severity levels of vulnerabilities existing on each machine detected by the attacker; system asset information, including the important assets existing on each machine confirmed by the attacker; system access permission information, including the access permissions of each user obtained by the attacker. In this embodiment, the attacker's state space is a 52-bit vector. It includes the number of machines N attacker-host = 13, the number of connections Nattacker-connections = 122 etc., can be represented by the vector [13, 122]. The remaining 50-bit vector represents the basic status information of the machines that can be obtained, including whether there is abnormal activity on the machines and the degree of harm to the machines.
[0083] The action space of the attacker: port scanning, vulnerability exploitation, privilege escalation, data theft.
[0084] The state space of the defender is divided into five parts. Network topology information, including the number of machines and the number of connections mastered by the defender, etc.; system vulnerability information, including the types and severity levels of vulnerabilities on each machine discovered by the defender; system security configuration information, including the security protection measures deployed on each machine by the defender; asset information, including the important assets existing on each machine confirmed by the defender; system access permission information, including the access permissions of each user mastered by the defender. In this embodiment, the state space of the defender is a 40-bit vector. The network topology information includes the number of machines N defender-host = 13 and the number of connections N defender-connections = 122. The remaining 38-bit vector represents whether the previous attack was successful, which machines were attacked, and the privilege level on the machines.
[0085] The action space of the defender: analyze the machine, terminate malicious processes, terminate malicious processes, and deploy honeypot services.
[0086] Step S3: Attack and defense game
[0087] The optimal defender, optimal attacker, and game equilibrium in network security attack and defense game: We use π attacker to represent the attacker's strategy and π defender to represent the defender's strategy. Here, the strategy refers to the rule or method of choosing an action in a given state. It can be a deterministic strategy, that is, always choosing the same action in a specific state; or a random strategy, that is, choosing different actions with a certain probability in a specific state. Use J attAcKEr to represent the attacker's payoff under a specific attacker's strategy and defender's strategy, that is, the time-accumulated RA (attacker state, attacker action). Use J deFender to represent the defender's payoff under a specific attacker's strategy and defender's strategy, that is, the time-accumulated RD (defender state, defender action).
[0088] Optimal Defender. The ultimate goal of the attack and defense game is to find the optimal strategy of the defender such that the defender's payoff J defender reaches the maximum value.
[0089] in is the attacker's optimal strategy. The optimal defender will choose the optimal defender strategy based on the defender state space and the defender action space. To maximize your own benefits defender .
[0090] Optimal Attacker. When we find the optimal strategy for the defender, we will also find the optimal strategy for the attacker. This is a method to help us find the optimal strategy for the defender, so that the attacker's profit J attac Reaching the maximum value,
[0091] in is the optimal strategy of the defender. The optimal attacker will choose the optimal attacker strategy based on the attacker state space and attacker action space. To maximize your own benefits attacker .
[0092] When the game is in equilibrium, we find the optimal attacker strategy and defender strategies The following conditions are met.
[0093]
[0094] In the state of attack-defense game equilibrium, neither the attacker nor the defender can increase their own benefits by unilaterally changing their strategy. Specifically:
[0095] Create two PPO agents, the attacker PPO agent and the defender PPO agent. The attacker PPO agent is used to determine the attacker's environment state s according to the t-th time step. attacker (t) Get the attacker's action a attacker (t), the network parameters of the attacker’s PPO agent are the attacker’s strategy π attacker , the defender PPO agent is used to determine the defender's environment state s at the t-th time step defende (t), the defender's reward r defender (t), get the defender's action a defender (t), the network parameters of the defender PPO agent are the defender strategy π defender .
[0096] Build a simulation environment E for the network. The input of the simulation environment E is the attacker action a at the tth time step. attacker (t) and the defender action a defender(t), the output of the simulation environment E is the attacker's environmental state s attacker (t + 1) and the reward r attacker (t + 1) and the defender's environmental state s defender (t + 1) and the reward r attacker (t + 1).
[0097] In this embodiment, as Figure 3 shown, it shows the relationship between two agents created using the PPO algorithm and the environment. These two agents need to learn appropriate strategies through reinforcement learning to cope with the challenges of attack and defense. These agents can continuously optimize their strategies to improve the execution effect of attack and defense tasks and promote the competition and learning process between agents: alternately training the attacker's strategy and the defender's strategy, as Figure 4 shown, including
[0098] Step S3.1: Initialization
[0099] Initialize the attacker's policy pool Pool attacker and the defender's policy pool Pool defender , set the maximum number of iterations to K, and initialize the iteration number k = 0. In this embodiment
[0100] Step S3.2: Defender policy optimization
[0101] In this embodiment, as Figure 5 shown, the defender policy optimization includes the following steps:
[0102] Step S3.2.1: Initialization
[0103] Empty the experience buffer, initialize the simulation environment E, and initialize the network parameters of the defender's PPO agent, i.e., the defender's policy π defender ;
[0104] Step S3.2.2: Adversarial training
[0105] Sample and select an attacker's policy π attacker from the attacker's policy pool Pool attacker for the attacker's PPO agent. Among them, there is an 80% probability of sampling and selecting an attacker's policy π attacker randomly from the first L of the attacker's policy pool Pool attacker , and a 20% probability of randomly selecting an attacker's policy π attacker from the remaining attacker's policies in the attacker's policy pool Pool attacker ;
[0106] According to the simulation environment E, obtain the attacker's environmental state s at the first time step attacker(1) and the attacker's reward r attacker (1) and the defender's environmental state s defender (1) and the defender's reward r defender (1), and send the environmental state s attacker (1) to the attacker's PPO agent to obtain the attacker's action a attacker (1), send the defender's environmental state s defender (1) to the defender's PPO agent to obtain the defender's action a defender (1), send the attacker's action a attacker (1), the defender's action a defender (1) into the simulation environment E to obtain the attacker's environmental state s at the second time step attacker (2) and the attacker's reward r attacker (2) and the defender's environmental state s defender (2) and the defender's reward r defender (2);
[0107] Sample and select an attacker's policy π from the attacker's policy pool Pool attacker for the attacker's PPO agent, send the environmental state s attacker (2) to the attacker's PPO agent to obtain the attacker's action a attacker (2), send the defender's environmental state s attacker (2) to the defender's PPO agent to obtain the defender's action a defender (2), send the attacker's action a defender (2), the defender's action a attacker (2) into the simulation environment E to obtain the attacker's environmental state s at the third time step defender (3) and the attacker's reward r attacker (3) and the defender's environmental state s attacker (3) and the defender's reward r defender (3); defender (3);
[0108] Perform adversarial training in this way to obtain a series of adversarial samples:
[0109] (s defender (t), a defender (t), r defender (t), s defender (t + 1))
[0110] Step S3.2.3: Update the network parameters of the defender's PPO agent
[0111] The adversarial examples are put into the experience buffer. When the buffer is full, all the adversarial examples are taken out from the experience buffer, and the PPO algorithm is used to update the network parameters of the defender PPO agent, i.e., the defender policy π. defender ;
[0112] Step S3.2.4: Determine whether the updated defender policy π defender is improved compared with the current defender policy π defender If so, clear the experience buffer, continue to put adversarial examples. When the buffer is full, take out and update the defender policy π defender , continue to determine whether it is improved. If so, continue to clear the experience buffer and continue to put adversarial examples until there is no improvement. Then put the new defender policy π defender at the front of the defender policy pool Pool defender .
[0113] Step S3.3: Optimize the attacker's policy
[0114] In this embodiment, as Figure 5 shown, the optimization of the attacker's policy includes the following steps:
[0115] Step S3.3.1: Initialization
[0116] Clear the experience buffer, initialize the simulation environment E, and initialize the network parameters of the attacker PPO agent, i.e., the attacker policy π attacker ;
[0117] Step S3.3.2: Adversarial training
[0118] Sample and select a defender policy π defender from the defender policy pool Pool defender for the defender PPO agent. Among them, there is an 80% probability of randomly selecting a defender policy π defender from the first L in the defender policy pool Pool ddefender , and a 20% probability of randomly selecting a defender policy π defender from the remaining defender policies in the defender policy pool Pool defender ;
[0119] According to the simulation environment E, obtain the environmental state s attacker (1) of the attacker at the first time step and the reward r attacker (1) of the attacker, as well as the environmental state s defender (1) of the defender and the reward r defender (1) of the defender, and send the environmental state s attacker (1) to the attacker PPO agent to obtain the attacker's action a attacker(1) Input the environmental state s of the defender defender (1) into the defender's PPO agent to obtain the defender's action a defender (1) Input the attacker's action a attacker (1) and the defender's action a defender (1) into the simulation environment E to obtain the attacker's environmental state s at the second time step attacker (2) and the attacker's reward r attacker (2) as well as the defender's environmental state s defender (2) and the defender's reward r defender (2);
[0120] Sample and select a defender policy π from the defender's policy pool Pool defender for the defender's PPO agent. Input the environmental state s defender (2) into the attacker's PPO agent to obtain the attacker's action a attacker (2). Input the defender's environmental state s attacker (2) into the defender's PPO agent to obtain the defender's action a defender (2). Input the attacker's action a defende (2), the defender's action a attacker (2) into the simulation environment E to obtain the attacker's environmental state s at the third time step defender (3) and the attacker's reward r attacker (3) as well as the defender's environmental state s attacker (3) and the defender's reward r defender (3); defender (3);
[0121] Conduct adversarial training in this way to obtain a series of adversarial samples:
[0122] (s attecker (t), a attacker (t), r attacker (t), s attacker (t + 1))
[0123] Step S3.3.3: Update the network parameters of the attacker's PPO agent
[0124] Put the adversarial samples into the experience buffer. When it is full, take out all the adversarial samples from the experience buffer and use the PPO algorithm to update the network parameters of the attacker's PPO agent, that is, the attacker's policy π attacker ;
[0125] Step S3.3.4: Judge whether the updated attacker's policy π attacker is the same as the current attacker's policy π attackerCompare whether there is an improvement. If so, clear the experience buffer, continue to put adversarial examples, and when it is full, continue to take them out and update the attacker's policy π attacker , continue to judge whether there is an improvement. If so, continue to clear the experience buffer, continue to put adversarial examples, and continue to judge until there is no improvement. Then put the current attacker's policy π attacker into the attacker's policy pool Pool attacker at the front.
[0126] Step S3.4: k = k + 1, judge whether the iteration number k is less than the maximum iteration number K. If it is not less than, the iteration ends, and the obtained defender's policy π defender is the optimal network security defense policy. Otherwise, return to Step S3.2.
[0127] Although the above illustrative specific embodiments of the present invention have been described to facilitate those skilled in the art of the present technology to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
Claims
1. A network security defense method based on game reinforcement learning, characterized in that It includes the following steps: (1), Establish the attack and defense game objectives Determine key assets and importance weights: S = {s1, s2, … s J}, W = {w1, w2, … w J}, J is the number of key assets; Determine the attacker's action space A attacker = {a attacker1 , a attacker2 , … a attackerM}, the action occurrence probability P = {p1, p2, … p M}, the influence degree I = {i1, i2, … i M}, where M is the number of attacker actions; Determine the defender's action space: A defender = {a defender1 , a defender2 , … a defenderN}, the action occurrence probability Q = {q1, q2, … q N} and the cost C = {c1, c2, … c N}, where N is the number of defender actions; Determine the defense objective function F: where T is the time step of one round of attack and defense interaction, w(t) is the weight of the critical asset targeted by the attacker's action at the t-th time step, w(t)=w j , j = 1, 2, …, J, p(t) and i(t) are the occurrence probability and the impact degree on the critical asset corresponding to the attacker's action at the t-th time step respectively, p(t)=p m , i(t) = i m , m = 1, 2, …, M, q(t) and c(t) are the occurrence probabilities and costs of the defender's actions at the t-th time step respectively, q(t) = q n , c(t) = c n , n = 1, 2, …, N; (2), Set the game environment The attacker's reward r attacker (t) = RA suc (Attacker state, attacker action) - RA cost (Attacker state, attacker action), where RA suc (Attacker state, attacker action) = p(t) × i(t), RA cost For taking a attacker (t), the risk detected by the system after a attacker (t) = a attackerm ; Reward r of the defender defender (t) = RD suc (Defender state, Defender action) - RD cost (Defender state, Defender action), where RD suc is the reward for successful defense after taking a defender (t), RD cost is the cost of taking the defender action, a defender (t) = a defendern , RD cost (Defender state, Defender action) = q(t) × c(t); (3), Attack and defense game Alternately train the attacker's strategy and the defender's strategy: 3.1), Initialization Initialize the attacker's policy pool Pool attacker and the defender's policy pool Pool defender , set the maximum number of iterations to K, and initialize k = 0; 3.2), Optimize the defender's strategy Clear the experience buffer, initialize the simulation environment E, and the defender PPO agent policy π defender , select a policy π from Pool attacker and give it to the attacker PPO agent; attacker According to E, obtain the attacker's environmental state s at the first time step attacker (1) and the attacker's reward r attacker (1) and the defender's environmental state s defende (1) and the defender's reward r defender (1), and send s attacker (1) into the attacker's PPO agent to obtain a attacker (1), send s defender (1) into the defender's PPO agent to obtain a defender (1), send a attacker (1), a defender (1) into E to obtain s at the second time step attacker (2), r attacker (2), s defender (2), r defender (2); Select a π from Pool attacker and give it to the attacker's PPO agent, input s attacker (2) into the attacker's PPO agent to get a attacker (2), input s attacker (2) into the defender's PPO agent to get a defender (2), input a defender (2) and a attacker (2) into E to get s defende (3), r attacker (3), s attacker (3), r defender (3) at the third time step; defender (3); Perform adversarial training in this way to obtain a series of adversarial samples: (s defender (t),a defender (t),r defender (t),s defender (t + 1)) The adversarial examples are put into the experience buffer. When the buffer is full, all the adversarial examples are taken out from it, and the PPO algorithm is used to update π defender ; Judge the updated π defender and the current π defender to see if it has improved. If so, clear the experience buffer, continue to put in adversarial samples, and when it is full, continue to take out and update π defender , judge if it has improved. If so, continue to clear it and continue to put in adversarial samples until there is no improvement, and then put the new π defender at the front of Pool defende ; Optimize the attacker's strategy according to the above defender's strategy optimization method; 3.3), k = k + 1, determine whether k is less than K. If not, end and the obtained π defender is the optimal network security defense strategy. Otherwise, return to step 3.2).
2. The network security defense method of game reinforcement learning according to claim 1, characterized in that The specific steps for optimizing the attacker's strategy are as follows: Clear the experience buffer, initialize the simulation environment E, and initialize the network parameters of the attacker PPO agent, i.e., the attacker's policy π attacker , select a defender policy π from the defender policy pool Pool defender ; attacker Give it to the defender PPO agent. Among them, when sampling and selecting, there is an 80% probability of randomly selecting a defender policy π from the first L defender policies in the defender policy pool Pool defender ; defender , and a 20% probability of randomly selecting a defender policy π from the remaining defender policies in the defender policy pool Pool defender ; defender ; According to the simulation environment E, obtain the environmental state s of the attacker at the first time step attacker (1) and the reward r of the attacker attacker (1) and the environmental state s of the defender defender (1) and the reward r of the defender defender (1), and send the environmental state s attacker (1) to the attacker's PPO agent to obtain the attacker's action a attacker (1), send the environmental state s of the defender defender (1) to the defender's PPO agent to obtain the defender's action a defender (1), send the attacker's action a attacker (1), the defender's action a defender (1) into the simulation environment E to obtain the environmental state s of the attacker at the second time step attacker (2) and the reward r of the attacker attacker (2) and the environmental state s of the defender defender (2) and the reward r of the defender defender (2); Sample a defender strategy π from the defender strategy pool Pool defender Feed the defender PPO agent with the environmental state s defender and send it to the attacker PPO agent to obtain the attacker's action a attacker (2), Feed the defender's environmental state s attacker (2) into the defender PPO agent to obtain the defender's action a defender (2), Feed the attacker's action a defender (2), the defender's action attacker (2), the attacker's action a and the defender's action a defender (2) Input into the simulation environment E to obtain the environmental state of the attacker at the 3rd time step s attacker (3) and the attacker's reward r attacker (3) and the defender's environmental state s defender (3) and Reward r of the defender defender (3); Perform adversarial training in this way to obtain a series of adversarial samples: (s attacker (t), a attacker (t), r attacker (t), s attacker (t + 1)) The adversarial samples are put into the experience buffer. When the buffer is full, all the adversarial samples are taken out from the experience buffer, and the PPO algorithm is used to update the network parameters of the attacker's PPO agent, that is, the attacker's policy π attacker ; Judge the updated attacker strategy π attacker and the current attacker strategy π attacker to see if it has improved. If so, clear the experience buffer, continue to put in adversarial samples, and when it is full, continue to take out and update the attacker strategy π attacker , and continue to judge if it has improved. If so, continue to clear the experience buffer, continue to put in adversarial samples, and continue to judge until there is no improvement. Then put the current attacker strategy π attacker at the front of the attacker strategy pool Pool attacker .
3. The network security defense method for game reinforcement learning according to claim 1, characterized in that The attacker's action space includes four attacker actions: port scanning, vulnerability exploitation, privilege escalation, and data theft. The attacker's state space is divided into four parts: network topology information, system vulnerability information, system asset information, and system access privilege information; The defender's action space includes four defender actions: analyzing the machine, terminating malicious processes, and deploying honeypot services. The defender's state space is divided into five parts: network topology information, system vulnerability information, system security configuration information, asset information, and system access privilege information.
Citation Information
Patent Citations
Method for selecting optimal defense strategy for moving target defense based on game theory
CN109617863A
Network spoofing defense decision-making method and system based on Flipit intelligent game
CN116962050A