A method, device, and electronic device for defending against penetration attacks based on reinforcement learning
By modeling penetration testing as a Markov decision-making process and using deep Q network algorithm to train the agent, combining modifying sensitive host rewards and deploying honeypot hosts, the dynamic interaction and poisoning attack problems of penetration testing are solved, and defense against malicious penetration attacks is achieved and network security is improved.
Patent Information
- Application Number
- CN202210949674.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-08-09
AI Technical Summary
The existing penetration testing methods based on reinforcement learning cannot interact dynamically when facing network environment updates, and there is a problem of poisoning attacks causing policy failure, and malicious penetration attacks may cause damage to the target network system.
Model the penetration testing process as a Markov decision-making process, use deep Q network algorithm to train agents to find the optimal penetration path, and defend against malicious penetration attacks by modifying the reward value of sensitive hosts and deploying honeypot hosts.
It realizes dynamic interaction when network environment is updated and effective defense against malicious infiltration attacks, reduces the impact on the target network system and improves network security.
Smart Images

Figure CN115473677B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network space security and deep reinforcement learning defense, and particularly relates to a method and device for defending against penetration attacks based on reinforcement learning, and an electronic device. Background Art
[0002] Penetration testing refers to a common method for testing the security of a network environment by simulating hacker attacks. Generally, traditional penetration testing mainly relies on manual operations. By manually using penetration tools such as Namp (vulnerability scanning tool), relevant vulnerability information can be obtained, such as the services running on the vulnerabilities and the types of services. After the scanning stage is completed, a semi-automatic penetration tool such as Metasploit is used to perform specific penetration operations on the target host. However, with the increase in the types of host vulnerabilities and the expansion of the network scale, manually using penetration testing tools to penetrate the target network will cost a large amount of manpower and time. The birth of automated penetration testing tools, such as the method of constructing an attack graph, can solve this problem to a certain extent. An attack graph refers to a connection graph that shows the vulnerability exploitation relationships of each node in the target network. From the graph, the attack sequence of the target network and the impact generated by each corresponding penetration action can be clearly shown. However, the subsequent drawback is that in the process of each network topology update, the attack graph needs to be updated again and cannot interact dynamically with the network environment. Therefore, most of the current research hotspots focus on intelligent penetration testing, that is, based on the reinforcement learning algorithm framework, training an agent to complete the automated penetration testing process in a way of continuous trial and error to achieve intelligent penetration testing. And in the face of the update of the network environment, the agent can also generate corresponding strategies.
[0003] Therefore, intelligent penetration testing based on the reinforcement learning algorithm framework can effectively reduce costs for network space security testing and defense. Intelligent penetration testing combines the reinforcement learning algorithm and the penetration testing method, and models the penetration process as a Markov decision process. In order to further enhance the effect of intelligent penetration testing, it is necessary to optimize the decision-making of the penetration path, so as to timely discover the vulnerabilities existing in the network environment and the penetration paths that the attacker may take, in order to further strengthen network security protection. However, in the process of Markov decision modeling, the strategy optimization of the reinforcement learning algorithm depends on the training data. Once the training data is obtained and maliciously polluted by a malicious attacker, the strategy will fail, thus unable to effectively perform penetration testing on the security of the network and achieve the security protection of the system.
[0004] During the training process of the reinforcement learning model, the way that training data is obtained by malicious attackers and maliciously polluted, resulting in the failure of the agent's policy, is called a poisoning attack. Based on the different attack scenarios, the poisoning attack method based on reinforcement learning can be divided into the poisoning attack on a single agent and the poisoning attack on a multi-agent system. Both methods exploit the security vulnerabilities of the agent model by poisoning the model during the training phase. Although the two targeted poisoning methods are slightly different, they both achieve the purpose of model poisoning attack by poisoning the training data. For example, a high reward value will be given during the execution of the poisoning strategy to strengthen the strategy, thus leaving a backdoor security risk.
[0005] Penetration testing is generally regarded as a positive attack based on the "blue side". However, in some scenarios, if the purpose of penetration testing is not to detect and evaluate the security vulnerabilities of the target network system, this method will also have a certain impact on different hosts of the target network. For example, a normal host or a key host is regarded as an attacking machine, and after obtaining permission through penetration attack, malicious attacks are launched on other hosts through lateral movement, which may cause fatal destruction to the target network system in severe cases. But currently, there is no method to defend against intelligent penetration attacks based on reinforcement learning. Summary of the Invention
[0006] In view of the deficiencies of the prior art, the purpose of the embodiments of the present application is to provide a method, device, and electronic device for defending against penetration attacks based on reinforcement learning. By modeling the penetration process of the target network as a Markov decision process and using the deep Q-network algorithm to train the intelligent agent to find the optimal penetration path, the intelligent penetration testing process is realized. In addition, during the training process, by modifying the reward value of the target host for penetration attack, the intelligent agent is regarded as the penetration attack party. By modifying the training data, the intelligent agent cannot perform attack operations on the target sensitive host, thus achieving the purpose of defending the target sensitive host.
[0007] According to the first aspect of the embodiments of the present application, a method for defending against penetration attacks based on reinforcement learning is provided, including:
[0008] (1) Model the penetration testing process as a Markov decision process, where the Markov decision process includes states, actions, and reward values;
[0009] (2) Train an intelligent agent based on the deep Q-network algorithm, where the intelligent agent acts as the penetration attack party, and the training objective is to generate the current optimal penetration attack path process, where the current optimal penetration attack path process is that in the current network environment, the intelligent agent uses as few actions as possible to attack the target sensitive host with the maximum value in the network environment;
[0010] (3) Flip the value of the sensitive host in the network environment, and set that the penetration attack round does not end when obtaining the Root permission of the target sensitive host. Modify the end condition of the penetration attack round to that the number of training steps in the round reaches a predetermined threshold;
[0011] (4) Update the modification of the host value in step (3) to the reward value in step (1). Use the trained agent in step (2) to perform defense training on the network environment in step (3). Repeat the process of defense training until the number of training rounds reaches a predetermined threshold to obtain a strategy for defending against penetration attacks.
[0012] Further, the state is the information observed by the agent in the network environment, including the subnets in the network environment, the hosts in each subnet, the operation type of each host, the penetration service corresponding to each host, and the processes in the network joint that can implement privilege escalation operations;
[0013] The action refers to the vulnerability scanning, penetration, and privilege escalation operations performed by the agent on the target host;
[0014] The reward value is the sum of the values of the sensitive hosts in the network environment minus the sum of the costs of the actions performed by the agent.
[0015] Further, step (2) includes:
[0016] (2.1) Store the state transition process (state s t , action a t , reward r t , next state s t+1 ) in the experience replay pool Buff as the training dataset;
[0017] (2.2) In the form of random sampling, sample N training data from Buff, input the N training data into the DQN to obtain the predicted Q value of the current value network and the target Q value of the target value network. With the goal of minimizing the loss function, update the network parameters of the current value network through the backpropagation of the reverse gradient of the neural network, where the loss function is the mean square error of the predicted Q value and the target Q value:
[0018]
[0019] Among them, is the target Q value;
[0020] (2.3) Repeat step (2.2). During the repetition process of step (2.2), copy the network parameters of the current value network to the target value network every predetermined time to update the target value network until the end of the round training to obtain the current optimal penetration attack path process:
[0021]
[0022] Further, the optimal penetration attack path process in step (2) is to attack the target sensitive host with the greatest value using as few operations as possible:
[0023]
[0024] where s t is the state at time t, a t is the action at time t, s t+1 is the state at the next time after executing a t , γ is the discount factor, and R represents the reward obtained under the current penetration strategy π.
[0025] Further, step (3) further includes:
[0026] Deploy one or more normal hosts in the subnet with more than 5 hosts in the network environment, the previous subnet of the subnet where the target sensitive host is located, or the subnet where the target sensitive host is located as honeypot hosts, set the value of the honeypot hosts to the highest value in the network environment, and modify the end condition of the penetration attack round to obtain the User permission of the honeypot host or the number of training steps in the round reaches a predetermined threshold.
[0027] According to the second aspect of the embodiments of the present application, there is provided a device for defending against penetration attacks based on reinforcement learning, including:
[0028] A modeling module, configured to model the penetration testing process as a Markov decision process, where the Markov decision process includes a state, an action, and a reward value;
[0029] A training module, configured to train an intelligent agent based on the deep Q-network algorithm, where the intelligent agent is used as the penetration attack party, and the training objective is to generate the current optimal penetration attack path process, where the current optimal penetration attack path process is that the intelligent agent uses as few actions as possible to attack the target sensitive host with the greatest value in the current network environment;
[0030] A network environment modification module, configured to flip the sign of the value of the sensitive host in the network environment, and set that the penetration attack round does not end when obtaining the Root permission of the target sensitive host, and modify the end condition of the penetration attack round to that the number of training steps in the round reaches a predetermined threshold;
[0031] A defense training module, configured to update the modification of the host value in the network environment modification module to the reward value in the modeling module, and use the trained agent in the training module to perform defense training on the network environment in the network environment modification module. The process of defense training is repeated until the number of training rounds reaches a predetermined threshold, so as to obtain a strategy for defending against penetration attacks.
[0032] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including:
[0033] One or more processors;
[0034] A memory, configured to store one or more programs;
[0035] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect.
[0036] According to a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in the first aspect are implemented.
[0037] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0038] As can be seen from the above embodiments, in the present application, 1) the penetration test based on the network scenario is modeled as a Markov decision process, and then the agent is trained through the deep Q-network algorithm to perform penetration path optimization; 2) considering the influence of malicious penetration attacks, in order to defend against malicious penetration attacks, combined with the idea of poisoning attacks on the reinforcement learning model, the training data of the model is poisoned. Specifically, by modifying the reward values of sensitive hosts and vulnerable hosts, the purpose of invalidating the strategies of penetration attackers is achieved.
[0039] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0041] Figure 1 Is a flowchart of a method for defending against penetration attacks based on reinforcement learning shown according to an exemplary embodiment.
[0042] Figure 2 Is a schematic diagram of a network scenario shown according to an exemplary embodiment.
[0043] Figure 3 Is a schematic diagram of the DQN algorithm structure used in the method of the present invention.
[0044] Figure 4 is a block diagram of a device for defending against penetration attacks based on reinforcement learning shown according to an exemplary embodiment.
[0045] Figure 5 is a schematic diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0046] Here, the exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application.
[0047] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0048] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0049] Figure 1 is a flowchart of a method for defending against penetration attacks based on reinforcement learning shown according to an exemplary embodiment. As Figure 1 shown, the method may include the following steps:
[0050] (1) Model the penetration testing process as a Markov decision process, where the Markov decision process includes states, actions, and reward values;
[0051] (2) Train an agent based on the deep Q-network algorithm, where the agent acts as a penetration attacker, and the training objective is to generate the current optimal penetration attack path process, where the current optimal penetration attack path process is for the agent to attack the target sensitive host with the maximum value in the current network environment using as few actions as possible in the current network environment;
[0052] (3) Flip the value of the sensitive host in the network environment, and set that the penetration attack round does not end when obtaining the Root permission of the target sensitive host, and modify the end condition of the penetration attack round to that the number of training steps in the round reaches a predetermined threshold;
[0053] (4) Update the modification of the host value in step (3) to the reward value in step (1), and use the trained agent in step (2) to perform defense training on the network environment in step (3). Repeat the process of defense training until the number of training rounds reaches a predetermined threshold to obtain a strategy for defending against penetration attacks.
[0054] As can be seen from the above embodiments, in this application, 1) model the penetration test based on the network scenario as a Markov decision process, and then train the agent to perform penetration path optimization through the deep Q-network algorithm; 2) considering the impact of malicious penetration attacks, in order to defend against malicious penetration attacks, combine the idea of poisoning attacks on the reinforcement learning model, and poison the training data of the model. Specifically, by modifying the reward values of sensitive hosts and vulnerable hosts, the purpose of invalidating the strategy of the penetration attacker is achieved.
[0055] In the specific implementation of step (1), the state is the information observed by the agent in the network environment, including the subnets in the network environment, the hosts in each subnet, the operation types of each host, the penetration services corresponding to each host, and the processes in the network joints that can implement privilege escalation operations; specifically, Figure 2 is a schematic diagram of a network scenario shown according to an exemplary embodiment. Specifically, in this network environment, the state is that there are a total of 8 hosts and 4 different subnets, including hosts of two different operation types, Linux and Windows, and each host runs the corresponding penetration service. This scenario includes three different types of vulnerability services, HTTP, SSH, and FTP. Finally, in this scenario, privilege escalation operations can be achieved through two processes, Tomact and Daclsvc;
[0056] The action refers to the vulnerability scanning, penetration, and privilege escalation operations performed by the agent on the target host; an action can be regarded as an attack vector defined by <c, m>, indicating that the agent performs operation m on host c; in each training step, the size of the operation space of the agent reaches Z(P×Q), where P is the number of hosts in the network and Q represents the number of operations that the attack target can execute;
[0057] The reward value is the sum of the values of the sensitive hosts in the network environment minus the sum of the costs of the actions performed by the agent; specifically, after the agent takes the corresponding operation, the environment will feedback a corresponding reward value, and the calculation of this reward value is as follows:
[0058]
[0059] In the formula, C represents the set of sensitive hosts in the network, M represents the set of actions of the agent, value represents the value of the host, and cost represents the cost for the agent to execute an action.
[0060] In this embodiment, the values of the sensitive and vulnerable hosts at (2, 0) and (4, 0) are defined as 100, while the values of other ordinary hosts are set to 0. In addition, for the penetration actions of different service vulnerabilities, different probabilities and costs are set according to CVSS. Among them, the probability of penetrating the SSH service vulnerability is 0.9, and the cost consumed is 3; the probability of penetrating the FTP service vulnerability is 0.6, and the cost consumed is 1; the probability of penetrating the HTTP service vulnerability is 0.9, and the cost consumed is 2. For other vulnerability scanning and privilege escalation operations, the cost consumed is 1.
[0061] In the specific implementation of step (2), the penetration attacker is trained based on the Deep Q-Network algorithm (DQN) in reinforcement learning. The agent is regarded as the penetration attacker, and the training goal is to successfully penetrate to the target sensitive host in subnet 4 with the fewest steps, so as to obtain an optimal penetration path. This step may specifically include:
[0062] (2.1) Store the state transition process (state s t , action a t , reward r t , next state s t+1 ) in the experience replay pool Buff as the training data set;
[0063] Specifically, DQN improves the traditional Q-learning algorithm during training to alleviate the problem of unstable representation functions in non-linear networks. For example, DQN adopts an experience replay mechanism and uses an experience replay pool to store transition samples. At each time step t, the transition samples obtained by the agent interacting with the environment are stored in the experience replay pool. The state transition process (state s i , action a i , reward r i , next state s' i ) is stored in the experience replay pool Buff as the training data set of the DQN model, and batch learning is performed in the form of random sampling. For the state transition process (state s t , action a t , reward r t , next state s t+1) The agent interacts with the environment to obtain the state. Based on the current state, the agent then performs an action through model training. After the action is executed, the environment will feedback a reward to the agent, and then the agent will obtain the state at the next moment. The process of obtaining these four data can be defined as a training process for one step, which is actually the definition of the Markov decision process in step (1).
[0064] (2.2) In the form of random sampling, sample N training data from the Buff, input the N training data into the DQN, obtain the predicted Q value of the current value network and the target Q value of the target value network. Aiming to minimize the loss function, update the network parameters of the current value network through the backpropagation of the neural network's reverse gradient, where the loss function is the mean square error of the predicted Q value and the target Q value:
[0065]
[0066] Among them, is the target Q value, s' represents the next state that appears after taking action a, and a′ is the possible action in the s′ state. γ is the discount factor, and the larger the discount factor, the more attention is paid to the long-term return.
[0067] Specifically, as Figure 3 shown, DQN combines Q-learning with a convolutional neural network to construct a reinforcement learning training model. Its algorithm feature is to combine the convolutional neural network (CNN) with the Q-Learning algorithm in traditional reinforcement learning to create a new DQN model. The input of the DQN model is the current state. After non-linear transformation through 3 convolutional layers and 2 fully connected layers, finally, a Q value is generated for each action at the output layer. DQN uses the target network mechanism, that is, on the basis of the current value network structure, a target value network with exactly the same structure is built to form the overall model framework of DQN. During the training process, the predicted Q value output by the current value network is used to select action a, and another target value network is used to calculate the target Q value.
[0068] (2.3) Repeat step (2.2). During the repetition of step (2.2), copy the network parameters of the current value network to the target value network every predetermined time to update the target value network until the end of the episode training, obtaining the current optimal penetration attack path process:
[0069]
[0070] Specifically, for the target value network, its network parameters do not need to be iteratively updated. Instead, the network parameters are copied from the current value network every once in a while, that is, delayed update, and then the next round of learning is carried out. According to the Bellman optimal equation theory, as long as step (2.2) is continuously iteratively updated, the target Q value and the predicted Q value can be made infinitely close, so as to finally obtain the optimal penetration strategy. This method reduces the impact of each change in the Q value on the policy parameters, that is, reduces the correlation between the target Q value and the predicted Q value, and increases the stability of policy training.
[0071] In this embodiment, the condition for the end of round training can be reaching a predetermined number of training steps. For a small-scale network with the number of hosts less than or equal to 10, the number of training steps per round can be set to 1000 steps. If the scale of the host increases, the number of training steps per round should also be increased accordingly. This setting is a conventional setting in this field and will not be elaborated here.
[0072] In the specific implementation of step (3), the DQN model can be poisoned by modifying the training data, and the value symbols of sensitive hosts are flipped to reduce the attack performance of the attacker. In this embodiment, the values of sensitive host (2, 0) and sensitive host (4, 0) can be changed from 100 to -100, and when the Root permission of the target sensitive host is obtained, it does not mean the end of the round. 3.2) Even if the intelligent agent penetrates into a sensitive host during the training process, there will be no penetration success flag, and the obtained reward return is very low, making the intelligent agent mistakenly think that this action strategy is an ineffective penetration path optimization strategy, so as to achieve the purpose of reducing the performance of the intelligent agent and affecting the success rate of the penetration test of sensitive hosts;
[0073] In order to further achieve the effect of directly defending against malicious penetration attackers, honeypots can be set up in this network, and high positive rewards can be set for the hosts at the honeypots, so that the penetration attack party will be trapped in the honeypots, achieving the effect of protecting other sensitive hosts and vulnerable hosts. Therefore, step (3) can be further optimized as:
[0074] Flip the signs of the values of sensitive hosts in the network environment, and set that the round of penetration attack does not end when the Root permission of the target sensitive host is obtained. Deploy one or more normal hosts in the subnet with the number of hosts in the network environment exceeding 5, the previous subnet of the subnet where the target sensitive host is located, or the subnet where the target sensitive host is located as honeypot hosts, and set the value of the honeypot hosts to the highest value in the network environment. Modify the end condition of the round of penetration attack to obtaining the User permission of the honeypot host or the number of training steps in the round reaching a predetermined threshold.
[0075] In this embodiment, the (3, 2) normal host in subnet 3 is deployed as a honeypot, and the reward value of the honeypot host is set to 200, while the reward value of the sensitive host is modified to -100. This is because there are two sensitive hosts in the network, and the reward value obtained by finding all honeypot hosts is close to 200. When the reward value of the honeypot host is modified to 200, the training effect of the reward value obtained by the agent during the poisoning training process will be close to the model performance obtained by normal training, making it impossible for the attacker to judge whether the model is poisoned by observing the training process effect.
[0076] The presence of honeypot hosts in the network with honeypot hosts can be used to induce the attacker's strategy. By increasing the value of the honeypot host, the agent is induced to perform attack penetration on the honeypot host, so that it will fall into the honeypot host during the test process, rather than bypassing it to attack the sensitive host, effectively reducing the success rate of the sensitive host penetration test and achieving the purpose of network security protection.
[0077] In the further optimization of step (3), two methods of modifying the training data (i.e., the reward value) and deploying the honeypot are used to attack the malicious penetrator from two dimensions, the surface layer and the deep layer.
[0078] In the specific implementation of step (4), the modification of the values of the sensitive host and the normal host in step (3) is updated to the reward value in step (1), and the defense training is carried out again with the new reward value. Specifically, the agent can only perform penetration between connected subnets or hosts in the network. The end condition for the agent in each round is to obtain the Root permission of all target sensitive hosts or the number of steps in the round training reaches the set maximum value; repeat the training several times until the training round ends to obtain the strategy for defending against penetration attacks. In this embodiment, the number of training rounds is set to 20,000 rounds, which is a conventional setting in this field and will not be elaborated here. Specifically, the strategy for defending against penetration attacks can be obtained by training the DQN model. Through this strategy, the agent trained by the DQN model will choose to bypass the sensitive host or fall into the honeypot host to penetrate the network, thus achieving the purpose of protecting the sensitive host.
[0079] In summary, this application conducts research on the network security of intelligent penetration testing. First, it uses a reinforcement learning algorithm to model the target network environment, and searches for the optimal penetration path through simulating hacker attacks to achieve penetration attacks on the network system. For network security protection, this paper adopts the method of poisoning attacks by modifying the training reward data, and achieves the purpose of security protection by invalidating the strategies of malicious penetration attackers. Then, further use honeypot hosts in the network, modify the reward values in their training data during the intelligent penetration path optimization training process to conduct poisoning attacks on them, induce attackers to penetrate the honeypots, and obtain incorrect target penetration paths, thereby achieving the purpose of protecting the security of vulnerable sensitive hosts in the network. The experimental results show that through the poisoning method of flipping the reward values of sensitive hosts and honeypot hosts, penetration attacks can be effectively defended. In a network scenario with honeypots, the purpose of defending penetration attacks can be achieved by luring attackers to the honeypots.
[0080] Corresponding to the foregoing embodiments of the method for defending against penetration attacks based on reinforcement learning, this application also provides an embodiment of a device for defending against penetration attacks based on reinforcement learning.
[0081] Figure 4 It is a block diagram of a device for defending against penetration attacks based on reinforcement learning shown according to an exemplary embodiment. Refer to Figure 4 , this device may include:
[0082] A modeling module 21, configured to model the penetration testing process as a Markov decision process, where the Markov decision process includes states, actions, and reward values;
[0083] A training module 22, configured to train an agent based on the deep Q-network algorithm, where the agent acts as a penetration attacker, and the training objective is to generate the current optimal penetration attack path process, where the current optimal penetration attack path process is that in the current network environment, the agent uses as few actions as possible to attack the target sensitive host with the maximum value in the network environment;
[0084] A network environment modification module 23, configured to flip the sign of the value of the sensitive host in the network environment, and set that the penetration attack round does not end when obtaining the Root permission of the target sensitive host, and modify the end condition of the penetration attack round to that the number of training steps in the round reaches a predetermined threshold;
[0085] A defense training module 24, configured to update the modification of the host value in the network environment modification module to the reward value in the modeling module, use the trained agent in the training module to perform defense training on the network environment in the network environment modification module, and repeat the defense training process until the number of training rounds reaches a predetermined threshold to obtain a strategy for defending against penetration attacks.
[0086] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0087] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0088] Correspondingly, this application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method for defending against penetration attacks based on reinforcement learning as described above. As Figure 5 shown, it is a hardware structure diagram of a device with any data processing ability where the method for defending against penetration attacks based on reinforcement learning provided by the embodiments of the present invention is located. Except for Figure 5 the processors, memory, and network interfaces shown, any device with data processing ability where the device in the embodiments is located usually includes other hardware according to the actual functions of the device with any data processing ability, which will not be elaborated here.
[0089] Correspondingly, this application also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the method for defending against penetration attacks based on reinforcement learning as described above is implemented. The computer-readable storage medium can be an internal storage unit of any device with data processing ability described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of the wind turbine, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium can also include both the internal storage unit of any device with data processing ability and the external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing ability, and can also be used to temporarily store the data that has been output or will be output.
[0090] Other embodiments of the present application will be readily contemplated by those skilled in the art upon consideration of the specification and practice of the disclosure herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0091] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A method for defending against penetration attacks based on reinforcement learning, characterized in that, Comprising: (1) Modeling the penetration testing process as a Markov decision process, where the Markov decision process includes states, actions, and reward values; (2) Training an agent based on the deep Q-network algorithm, where the agent acts as the penetration attacker, and the training objective is to generate the current optimal penetration attack path process, where the current optimal penetration attack path process is for the agent to use as few actions as possible to attack the most valuable target sensitive host in the current network environment; (3) Reversing the sign of the value of the sensitive host in the network environment, and setting that the penetration attack round does not end when obtaining the Root permission of the target sensitive host, and modifying the end condition of the penetration attack round to that the number of training steps in the round reaches a predetermined threshold; (4) Updating the modification of the host value in step (3) to the reward value in step (1), and using the trained agent in step (2) to perform defense training on the network environment in step (3), and repeating the defense training process until the number of training rounds reaches a predetermined threshold to obtain the strategy for defending against penetration attacks.
2. The method according to claim 1, wherein The state is the information observed by the agent in the network environment, including the subnets in the network environment, the hosts in each subnet, the operation type of each host, the penetration service corresponding to each host, and the processes in the network environment that can achieve privilege escalation operations; The action refers to the vulnerability scanning, penetration, and privilege escalation operations performed by the agent on the target host; The reward value is the sum of the values of the sensitive hosts in the network environment minus the sum of the costs of the actions performed by the agent.
3. The method according to claim 1, wherein Step (2) includes: (2.1) Store the state transition process (state s t , action a t , reward r t , and next state s t+1 ) in the experience replay pool Buff as the training dataset; (2.2) Sampling N training data from the Buff in a random sampling manner, inputting the N training data into the DQN to obtain the predicted Q value of the current value network and the target Q value of the target value network, and aiming to minimize the loss function, updating the network parameters of the current value network through the backpropagation of the reverse gradient of the neural network, where the loss function is the mean square error of the predicted Q value and the target Q value: Among them, is the target Q value; (2.3) Repeating step (2.2), and during the repetition process of step (2.2), copying the network parameters of the current value network to the target value network every predetermined time to update the target value network until the end of the round training to obtain the current optimal penetration attack path process:
4. The method according to claim 1, wherein The optimal penetration attack path process in step (2) is to attack the most valuable target sensitive host with as few operations as possible: where s t is the state at time t, a t is the action at time t, s t+1 is the state at the next time step after executing a t , γ is the discount factor, and R represents the reward obtained under the current penetration strategy π.
5. The method according to claim 1, characterized in that, Step (3) further includes: Deploying the subnets in the network environment with more than 5 hosts, the previous subnet of the subnet where the target sensitive host is located, or one or more normal hosts in the subnet where the target sensitive host is located as honeypot hosts, setting the value of the honeypot host to the highest value in the network environment, and modifying the end condition of the penetration attack round to obtaining the User permission of the honeypot host or the number of training steps in the round reaching a predetermined threshold.
6. A method for defending against penetration attacks based on reinforcement learning, characterized in that, Comprising: A modeling module for modeling the penetration testing process as a Markov decision process, where the Markov decision process includes states, actions, and reward values; A training module for training an intelligent agent based on the deep Q-network algorithm, where the intelligent agent acts as a penetration attacker, and the training objective is to generate the current optimal penetration attack path process, and the current optimal penetration attack path process is that in the current network environment, the intelligent agent uses as few actions as possible to attack the most valuable target sensitive host in the network environment; A network environment modification module for flipping the value of the sensitive host in the network environment, setting that the penetration attack round does not end when obtaining the Root permission of the target sensitive host, and modifying the end condition of the penetration attack round to that the number of training steps in the round reaches a predetermined threshold; A defense training module for updating the modification of the host value in the network environment modification module to the reward value in the modeling module, using the trained intelligent agent in the training module to perform defense training on the network environment in the network environment modification module, and repeating the defense training process until the number of training rounds reaches a predetermined threshold to obtain a strategy for defending against penetration attacks.
7. An electronic device, characterized in that, Comprising: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-5.
8. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the instruction is executed by the processor, the steps of the method according to any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Malicious software image format detection model-oriented black box attack defense method and device thereof
CN110826059A
Network space safety defense method based on dynamic defense graph and reinforcement learning
CN113810406A