Method and device for realizing intelligent penetration test based on deep reinforcement learning, processor and computer readable storage medium thereof

By combining deep reinforcement learning and the Noisy Dueling DQN algorithm, automated penetration path discovery and planning under incomplete information conditions are achieved, solving the problem of low efficiency of traditional penetration testing in complex network environments and improving the automation and accuracy of penetration testing.

CN121077795APending Publication Date: 2025-12-05THE THIRD RES INST OF MIN OF PUBLIC SECURITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511392680.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Traditional penetration testing methods are time-consuming and costly when facing large-scale and complex network environments, making it difficult to adapt to rapidly changing network threats. Furthermore, due to incomplete information, the coverage and accuracy of penetration testing are limited, making it difficult to effectively deal with complex and ever-changing network attack scenarios.

Method used

We employ an intelligent penetration testing method based on deep reinforcement learning. By constructing a partially observable Markov decision process (POMDP) ​​model and combining deep reinforcement learning with the Noisy Dueling DQN algorithm, we achieve automated penetration path discovery and intelligent planning. We use Nmap for port scanning and machine learning algorithms to identify product characteristics, execute vulnerability exploitation, and generate a security test report.

Benefits of technology

It enables efficient penetration path planning under incomplete information conditions, reduces reliance on manual intervention, improves the automation level of penetration testing and the accuracy of security assessment, and can automatically identify hidden vulnerabilities and generate structured reports in complex network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121077795A_ABST
    Figure CN121077795A_ABST
Patent Text Reader

Abstract

The invention relates to a method for realizing an intelligent penetration test based on deep reinforcement learning, and the method comprises the following steps: learning an automatic penetration path discovery algorithm and a penetration path intelligent planning algorithm under an incomplete information condition on a training server through a machine learning model, and autonomously learning a vulnerability utilization strategy; performing port scanning on a target server by using Nmap, and identifying product features which cannot be directly identified through signatures in combination with a machine learning algorithm; initiating a vulnerability utilization attack to the target server; and executing vulnerability utilization and establishing session connection with the target server. By adopting the method and the device for realizing the intelligent penetration test based on deep reinforcement learning, the processor and the computer readable storage medium, automatic penetration path discovery is realized, manual dependence is reduced, efficiency and expandability are improved, and a penetration path of a large-scale network can be efficiently and automatically planned; according to the multi-thread learning synergy method, the learning process of the intelligent agent is accelerated, and the training and learning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of network security, especially to the field of penetration testing, and particularly to a method and device for intelligent penetration testing based on deep reinforcement learning, a processor and a computer readable storage medium thereof. BACKGROUND

[0002] In the field of network security, penetration testing is an effective security assessment method, which aims to evaluate the security of target systems by simulating the behavior of attackers, so as to discover and fix potential security vulnerabilities before real attacks occur.

[0003] Traditional penetration testing methods mainly rely on the experience and manual operation of security experts, which have obvious limitations in the face of large-scale and complex network environments. First, manual penetration testing is time-consuming and costly, and it is difficult to adapt to rapidly changing network threats. Second, due to the incompleteness of information, security experts are difficult to fully understand the state of the target network, which limits the coverage and accuracy of penetration testing. In addition, with the evolution of attack technology, especially the popularity of automated attack tools, traditional penetration testing methods have been difficult to effectively cope with complex and variable network attack scenarios.

[0004] Machine learning is a branch of artificial intelligence, which enables computer systems to automatically learn and improve their performance using data and algorithms without explicit programming instructions. The core of machine learning is to build algorithms that can learn patterns and rules from data and use this knowledge to predict future events or make decisions. In recent years, researchers have begun to explore the use of machine learning and artificial intelligence technology to automate and intelligentize the penetration testing process. Deep reinforcement learning, as an advanced machine learning method, is widely used in automated penetration testing research due to its advantages in handling sequential decision-making problems. We combine deep reinforcement learning technology with automated penetration testing to achieve intelligent discovery and planning of penetration paths in an incomplete information environment, in order to improve the automation level of penetration testing and the accuracy of security assessment. SUMMARY

[0005] The purpose of the present application is to overcome the shortcomings of the prior art, and to provide a method, device, processor and computer readable storage medium for intelligent penetration testing based on deep reinforcement learning, which meets the requirements of good accuracy, high automation level and wide application range.

[0006] In order to achieve the above purpose, the method, device, processor and computer readable storage medium for intelligent penetration testing based on deep reinforcement learning of the present application are as follows: The main feature of the method for intelligent penetration testing based on deep reinforcement learning is that the method comprises the following steps: (1) Running an automatic penetration path discovery algorithm through a machine learning model on the training server, then running an intelligent penetration path planning algorithm under incomplete information conditions, and autonomously learning vulnerability exploitation strategies, while the machine learning model continuously updates neural network weights; (2) Using Nmap to perform port scanning on the target server to obtain information such as operating system type, open ports, product name, and product version, and combining machine learning algorithms to identify product features that cannot be directly identified through signatures; (3) Based on the trained data model and identified product information, launching vulnerability exploitation attacks on the target server; if the target server is not directly accessible, penetrating to the target step by step through intermediate nodes according to the planned path; (4) Executing vulnerability exploitation and establishing a session connection with the target server, generating a security test report containing vulnerability details and exploitation results.

[0007] Preferably, step (1) specifically includes the following steps: (1.1) Running an automatic penetration path discovery algorithm to learn a general attack chain template from any entry to any target in a simulation environment; (1.2) Running an intelligent penetration path planning algorithm under incomplete information conditions, updating state observation every time a host is detected, and dynamically adjusting the subsequent path using the planning algorithm.

[0008] Preferably, step (1.1) specifically includes the following steps: (1.1.1) Constructing a state space for a penetration test automatic path discovery model, which includes the state, configuration information, and termination state of each device in the network; (1.1.2) Constructing an action space, which includes scanning actions and exploitation actions; (1.1.3) Constructing a reward function to calculate the immediate reward of scanning actions or exploitation actions; (1.1.4) Discovering an automatic path based on noise parameterized competitive deep Q learning.

[0009] Preferably, step (1.1.4) specifically includes the following steps: (1.1.4.1) Using the neural network structure of Dueling DQN to change the output of the final Q value into a linear combination of the value function and the advantage function; (1.1.4.2) Adding noise parameters to the network fully connected layer; (1.1.4.3) Using the maximum Q value action obtained by the current Q network to obtain the target Q value in the target Q network.

[0010] Preferably, the step (1.2) specifically comprises the following steps: (1.2.1) Calculate the network information gain and use it as an evaluation index for the actions taken by the penetration test; (1.2.2) Calculate and action cost , For information gain , calculate the reward r of the action taken by the penetration agent; (1.2.3) Learn using the Noisy Double Dueling DQN algorithm.

[0011] Preferably, the step (2) specifically comprises the following steps: (2.1) Call the Nmap script to obtain the port, operating system and service fingerprint raw data of the target server; (2.2) Input the raw data into the service identification sub-model to output a structured vector; (2.3) If the service signature confidence is lower than the preset threshold δ, call the CNN-based implicit feature identifier for secondary classification to complete the missing fields; if it is higher than the preset threshold, directly output the vector; (2.4) Use the final structured vector as the input of the subsequent exploit.

[0012] Preferably, the step (3) specifically comprises the following steps: (3.1) Input the obtained structured vector into the vulnerability matching engine to search the local mapping table; (3.2) If there is a directly exploitable vulnerability, generate a one-time exploit script; if the target is not directly accessible, call the path planner to calculate the shortest attack path; (3.3) If the exploit is successful, mark the host as "controlled" and update the global network state.

[0013] Preferably, the step (4) specifically comprises the following steps: (4.1) Aggregate the exploit chain and record the complete attack path from the entry point to the target host; (4.2) Convert the attack data into a PDF report and output.

[0014] The device for implementing intelligent penetration testing based on deep reinforcement learning, its main features are that the device comprises: a processor configured to execute computer executable instructions; A memory storing one or more computer-executable instructions that, when executed by the processor, implement the steps of the method for intelligent penetration testing based on deep reinforcement learning.

[0015] The processor for intelligent penetration testing based on deep reinforcement learning, wherein the processor is configured to execute computer-executable instructions that, when executed by the processor, implement the steps of the method for intelligent penetration testing based on deep reinforcement learning.

[0016] The computer-readable storage medium, wherein the computer-readable storage medium stores a computer program that can be executed by a processor to implement the steps of the method for intelligent penetration testing based on deep reinforcement learning.

[0017] The method, device, processor and computer-readable storage medium for intelligent penetration testing based on deep reinforcement learning of the present application, by constructing a model based on a partially observable Markov decision process (POMDP) and combining deep reinforcement learning, realizes automatic penetration path discovery, reduces artificial dependence, and improves efficiency and scalability. Secondly, the improved Noisy Dueling DQN algorithm enhances learning efficiency and robustness, so that the agent can learn more stably and effectively in the face of uncertainty and complex environment. Thirdly, the intelligent penetration path planning algorithm of the present application does not require prior knowledge such as network topology and software configuration, and can efficiently and automatically plan the penetration path of a large-scale network. Finally, the multi-thread learning efficiency method speeds up the learning process of the agent and improves the training and learning efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 A schematic diagram of the method for intelligent penetration testing based on deep reinforcement learning of the present application.

[0019] Figure 2 A model structure of the Noisy Duel DQN algorithm of the method for intelligent penetration testing based on deep reinforcement learning of the present application.

[0020] Figure 3 A schematic diagram of the method for intelligent penetration testing based on deep reinforcement learning of the present application.

[0021] Figure 4 A schematic diagram of the method for intelligent penetration testing based on deep reinforcement learning of the present application.

[0022] Figure 5The schematic diagram of the change of the probability of the simulation experiment honeypot host being intruded with the number of steps for the method for realizing intelligent penetration testing based on deep reinforcement learning of the application.

[0023] Figure 6 The schematic diagram of the change of the number of steps used for each round of training for the method for realizing intelligent penetration testing based on deep reinforcement learning of the application. DETAILED DESCRIPTION

[0024] In order to make the technical content of the application clearer, further description will be made below in combination with specific embodiments.

[0025] The method for realizing intelligent penetration testing based on deep reinforcement learning of the application, which comprises the following steps: (1) running an automatic penetration path discovery algorithm on a training server through a machine learning model, then running a penetration path intelligent planning algorithm under incomplete information conditions, and autonomously learning a vulnerability exploitation strategy, while the machine learning model continuously updates neural network weights; (2) performing port scanning on a target server using Nmap to obtain information about the operating system type, open ports, product name and product version, and combining a machine learning algorithm to identify product features that cannot be directly identified through signatures; (3) based on the trained data model and the identified product information, launching a vulnerability exploitation attack on the target server; if the target server cannot be directly accessed, penetrating to the target step by step through an intermediate node according to the planned path; (4) performing vulnerability exploitation and establishing a session connection with the target server, and generating a security test report containing vulnerability details and exploitation results.

[0026] As a preferred embodiment of the application, the step (1) specifically comprises the following steps: (1.1) running an automatic penetration path discovery algorithm to learn a universal attack chain template from any entry to any target in a simulation environment; (1.2) running a penetration path intelligent planning algorithm under incomplete information conditions, updating state observation every time a host is detected, and dynamically adjusting the subsequent path with the planning algorithm.

[0027] As a preferred embodiment of the application, the step (1.1) specifically comprises the following steps: (1.1.1) constructing a state space of a penetration testing automatic path discovery model, wherein the state space comprises the state, configuration information and termination state of each device in the network; (1.1.2) constructing an action space, wherein the action space comprises scanning actions and exploitation actions; (1.1.3) constructing a reward function to calculate the immediate reward of the scanning action or the exploitation action; (1.1.4) discovering an automatic path based on noise parameterized competitive deep Q-learning.

[0028] As a preferred embodiment of the present application, the step (1.1.4) specifically comprises the following steps: (1.1.4.1) adopting the neural network structure of Dueling DQN to change the output of the final Q value into a linear combination of the value function and the advantage function; (1.1.4.2) adding noise parameters to the full connection layer of the network; (1.1.4.3) using the maximum Q value action obtained by the current Q network to obtain the target Q value in the target Q network.

[0029] As a preferred embodiment of the present application, the step (1.2) specifically comprises the following steps: (1.2.1) calculating the network information gain and taking it as an evaluation index of the action taken by the penetration test; (1.2.2) calculating and the action cost , as the information gain , calculating the reward r of the action taken by the penetration intelligent agent; (1.2.3) learning using the Noisy Double Dueling DQN algorithm.

[0030] As a preferred embodiment of the present application, the step (2) specifically comprises the following steps: (2.1) calling the Nmap script to obtain the port, operating system and service fingerprint original data of the target server; (2.2) inputting the original data into the service identification sub-model to output a structured vector; (2.3) if the service signature confidence is lower than a preset threshold δ, calling the CNN-based implicit feature identifier for secondary classification to complete the missing fields; if it is higher than the preset threshold, directly outputting the vector; (2.4) taking the final structured vector as the input of the subsequent vulnerability exploitation.

[0031] As a preferred embodiment of the present application, the step (3) specifically comprises the following steps: (3.1) inputting the obtained structured vector into the vulnerability matching engine to search the local mapping table; (3.2) If there is a directly available vulnerability, a one-time use script is generated; if the target is not directly accessible, a path planner is called to calculate the shortest attack path; (3.3) If the vulnerability is successfully exploited, the host is marked as "controlled", and the global network state is updated.

[0032] As a preferred embodiment of the present application, the step (4) specifically comprises the following steps: (4.1) Summarize the vulnerability exploitation chain and record the complete attack path from the entry point to the target host; (4.2) Convert the attack data into a PDF report and output.

[0033] The device for intelligent penetration testing based on deep reinforcement learning of the present application, wherein the device comprises: a processor configured to execute computer executable instructions; a memory storing one or more computer executable instructions, which, when executed by the processor, implement the steps of the above-mentioned method for intelligent penetration testing based on deep reinforcement learning.

[0034] The processor for intelligent penetration testing based on deep reinforcement learning of the present application, wherein the processor is configured to execute computer executable instructions, which, when executed by the processor, implement the steps of the above-mentioned method for intelligent penetration testing based on deep reinforcement learning.

[0035] The computer readable storage medium of the present application, having a computer program stored thereon, which can be executed by a processor to implement the steps of the above-mentioned method for intelligent penetration testing based on deep reinforcement learning.

[0036] The present application provides an intelligent penetration testing tool, which integrates an automatic penetration path discovery algorithm based on deep reinforcement learning and an intelligent penetration path planning algorithm under incomplete information conditions into a penetration testing agent, and realizes dynamic optimization of penetration paths and efficient decision-making under incomplete information through deep reinforcement learning and intelligent planning technology, thereby significantly improving the penetration testing capability in a complex network environment; combined with machine learning feature analysis and multi-hop path planning, it can automatically identify implicit vulnerabilities and generate a structured report, thereby greatly reducing manual intervention and testing period, and providing an intelligent and highly adaptive solution for network security attack and defense, and completing automatic penetration testing based on machine learning.

[0037] The intelligent penetration testing tool of the present application integrates an automatic penetration path discovery algorithm based on deep reinforcement learning and an intelligent penetration path planning algorithm under incomplete information conditions, and comprises the following steps: Step 1: The system learns an automatic penetration path discovery algorithm and an intelligent penetration path planning algorithm under incomplete information conditions through a machine learning model on a training server, and autonomously learns vulnerability exploitation strategies; the model updates neural network weights to achieve performance optimization; Step 2: The system performs port scanning on the target server using Nmap, obtains information such as operating system type, open ports, product name, and product version, and combines machine learning algorithms to identify product features that cannot be directly identified by signatures; Step 3: The system initiates a vulnerability exploitation attack on the target server based on the trained data model and identified product information; if the target server cannot be directly accessed, it penetrates step by step through intermediate nodes according to the planned path; Step 4: Perform vulnerability exploitation and establish a session connection with the target server, and generate a security test report containing vulnerability details and exploitation results.

[0038] In step 1, the automatic penetration path discovery method is modeled as a partially observable Markov decision process. The principles of partially observable Markov decision are as follows: A partially observable Markov decision process (POMDP) is defined by the tuple . If the system is in state (state space) and the agent performs action (action space), the following results will occur: 1) according to the transition function , the state is transferred; 2) according to the observation function , the observation (observation space) is obtained, and 3) a scalar reward . The agent must find a policy to choose the best action at each step based on its past observations and actions to maximize its future earnings. The expected value of the optimal policy is denoted by .

[0039] 3. The penetration testing automatic path discovery modeling uses the Noisy Dueling DQN algorithm to achieve penetration testing automatic path discovery. The principles of the Noisy Dueling DQN algorithm are as follows: The noisy dueling DQN algorithm combines the techniques of dueling DQN and noisy net to improve the decision-making and exploration ability of the agent in a complex environment. The algorithm decomposes the Q network into a state value function and an advantage function, so that the agent can more accurately evaluate the value of each action relative to other actions, and thus more effectively distinguish the pros and cons of actions when making decisions.

[0040] In step 1, when the information is insufficient, the system realizes the intelligent planning of the penetration path under the condition of incomplete information.

[0041] The intelligent planning of the penetration path, the target of the attacker is to minimize the information entropy of the target network composed of host information entropy and network information entropy.

[0042] The intelligent planning of the penetration path, in step 1, the intelligent planning of the penetration path uses deep Q learning. The principle of deep Q learning is as follows: Q learning (Q-Learning) is one of the most representative algorithms in reinforcement learning. Q represents the quality function of the strategy . , refers to the expected return that can be obtained by taking action in a certain state at a certain time, the algorithm associates all states and actions to form a Q table to store the Q value of state-action pair , and the strategy is formed by selecting the action with the highest Q value in each state.

[0043] In step 3, the bottom layer calls Metasploit to perform penetration testing, and the upper layer uses reinforcement learning technology to improve the penetration success rate and efficiency, realizing highly automated penetration testing.

[0044] The penetration testing agent designs a multi-threaded learning mode, interacts with the target system in each thread to learn, and aggregates and saves the learning results of each thread, thereby improving the training and learning efficiency of the penetration testing agent.

[0045] In the specific embodiments of the present application, the following examples are provided: The present application provides an intelligent penetration testing tool, which mainly includes three parts: an automatic penetration path discovery algorithm based on deep reinforcement learning, an intelligent penetration path planning algorithm under the condition of incomplete information, and an automatic penetration testing system integrating the two theoretical methods.

[0046] First part: Partially Observable Markov Decision Processes: In penetration testing, the actions of an attacker and the transitions of the network state follow Markov properties, i.e., the future state transition only depends on the current state and action, and is independent of the history states. However, in real-world scenarios, an attacker often cannot fully observe the entire state information of the network (e.g., host configurations, security protection policies, or vulnerability details), which makes the penetration testing process a partially observable Markov decision process.

[0047] A partially observable Markov decision process (POMDP) process is defined by a tuple (S, A, T, O, R, γ). If the system is in a state (state space) and an agent performs an action (action space), the following results are produced: 1) according to the transition function , the state is transferred to; 2) according to the observation function , the observation (observation space) is obtained, and 3) a scalar reward . The agent must find a policy to select the best action in each step according to its past observations and actions to maximize its future rewards. The expected value of the optimal policy is denoted by .

[0048] State space construction: The present invention constructs a state space for a penetration testing automatic path discovery model, focusing on dynamic and observable states in the target system rather than static information such as network structure and firewall rules. The state space is mainly used to describe the state of each device in the network, including controlled, reached, and un-reached states to reflect the progress of the attacker's penetration of the network. For uncontrolled hosts, the state space includes software configuration information such as operating systems, servers, and open ports, as well as indicators of vulnerable programs and devices or program crashes. In addition, the present invention introduces a terminal state to represent the abandonment of penetration attacks on the target system, i.e., when the expected return of all potential attack paths is lower than the cost of the attack. Specifically, the state of each host can be regarded as a tuple containing the state values of related programs. The global system state is composed of these host state tuples, with each tuple corresponding to a host in the network. The state space is constructed by enumerating these tuples, thus decomposing the state space in a natural way of programs and machines. This construction method enables the state space to be dynamically updated to reflect changes in the network, such as newly discovered vulnerabilities or patched security patches, thus providing a clear, dynamic, and scalable framework for penetration testing.

[0049] Action Space Construction: The action space defines all possible actions an agent can perform in a network penetration test. It is mainly divided into two categories: scanning and exploitation. Both types of actions target reachable hosts. Scanning actions include operating system detection and port scanning, whose primary purpose is to obtain the target host's configuration information without affecting the host's state. Exploitation actions aim to control the target host by exploiting known vulnerabilities to gain control. For all actions, the outcome is deterministic; each action returns a definite observation indicating whether the exploit was successful, failed, or caused a crash. These outcomes are entirely determined by the target host's configuration.

[0050] Reward function construction: The immediate reward of an action depends on the transition that action causes in the current state. Any scan action or action that utilizes the immediate reward of an action. It can be broken down into three components: ; in, Represents the state Next action And successfully transitioned to the state The reward is the value gained from successfully exploiting the vulnerability. If the conversion does not involve a successful exploit, the reward is 0. Related to the duration of an action, it reflects the time resources required to perform the action, while This is related to the risk of the action being detected by the target system; the higher the risk, the greater the cost. Although and While there may be a correlation between them, there is no direct one-to-one correspondence between them. Therefore, this invention defines these two cost factors separately.

[0051] Automatic path discovery methods based on noise-parameterized competitive deep Q-learning: This project proposes the Noisy Dueling DQN algorithm, such as... Figure 2 When calculating the target Q-value, it employs a Dueling DQN neural network structure, transforming the final Q-value output into a linear combination of the value function and the advantage function, expressed as: ; Based on this, the noise parameters are... Added to the fully connected layer of the network to enhance the model's exploratory capabilities, using This represents the parameters of the current Q-network. The parameters of the target Q-network are represented by the action that yields the maximum Q-value from the current Q-network. This action is then fed into the target Q-network to obtain the target Q-value, calculated using the following formula: ; The loss function is: ; The model structure of the whole algorithm is shown in Figure 2 .

[0052] The input of the Noisy Dueling DQN algorithm is an experience replay pool D, a capacity N, a number of training rounds V, a number of training steps T per round, a target network update frequency C, a set of random noises , Q-network parameters , and target-network parameters .

[0053] The output is an action-value function.

[0054] The specific code is as follows: For Episode = 1 to V do: Initialize the environment state; For step = 1 to T do: Q-network noise sampling ; Select an action ; Execute the action , observe the reward and the next state obtained; Store the sequence to the experience replay pool D; From the experience replay pool D, according to ; Q-network noise sampling , target-network noise sampling ; Calculate the target Q value according to formula (6) ; Calculate the gradient descent and the loss function according to formula (7); Update the target-network parameters every frequency C; ; End For; End For; Second part: Network information gain based penetration path planning modeling: the target of the attacker is to minimize the information entropy of the target network composed of host information entropy and network information entropy. The host information entropy is composed of four parts: 1) operating system (OS) information for describing the information of operating system type, version and language package; 2) application information for describing the information of installed software; 3) port information for describing the services opened by the target host; 4) protection mechanism information for describing the protection mechanism information activated by the target host. The detailed host information can be formalized as a vector , wherein represents the probability distribution of the operating system, and each element represents the probability of the corresponding operating system installed on the host. The other three vectors are not mutually exclusive, which means that a specific computer can have multiple elements at the same time (for example, a host can open ports 80 and 22 at the same time and install Firefox and IE). Therefore, it is necessary to normalize the vector. Given a vector, the exposure state of the target computer can be represented by the information entropy, and the calculation formula is: , wherein is the operating system vector, is the set of the other three vectors. According to the above formula, when the attacker has no information about the target device, the information entropy is very high, and as the attacker learns more and more device information through scanning or attack, the information entropy gradually decreases. When the attacker takes over the device, the uncertainty in the vector no longer exists, and the vector only contains 1 or 0, and the information entropy drops to 0. According to the above process, the information entropy after any attack action will not exceed the information entropy before the action, because the penetration testing action can reduce the uncertainty. Assuming that the exposure state of the target device is . Let and represent the device vector before and after the attack action , and and are the corresponding information entropy. Then the information gain about the attack action is always true.

[0055] Network information gain based penetration path planning based on deep reinforcement learning: the present application uses a deep reinforcement method to realize network information gain based penetration path planning method under the condition of incomplete information. The present application uses Noise Double Dueling DQN algorithm for learning.

[0056] ​The key to implementing network information gain-based penetration path planning using deep reinforcement learning technology is to develop an appropriate reward function. To learn the best attack path planning, the penetration agent rewards attack actions that obtain more information and punishes attack actions that do not contribute to further penetration. Based on the definition of network information entropy, this project uses network information gain as an evaluation index for actions taken in penetration testing, and its formula is as follows: ; Let H (S) be the network information entropy before the action is performed, Let H (S) be the network information entropy after the action is performed. When calculating the network information gain, there can be three situations: 1) After taking an operation (such as operating system detection and port scanning), the uncertainty of the victim computer is reduced but not completely eliminated, so the information gain brought by the action is the difference between the two probability distributions. 2) After the action, the device is controlled, such as a vulnerability exploit; in this case, the information gain is the information entropy of the state before the action. 3) The action has no effect on the state of the victim computer, and the probability distribution after the action is the same as before the action, so the information gain is 0.

[0057] ; The reward for the action taken by the penetration agent will consist of two parts, and , which satisfy . is the information gain , which is used to judge the action. Compared with the original fixed reward, is more flexible, which will guide the agent program of the present application to choose better actions to obtain more cumulative rewards. is the action cost. There are two reasons for setting this item: one is to limit the number of actions to avoid infinite loops; the second is to guide the agent program to find the best attack path as much as possible. Based on the Common Vulnerability Scoring System (CVSS) calculation, the formula is as follows: ; ; where, is the CVSS value, represents the timeliness of the action, is the vulnerability disclosure time. Network information entropy mainly involves the connectivity between computers, and scanning operations are usually used to test the connectivity of computers. Therefore, when the agent program has no available valid action, the present application uses a scanning operation, and when a computer that can be penetrated is available, the total information gain will increase.

[0058] Then the application uses a Noisy Double Dueling DQN algorithm for learning, which combines the reinforcement learning algorithms of Noisy DQN and Dueling DQN, enhances the exploration ability of the traditional D3QN algorithm by introducing a noise network and modifying the reward function, and improves the adaptability of the algorithm to environmental changes, thereby helping to obtain a better strategy.

[0059] However, the above method aims to find effective utilization actions on a specific device and cannot be extended to a network scenario. In order to extend a specific device to a network scenario, an observed device set is created to save the detected devices. When is empty or the action selected by the penetration agent has no effect on the information gain, a scanning operation is performed to discover new available devices. When there are multiple actions that affect the information gain, the action that contributes most to the cumulative reward can be selected according to the strategy . It is difficult to calculate the state transition probability in a manual setting, so the Monte Carlo method is used to estimate the state transition probability during the training phase. In the initial stage of training, the penetration agent will perform a large number of random actions to cover as much of the state space as possible. During the training process, the penetration agent will constantly try new actions to find the best strategy while updating the strategy and value until the optimal strategy is learned.

[0060] An experimental environment is built for experimental verification. A simulated network such as Figure 3 has a total of 7 subnets and 17 hosts. Among them, subnet 1, subnet 4, and subnet 6 are connected to the external network, subnet 2 and subnet 3 are internal networks, and subnet 7 is a honeypot. Firewalls are set up between subnets to block certain traffic, and hosts within a subnet can communicate with each other. Some of the hosts run software with vulnerabilities that attackers can exploit to gain privileges.

[0061] In a simulated attacker behavior scenario, the agent cannot obtain the network topology and host configuration information during the exploration phase, and the agent needs to perform a scanning operation to collect relevant information about the target network and hosts. The space searched by the agent for penetration paths during the training process grows exponentially with the number of hosts in the network, which can be represented by the formula S∈O(N^H), where H represents the number of hosts in the network, and N represents the number of scanning, vulnerability exploitation, and privilege escalation operations that can be performed on each host.

[0062] According to the service type, representative vulnerabilities are selected, and according to the access complexity index of the Common Vulnerability Scoring System, "high", "medium", and "low" are set to 0.2, 0.5, and 0.8 probability values, respectively. The privilege escalation and scanning operations are assumed to have a success probability of 1 in this scenario.

[0063] The basis for the agent's learning is the reward value, which is defined as the value of the host minus the cost of the action. The cost of the action is a comprehensive quantification of the time, skill, and monetary costs. The reward value is calculated as follows: ; Where H represents the set of all hosts that the agent successfully compromises, and A represents the set of all actions taken by the agent. Based on such a reward value, the learning goal of the agent is to attack the highest value host with the least number of operations as possible.

[0064] In the initial stage, sensitive hosts are assigned positive reward values, while honeypot hosts are assigned negative reward values. The agent starts in an Internet environment and only knows the subnet topology structure connected to the Internet. Through successive exploration stages, the agent learns to simulate the behavior strategy of an attacker. Each round of exploration can end due to one of the following three situations: obtaining ROOT permission of all sensitive hosts; reaching the preset maximum number of training steps; penetrating into a honeypot host.

[0065] Under the incentive of the reward value, the agent learns how to achieve the maximum cumulative reward value strategy through repeated trial and error. At the beginning, due to lack of learning, the agent tends to randomly select actions in an unknown environment, which results in a large number of steps used in each round of training, a low cumulative reward value obtained, and a possible penetration into a honeypot host. However, with continuous exploration and learning, behaviors that result in a decrease in cumulative reward value (e.g., excessive operations or penetration into a honeypot host) will not be adopted. Ultimately, the agent will learn to take the least number of operations to attack sensitive hosts with positive reward values.

[0066] The six algorithms are compared under the same scenario using the same hyperparameters. The present invention uses DQN as a benchmark to measure the average reward value change of these algorithms under a fixed number of training steps (see Figure 4 ), the probability of a honeypot host being compromised (see 5), and the training step change required for each round (see Figure 6 ).

[0067] Part III: Metasploit: Metasploit is an open-source security vulnerability detection tool that can help identify security issues, verify vulnerability mitigation measures, and manage expert-driven security assessments to provide real security risk intelligence. These features include intelligent development, code auditing, web application scanning, social engineering, and more.

[0068] The intelligent penetration tool integrates the automatic penetration path discovery algorithm based on deep reinforcement learning and the intelligent penetration path planning algorithm under the condition of incomplete information, and provides a solution for complex penetration testing.

[0069] The technical solution adopts a partial observable Markov decision process (POMDP), and explicitly processes incomplete information, and the modeling level difference cannot be solved by simple splicing.

[0070] The Noisy Dueling series algorithm of the technical solution is a multi-dimensional improvement of the traditional DQN, that is, noise improves exploration, and a competitive structure improves value evaluation, and belongs to the innovation of the algorithm architecture.

[0071] The technical solution can realize POMDP modeling under incomplete information, extend MDP to POMDP, introduce an observation space O and an observation function O(s',a,o), and quantify the state uncertainty; the network information entropy (host information entropy + network connection entropy) is used to represent the information completeness, and the scanning action is used to supplement the information.

[0072] The technical solution does not need prior topology adaptive path planning, and does not need prior knowledge such as network topology and software configuration, and gradually constructs a state space through online scanning and utilization actions.

[0073] The technical solution has a parallel experience replay mechanism, adopts multiple independent threads to interact with the environment in parallel, shares an experience replay buffer, and improves sampling efficiency.

[0074] In summary, the intelligent penetration tool proposed in the present application integrates the automatic penetration path discovery algorithm based on deep reinforcement learning and the intelligent penetration path planning algorithm under the condition of incomplete information. In the simulation experimental environment, the target machine can be successfully attacked.

[0075] The specific implementation scheme of the embodiment can be referred to the related description in the above embodiment, which will not be repeated here.

[0076] It can be understood that the same or similar parts in the above embodiments can be mutually referred to, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.

[0077] It should be noted that, in the description of the present application, the terms "first", "second", etc. are only used for the purpose of description, and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality of" is at least two.

[0078] Any processes or methods described in the flow charts or elsewhere herein can be understood as representing one or more steps of a method implemented by one or more computers or computer modules, and the scope of embodiments of the present application includes additional implementation in which one or more steps are performed by a computer or computers, and / or in which one or more steps are performed by a human computer interface.

[0079] It should be understood that parts of the present application can be implemented in hardware, software, firmware or a combination thereof. In the above embodiments, several steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution device. For example, if implemented in hardware, and as in another embodiment, any of the following technologies, known in the art, or a combination thereof, can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.

[0080] Those skilled in the art can understand that all or part of the steps carried out by the above-mentioned embodiments can be instructed by a program to complete the relevant hardware, and the corresponding program can be stored in a computer readable storage medium, and the program when executed includes one of the steps of the method embodiment or a combination thereof.

[0081] In addition, each functional unit in each embodiment of the present application can be integrated into one processing module, or each unit can exist physically separately, or two or more units can be integrated into one module. The integrated module can be realized in the form of hardware or in the form of a software functional module. The integrated module, if realized in the form of a software functional module and sold or used as an independent product, can also be stored in a computer readable storage medium.

[0082] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0083] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in one or more embodiments or examples.

[0084] The method, device, processor and computer readable storage medium thereof for realizing intelligent penetration testing based on deep reinforcement learning adopt the method, device, processor and computer readable storage medium thereof for realizing intelligent penetration testing based on deep reinforcement learning, construct a model based on a partially observable Markov decision process (POMDP), and combine deep reinforcement learning to realize automatic penetration path discovery, reduce artificial dependence, and improve efficiency and scalability. Secondly, the improved Noisy Dueling DQN algorithm enhances learning efficiency and robustness, so that the agent can learn more stably and effectively in the face of uncertainty and complex environment. Thirdly, the intelligent penetration path planning algorithm of the present application does not require prior knowledge such as network topology and software configuration, and can efficiently and automatically plan the penetration path of a large-scale network. Finally, the multi-thread learning efficiency method speeds up the learning process of the agent and improves the training and learning efficiency.

[0085] In this specification, the present application has been described with reference to its specific embodiments. However, it is obvious that various modifications and changes can be made without departing from the spirit and scope of the present application. Therefore, the specification and drawings should be considered illustrative rather than limiting.

Claims

1. A method for implementing intelligent penetration testing based on deep reinforcement learning, characterized in that, The method comprises the following steps: (1) running an automatic penetration path discovery algorithm on a training server through a machine learning model, then running an intelligent penetration path planning algorithm under incomplete information conditions, and autonomously learning vulnerability exploitation strategies, while the machine learning model continuously updates neural network weights; (2) performing port scanning on the target server using Nmap to obtain information about the operating system type, open ports, product name and product version, and combining the machine learning algorithm to identify product features that cannot be directly identified through signatures; (3) based on the trained data model and the identified product information, launching a vulnerability exploitation attack on the target server; if the target server cannot be directly accessed, gradually penetrating to the target through intermediate nodes according to the planned path; (4) performing vulnerability exploitation and establishing a session connection with the target server to generate a security test report containing vulnerability details and exploitation results.

2. The method for implementing intelligent penetration testing based on deep reinforcement learning according to claim 1, characterized in that, The step (1) specifically comprises the following steps: (1.1) running an automatic penetration path discovery algorithm to learn a general attack chain template from any entry to any target in a simulation environment; (1.2) running an intelligent penetration path planning algorithm under incomplete information conditions, updating state observation every time a host is detected, and dynamically adjusting the subsequent path with the planning algorithm.

3. The method for implementing intelligent penetration testing based on deep reinforcement learning according to claim 2, characterized in that, The step (1.1) specifically comprises the following steps: (1.1.1) constructing a state space of a penetration test automatic path discovery model, wherein the state space comprises the state, configuration information and termination state of each device in the network; (1.1.2) constructing an action space, wherein the action space comprises scanning actions and exploitation actions; (1.1.3) constructing a reward function to calculate the immediate reward of the scanning action or exploitation action; (1.1.4) discovering an automatic path based on a noise parameterized competitive deep Q learning.

4. The method for implementing intelligent penetration testing based on deep reinforcement learning according to claim 3, characterized in that, The step (1.1.4) specifically comprises the following steps: (1.1.4.1) adopting a neural network structure of Dueling DQN to change the output of the final Q value into a linear combination of the value function and the advantage function; (1.1.4.2) noise parameters to the network fully connected layer; (1.1.4.3) using the maximum Q value action obtained by the current Q network to obtain the target Q value in the target Q network.

5. The method for implementing intelligent penetration testing based on deep reinforcement learning according to claim 2, characterized in that, The step (1.2) specifically comprises the following steps: (1.2.1) calculating network information gain and taking it as an evaluation index of the action taken in the penetration test; (1.2.2) Computing and action costs , as information gain , compute the reward r for the action taken by the penetrating agent; (1.2.3) learning using the Noisy Double Dueling DQN algorithm.

6. The method for implementing intelligent penetration testing based on deep reinforcement learning according to claim 1, characterized in that, The step (2) specifically comprises the following steps: (2.1) calling an Nmap script to obtain the port, operating system and service fingerprint raw data of the target server; (2.2) inputting the raw data into a service identification sub-model to output a structured vector; (2.3) if the service signature confidence is lower than a preset threshold δ, calling a CNN-based implicit feature identifier for secondary classification to complete the missing fields; if it is higher than the preset threshold, directly outputting the vector; (2.4) taking the final structured vector as the input of subsequent vulnerability exploitation.

7. The method for implementing intelligent penetration testing based on deep reinforcement learning according to claim 1, characterized in that, The step (3) specifically comprises the following steps: (3.1) input the structured vector into the vulnerability matching engine to search the local mapping table; (3.2) if there is a directly exploitable vulnerability, generate a one-time exploit script; if the target is not directly accessible, call the path planner to calculate the shortest attack path; (3.3) if the vulnerability exploitation is successful, mark the host as "controlled" and update the global network state.

8. The method for implementing intelligent penetration testing based on deep reinforcement learning according to claim 1, characterized in that, The step (4) specifically includes the following steps: (4.1) aggregate the vulnerability exploitation chain and record the complete attack path from the entry point to the target host; (4.2) convert the attack data into a PDF report and output.

9. An apparatus for implementing intelligent penetration testing based on deep reinforcement learning, characterized in that, The device comprises: a processor configured to execute computer executable instructions; a memory storing one or more computer executable instructions, which, when executed by the processor, implement the steps of the method for implementing intelligent penetration testing based on deep reinforcement learning according to any one of claims 1 to 8.

10. A processor for implementing intelligent penetration testing based on deep reinforcement learning, the processor comprising: The processor is configured to execute computer executable instructions, which, when executed by the processor, implement the steps of the method for implementing intelligent penetration testing based on deep reinforcement learning according to any one of claims 1 to 8.

11. A computer readable storage medium, characterized in that, A computer program is stored thereon, which can be executed by a processor to implement the steps of the method for implementing intelligent penetration testing based on deep reinforcement learning according to any one of claims 1 to 8.