An automated Windows domain penetration method based on reinforcement learning
By applying reinforcement learning technology in Windows domain environment, analyzing the interaction between penetration tools and the target environment, and optimizing attack path generation, the problem of difficulty in efficiently discovering attack paths in multi-target scenarios in the existing technology is solved, and a more efficient and trustworthy attack path generation is achieved.
Patent Information
- Application Number
- CN202210108140.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-01-28
AI Technical Summary
The prior art is difficult to efficiently discover attack paths in multi-target scenarios, and fails to fully consider the correlation of hosts in the domain environment, resulting in difficulty in meeting the solution efficiency and path trustworthiness in complex domain environments.
Using an automated Windows domain penetration method based on reinforcement learning, we analyze the attack paths generated by the interaction between the penetration tool and the target environment through reinforcement learning, and combine the strategic characteristics and internal structure of the domain environment to optimize the path generation efficiency and ensure the effectiveness of the attack path.
It improves the efficiency and effectiveness of attack path generation in Windows domain environments, can more accurately discover the optimal attack path, and enhances the degree of automation of the complete process from information collection to vulnerability attacks.
Smart Images

Figure CN114444086B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of network attack and defense, and specifically relates to an automated Windows domain penetration method based on reinforcement learning. Background Art
[0002] With the development of the information age, office networks are becoming increasingly complex and integrated. The workgroup model that manages each host independently is gradually unable to meet business needs. More and more companies and institutions use Windows domains to build office networks. Windows domain is an organizational form of computer network. While providing convenience, it also brings huge security risks. First, all hosts in the domain have the right to interact with the domain controller, and the location of the domain controller is easily exposed and attacked; second, each domain account usually has the authority of multiple domain hosts. Using a host with weaker protection as a springboard, an attacker can easily penetrate other hosts in the domain. When the host has domain administrator credentials (with domain controller permissions), the attacker will be able to completely penetrate the entire domain. In addition, native vulnerabilities in the domain environment may also lead to the loss of high privileges: the domain uses special authentication protocols (such as Kerberos, NTLM) and permission policies (such as delegation) to complete identity authentication and permission allocation. When the protocol has defects and the policy is abused (such as CVE-2014-6324, CVE-2014-8812, etc.), attackers can bypass the identity authentication process to steal the credentials of high-privileged users, thereby seriously threatening the security of the entire domain network.
[0003] In order to proactively discover security risks in the system and more clearly analyze the possible actions that intruders may take, researchers proposed using an attack tree to express the dependency between attack behaviors and attack steps. Each node of the tree represents an attack behavior or a sub-goal, and the root node represents the final goal of the attack, depicting the entire process of the intruder from launching an attack to obtaining the final goal. Its limitation is that it can only perform path discovery for a single target and is not suitable for multi-target scenarios. In order to solve this problem, researchers proposed a path discovery method based on an attack graph model, an intelligent solution based on the multi-stage vulnerability analysis language MulVAL and the Q learning algorithm, an attack path generation method based on the deep reinforcement learning algorithm DQN, and a path sequence search method that formalizes the penetration test process into an MDP. However, these existing methods still have the contradiction that it is difficult to simultaneously meet the solution efficiency and path credibility, and do not consider the correlation between hosts in the target network. They are not suitable for domain environments with close host connections and strong network integrity. Summary of the invention
[0004] In view of the problems existing in the above-mentioned background technology, the present invention proposes an automated Windows domain penetration method based on reinforcement learning, which uses reinforcement learning to analyze the attack path generated by the interaction between the penetration tool and the target environment, and fully considers the policy characteristics of the domain environment and the internal structure of the domain network, thereby improving the efficiency of path generation while ensuring the effectiveness of the attack path.
[0005] An automated Windows domain penetration method based on reinforcement learning, characterized in that the method comprises the following steps:
[0006] Step 1: Based on information collection tools, vulnerability scanning tools and vulnerability frameworks, supplemented by the extension modules of the corresponding tools, Python language is used as the connection script between tools and tools, and between tools and environments to build a penetration testing platform;
[0007] Step 2: Use the penetration testing platform to collect information and scan vulnerabilities in the target environment, and match the vulnerability attack module used by each host based on the results;
[0008] Step 3: Use the penetration testing platform to automatically perform reinforcement learning modeling;
[0009] Step 4: Use the penetration testing platform to automatically call the attack module to attack, and perform reinforcement learning training and state simplification based on the returned results to continuously optimize the optimal attack path.
[0010] Furthermore, step 1 is subdivided into the following sub-steps:
[0011] Step 1-1: The information collection tool is based on the existing tool nmap, the vulnerability scanning tool uses the commercial tool AWVS, and the vulnerability framework is based on the open source penetration testing tool set metasploit-framework (MSF);
[0012] Step 1-2, expand MSF through vulnerability exploit code, standardize the codes in different programming languages, and rewrite them in ruby;
[0013] Steps 1-3, use Python scripts to connect the reinforcement learning module with information collection, vulnerability scanning, and vulnerability attacks. When reinforcement learning decides to take an action, directly call the corresponding tool's related functions to implement it;
[0014] Steps 1-4 complete the establishment of the entire penetration testing platform, which has three tools: information collection, vulnerability scanning, and vulnerability exploitation. Reinforcement learning is connected to the above three types of tools through the Python interface. When reinforcement learning completes training, automated penetration testing is achieved.
[0015] Furthermore, step 2 is subdivided into the following sub-steps:
[0016] Step 2-1, use the real domain network as the experimental environment, access the penetration test platform built in step 1, and give the permission of a host in the target domain network. Through this host as a springboard, the platform will first use the tool nmap and its extended functions, and automatically call the python script to collect information and scan vulnerabilities of all hosts and domain users in the target domain network, and obtain the IP, operating system, port: service, and vulnerability information of each host;
[0017] Step 2-2, for the vulnerability scanning results of each host in step 2-1, use the corresponding vulnerability exploitation conditions to match according to the host's operating system, port: service, and determine whether the scanning tool has a false positive. Only the vulnerabilities that meet the vulnerability exploitation conditions are selected as real attack actions by reinforcement learning.
[0018] Furthermore, step 3 is subdivided into the following sub-steps:
[0019] Step 3-1, the reinforcement learning is expressed in the form of MDP quadruple (S, A, R, P). Through continuous exploration and trial and error in the environment, positive and negative rewards are obtained, and the priority of the action is determined based on the reward. The different action selections a under different states s of reinforcement learning are called strategies. The learning goal is to optimize the strategy π to maximize the cumulative reward. The reinforcement learning strategy is solved using the Q learning algorithm. The algorithm uses the action and value function Q(s,a) to represent the strategy, and selects actions based on the expected size of the function of Q(s,a);
[0020] Step 3-2, based on the description of reinforcement learning in step 3-1, construct a four-tuple (S, A, R, P), and use the Q-learning algorithm to solve the optimal strategy for the target environment.
[0021] Furthermore, step 3-2 is subdivided into the following sub-steps:
[0022] Step 3-2-1, construct state space S: use host information and currently obtained target host permissions to define state space. According to the information collected in step 2-1, the state of each host is expressed in the form of "IP-operating system-port: service-existing vulnerability-acquired permissions", where only permissions change with the penetration process. The state space S is enumerated in the form of "state number-acquired permissions". In addition, the state of the successfully penetrated domain controller is set to the target state G. If host n is a domain controller, the state space is enumerated as a set of "1:x-2:x-3:x-4:x-5:x...n-1:xn:y", where x is N, L, U; y is Null or G, where N, L, and U represent the three states that the host may be in: no permissions, local administrator permissions, and domain ordinary user permissions. G represents the current final state, and Null represents the final state that has not been reached.
[0023] Step 3-2-2, construct action space A: The action space represents the vulnerability exploitation actions used by the attacker, which comes from the penetration testing platform; the vulnerability exploitation module is subdivided into credential exploitation module and general vulnerability module; during automated penetration, in addition to exploiting traditional vulnerabilities for penetration, lateral movement is also performed through domain user credentials;
[0024] Step 3-2-3, construct the reward function R: Reinforcement learning relies on the reward after interacting with the environment to make decisions; after performing any action, the agent obtains the immediate reward of the action according to the reward function and predicts the expected cumulative reward; the reward comes from the positive authority increase ΔR and the negative vulnerability loss r cost ; If each permission acquisition action generates permission gain R PA , according to the different permission status of each host, its score r is:
[0025] The target host permission is not obtained: r=0;
[0026] Get the normal user rights of the target host: r=8;
[0027] Obtain local administrator privileges on the target host: r=10;
[0028] Obtain domain administrator privileges: r = 100 + r extra ;
[0029] If the permission scores of a host before and after the vulnerability exploitation action are r1 and r2, the positive reward of the action is its permission gain: ΔR = r2-r1 (r2>r1). Since the domain administrator permission is equivalent to the highest permission of all machines in the domain, the extra score compensation r is obtained when the target state is reached. extra , whose value is equal to the sum of the scores obtained by upgrading all hosts in the domain environment to local administrator privileges;
[0030] r cost By Vulnerability Level V rank 、Vulnerability Availability exploitble , the time when the vulnerability is exposed T vuln The three factors are: the base score of the Common Vulnerability Scoring System (CVSS), the usability score, and the time since the vulnerability was released. cost It is expressed by the following formula:
[0031] r cost =10sigmoid(T vuln )-V exploitble V rank / 10
[0032] where r cost The minimum value is 0;
[0033] The reward for the permission acquisition action is the permission gain ΔR and the vulnerability consumption r cost Difference:
[0034] R PA =ΔR-r cost
[0035] Step 3-2-4, construct state transition P: Based on the actual penetration testing tool interacting with the real environment, explore the state transition through the interaction of the experimental environment; when reinforcement learning selects an action and executes it through the penetration testing platform, the platform will receive state change information fed back by the environment and explore the state transition accordingly.
[0036] Furthermore, step 4 is subdivided into the following sub-steps:
[0037] Step 4-1, the reinforcement learning model constructed in step 1-3 is trained by using the penetration test platform, that is, optimizing the strategy for this environment until the optimal strategy is found; at the beginning of automated penetration, the reinforcement learning algorithm cannot effectively generate the optimal path. At this time, the action is randomly selected to attack the host in the target network through the vulnerability attack module connected by the python script, and the information of reinforcement learning rewards and state transfer is recorded, and the reinforcement learning strategy is updated; when the target state is reached and the penetration test is completed, the platform will initialize the environment information and start the next round of penetration;
[0038] Step 4-2, in order to improve the penetration efficiency of the algorithm, the penetration problem will be simplified during the reinforcement learning training process. In a domain environment, penetrating a host has two main purposes, namely, intranet penetration or stealing unknown domain account credentials. The former helps to break through the restrictions of firewalls and routing, and the latter is used for lateral movement between domains. Hosts that cannot provide the above two helps are difficult to advance the penetration process. In other words, the actual benefit of penetrating any host is the same as the benefit of penetrating all hosts. Therefore, this type of host is considered to meet the state simplification conditions, and all its related states and vulnerability exploitation actions will be deleted;
[0039] Step 4-3, after a certain number of rounds, the reinforcement learning algorithm will select actions. Similarly, the rewards and state transition information of the environment feedback after the attack will be used for reinforcement learning training. Finally, after a certain number of iterative cycles, the algorithm converges, and the platform outputs the optimal path that has been verified and optimized multiple times.
[0040] Compared with the prior art, this method has the following beneficial effects:
[0041] (1) The constructed penetration testing platform includes existing mainstream penetration tools, such as nmap, AWVS, MSF, etc. At the same time, some supplements have been made to these tools based on actual conditions, such as rewriting some ruby codes for 1day vulnerabilities and expanding the vulnerability exploitation module of the MSF platform. In addition, the above tools are linked together using python tools, and action selection is performed through reinforcement learning to achieve the complete process from information collection to vulnerability attack to completion of the test. In the process of interacting with the environment, experience is continuously accumulated, and finally the optimal attack path is obtained. This method can effectively increase the penetration efficiency and provide protection ideas for domain administrators from the attacker's perspective.
[0042] (2) The penetration method based on user credentials in domain penetration is combined with the traditional penetration method to solve the problem that existing research relies entirely on host vulnerabilities for path discovery, thereby improving the applicability and attack effectiveness of this method in the domain environment.
[0043] (3) Redundant hosts are defined based on their contribution to the penetration process, which reduces unnecessary states in reinforcement learning and improves the efficiency of reinforcement learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 4 is an overall flow chart of the infiltration method in an embodiment of the present invention.
[0045] Figure 2 It is a schematic diagram of the specific functions of different tools of the penetration platform in an embodiment of the present invention.
[0046] Figure 3The present invention provides a flowchart of a specific method for constructing a reinforcement learning model for the penetration platform in an embodiment of the present invention.
[0047] Figure 4 Schematic diagram of the general process of interaction between the infiltration platform and the environment in an embodiment of the present invention. DETAILED DESCRIPTION
[0048] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings.
[0049] The present invention is a penetration method based on reinforcement learning. Firstly, the penetration test process is modeled by reinforcement learning, and the attack path is discovered through the real interaction between reinforcement learning and the domain environment. Secondly, based on the difference in the contribution of the host to the penetration process, the unnecessary states in the reinforcement learning are reduced, and the path discovery efficiency is improved. Finally, the optimized Q-learning algorithm is used to screen the optimal attack path, and all possible high-risk security risks in the domain environment are automatically verified, providing a security protection basis for domain administrators.
[0050] Specifically, a Windows domain penetration method based on reinforcement learning, such as Figure 1 As shown, the following steps are included:
[0051] Step 1: Based on information collection tools, vulnerability scanning tools and vulnerability frameworks, collect extension modules of corresponding tools as supplements, use Python language as the connection script between tools and tools, tools and environments, and build a penetration testing platform. Specifically:
[0052] Step 1-1: The information collection tool is based on the existing tool nmap, the vulnerability scanning tool uses the commercial tool AWVS, and the vulnerability framework is based on the open source penetration testing toolset metasploit-framework (MSF).
[0053] Step 1-2: Since MSF cannot fully cover the mainstream vulnerabilities in the current environment, some vulnerability exploitation codes are collected from open source code storage websites such as GitHub to expand MSF. Based on the different programming languages used by these codes, these codes are standardized and rewritten in the Ruby language.
[0054] Steps 1-3, in order to achieve automation, the platform needs to call actions based on the current penetration status. This part can be well implemented by reinforcement learning, but reinforcement learning must also be able to automatically call other tools in the platform to perform related work. Therefore, a python script is used to connect the reinforcement learning module with information collection, vulnerability scanning, and vulnerability attacks. When reinforcement learning decides to take an action, it will be able to directly call the corresponding tool to implement the relevant functions.
[0055] Steps 1-4, in summary, the platform has three tools for information collection, vulnerability scanning and vulnerability exploitation. Through the Python interface, reinforcement learning is connected to the above three types of tools. When reinforcement learning completes training, it can realize automated penetration testing.
[0056] Step 2: Use the penetration testing platform to collect information and scan vulnerabilities in the target environment, and match the vulnerability attack modules that each host may use based on the results, such as Figure 2 As shown, specifically:
[0057] Step 2-1, use the real domain network as the experimental environment, connect to the penetration test platform built in step 1, and give the permission of a host in the target domain network. Through this host as a springboard, the platform will first use the tool nmap and its extended functions, and automatically call the python script to collect information and scan vulnerabilities of all hosts and domain users in the target domain network, and obtain the IP, operating system, port: service, and vulnerability information of each host.
[0058] Step 2-2, for the vulnerability scanning results of each host in step 2-1, use the corresponding vulnerability exploitation conditions to match according to the host's operating system, port: service, in order to determine whether the scanning tool has a false positive. Only the vulnerabilities that meet the conditions will be selected as real attack actions by reinforcement learning. For example, the information collection tool shows that the operating system of the target host A is Windows Server 2008 for x64-based Systems Service Pack 2, and it is detected that the host has opened the SMB service on port 445. In addition, the vulnerability scanning tool finds that host A may have the CVE_2017_0143 vulnerability. Based on this result, the platform will match the basic information of the current host A with the exploitation conditions of the CVE_2017_0143 vulnerability in the database, and determine that the exploitation conditions are consistent with its basic situation, so the vulnerability is listed as a possible action to be taken against host A; in the case of the same vulnerability scanning result, if host A only opens the SSH service on port 22, since this information does not match the exploitation conditions of the CVE_2017_0143 vulnerability, the vulnerability scanning result is listed as a false positive, and the vulnerability will not be used to attack host A.
[0059] Step 3: Use a penetration testing platform to automatically perform reinforcement learning modeling, such as Figure 3 As shown, specifically:
[0060] Step 3-1, based on step 2, reinforcement learning is represented in the form of an MDP quadruple (S, A, R, P), through which positive and negative rewards are obtained through continuous exploration and trial and error in the environment, and the priority of the action is determined based on the reward. The different action selections a of reinforcement learning under different states s are called strategies. The ultimate goal of learning is to optimize the strategy π to maximize the cumulative reward. The embodiment of the present invention uses a Q learning algorithm to solve the reinforcement learning strategy. The algorithm uses an action and value function Q(s, a) to represent the strategy, and selects actions based on the expected size of the function of Q(s, a).
[0061] Step 3-2, based on the description of reinforcement learning in step 3-1, construct a four-tuple (S, A, R, P), and use the Q-learning algorithm to solve the optimal strategy for the target environment.
[0062] (1) Constructing the state space S: Use the host information and the currently obtained permissions of the target host to define the state space. According to the information collected in step 2-1, the state of each host can be expressed in the form of "IP-operating system-port: service-existing vulnerability-acquired permissions". Among them, only the permission state will change with the change of the penetration process. Therefore, the state space S is enumerated in the form of "state number-acquired permissions". In addition, the state of the successfully penetrated domain controller is set to the target state G. If host n is a domain controller, the state space is enumerated as a set of "1:x-2:x-3:x-4:x-5:x...n-1:xn:y", where x is N, L, U; y is Null or G, where N, L, and U represent the three states that the host may be in: no permissions, local administrator permissions, and domain ordinary user permissions. G represents the current final state, and Null represents the final state that has not been reached.
[0063] (2) Constructing action space A: The action space represents the possible vulnerability exploitation actions that attackers may use. It comes from the penetration testing platform. In order to improve the efficiency of domain penetration and optimize action selection, the vulnerability exploitation module is further subdivided into: credential exploitation module and general vulnerability module. During automated penetration, in addition to exploiting traditional vulnerabilities for penetration, lateral movement can also be performed through domain user credentials.
[0064] (3) Constructing the reward function R: Reinforcement learning relies on the rewards after interacting with the environment to make decisions. After performing any action, the agent will get the immediate reward of the action according to the reward function and predict the expected cumulative reward. The reward comes from the positive authority increase ΔR and the negative vulnerability loss r cost If each permission acquisition action may generate permission gain R PA , according to the different permission status of each host, its score r is:
[0065] The target host permission is not obtained: r=0;
[0066] Get the normal user rights of the target host: r=8;
[0067] Obtain local administrator privileges on the target host: r=10;
[0068] Obtain domain administrator privileges: r = 100 + r extra ;
[0069] If the authority scores of a host before and after the vulnerability exploitation action are r1 and r2, the positive reward of the action is its authority gain: ΔR = r2-r1 (r2>r1). Since the domain administrator authority is equivalent to the highest authority of all machines in the domain, an additional score compensation r will be obtained when the target state is reached. extra , whose value is equal to the sum of the scores obtained by upgrading all hosts in the domain environment to local administrator privileges.
[0070] r cost By Vulnerability Level V rank 、Vulnerability Availability exploitble , the time when the vulnerability is exposed T vuln (years), whose values come from the basic score of the Common Vulnerability Scoring System CVSS, the usability score, and the length of time the vulnerability was released. cost It can be expressed by the following formula:
[0071] r cost =10sigmoid(T vuln )-V exploitble V rank / 10, where r cost The minimum value is 0.
[0072] The reward for the permission acquisition action is the permission gain ΔR and the vulnerability consumption r cost Difference:
[0073] R PA =ΔR-r cost
[0074] (4) Constructing state transition P: This method is based on the interaction between the actual penetration testing tool and the real environment, so it is necessary to explore possible state transitions through the interaction of the experimental environment. When reinforcement learning selects an action and executes it through the penetration testing platform, the platform will receive state change information fed back by the environment and explore state transitions accordingly.
[0075] Step 4: Based on step 3, use the penetration testing platform to automatically call the attack module to attack, and perform reinforcement learning training and state simplification based on the returned results to continuously optimize the optimal attack path, such as Figure 4 As shown, specifically:
[0076] Step 4-1: The reinforcement learning model constructed in step 1-3 is trained using the penetration testing platform, that is, the strategy for this environment is optimized until the optimal strategy is found. At the beginning of automated penetration, the reinforcement learning algorithm cannot effectively generate the optimal path. At this time, the algorithm mainly adopts the method of randomly selecting actions, and attacks the hosts in the target network through the vulnerability attack module connected by the python script, while recording its reinforcement learning rewards, state transfer and other information, and updating the reinforcement learning strategy. When the target state is reached and the penetration test is completed, the platform will initialize the environment information and start the next round of penetration.
[0077] Step 4-2: In order to improve the penetration efficiency of the algorithm, the penetration problem will be simplified during the reinforcement learning training process. In a domain environment, penetrating a host has two main purposes, namely, intranet penetration or stealing unknown domain account credentials. The former helps to break through the restrictions of firewalls and routing, and the latter is used for lateral movement between domains. Hosts that cannot provide the above two helps are difficult to advance the penetration process. In other words, the actual benefits of penetrating any host are the same as the benefits of penetrating all hosts. Therefore, this type of host is considered to meet the state simplification conditions, and all its related states and vulnerability exploitation actions will be deleted.
[0078] Step 4-3: When the algorithm reaches a certain number of rounds, the reinforcement learning algorithm will select actions. Similarly, the rewards, state transfers and other information fed back by the environment after the attack will be used for reinforcement learning training. Finally, after a certain number of iterative cycles, the algorithm converges, and the platform will output the optimal path that has been verified and optimized multiple times.
[0079] In order to maintain the security of Windows domain network, predict the attacks that intruders may launch, and provide security protection strategies for domain administrators, this paper invented a security assessment method that applies reinforcement learning to Windows domain environment. First, based on the existing penetration testing tools and frameworks, the tool modules are enriched to improve its ability to discover and exploit current popular vulnerabilities. At the same time, reinforcement learning modeling functions are added, and the tool is automatically called by scripting language to build a penetration testing platform. Secondly, the vulnerability information of the target network and the exploitation module of the vulnerability tool are reinforced by reinforcement learning modeling, and a reward function is constructed to evaluate the actual value of the vulnerability exploitation action; finally, the state simplification method in the definition domain penetration is used to improve the efficiency of path solving, and the optimal attack path is verified through the continuous interaction between reinforcement learning and the experimental environment.
[0080] The above description is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiment. Any equivalent modifications or changes made by ordinary technicians in this field based on the contents disclosed by the present invention should be included in the protection scope recorded in the claims.
Claims
1. An automated Windows domain penetration method based on reinforcement learning, characterized in that: the method comprises the following steps: Step 1: Based on information collection tools, vulnerability scanning tools and vulnerability frameworks, supplemented by the extension modules of the corresponding tools, Python language is used as the connection script between tools and tools, and between tools and environments to build a penetration testing platform; Step 2: Use the penetration testing platform to collect information and scan vulnerabilities in the target environment, and match the vulnerability attack module used by each host based on the results; Step 3: Use the penetration testing platform to automatically perform reinforcement learning modeling; Step 3 is broken down into the following sub-steps: Step 3-1, the reinforcement learning is expressed in the form of MDP quadruple (S, A, R, P). Through continuous exploration and trial and error in the environment, positive and negative rewards are obtained, and the priority of the action is determined based on the reward. The different action selections a under different states s of reinforcement learning are called strategies. The learning goal is to optimize the strategy π to maximize the cumulative reward. The reinforcement learning strategy is solved using the Q learning algorithm. The algorithm uses the action and value function Q(s,a) to represent the strategy, and selects actions based on the expected size of the function of Q(s,a); Step 3-2: Based on the description of reinforcement learning in step 3-1, construct a four-tuple (S, A, R, P), and use the Q-learning algorithm to solve the optimal strategy for the target environment; Step 3-2 is broken down into the following sub-steps: Step 3-2-1, construct state space S: use the host information and the current permissions of the target host to define the state space. According to the information collected in step 2-1, the state of each host is expressed in the form of "IP-operating system-port: service-existing vulnerability-acquired permissions", where: Only the permissions change with the penetration process. The state space S is enumerated in the form of "state number-acquired permissions". In addition, the state of the successfully penetrated domain controller is set to the target state G. If the host n is a domain controller, the state space is enumerated as a set of "1:x-2:x-3:x-4:x-5:x......n-1:xn:y", where x is N, L, U; y is Null or G, where N, L, and U represent the host's three states of no permissions, local administrator permissions, and domain ordinary user permissions, respectively. G represents the current final state, and Null represents the final state not reached. Step 3-2-2, construct action space A: The action space represents the vulnerability exploitation actions used by the attacker, which comes from the penetration testing platform; the vulnerability exploitation module is subdivided into credential exploitation module and general vulnerability module; during automated penetration, in addition to exploiting traditional vulnerabilities for penetration, lateral movement is also performed through domain user credentials; Step 3-2-3, construct the reward function R: Reinforcement learning relies on the reward after interacting with the environment to make decisions; after performing any action, the agent obtains the immediate reward of the action according to the reward function and predicts the expected cumulative reward; the reward comes from the positive authority increase ΔR and the negative vulnerability loss r cost ; If each permission acquisition action generates permission gain R PA , according to the different permission status of each host, its score r is: The target host permission is not obtained: r=0; Obtain the normal user permissions of the target host: r=8; Obtain local administrator privileges on the target host: r=10; Obtain domain administrator privileges: r = 100 + r extra ; If the permission scores of a host before and after the vulnerability exploitation action are r1 and r2, the positive reward of the action is its permission gain: ΔR = r2-r1, r2>r1. Since the domain administrator's permission is equivalent to the highest permission of all machines in the domain, the extra score compensation r is obtained when the target state is reached. extra , whose value is equal to the sum of the scores obtained by upgrading all hosts in the domain environment to local administrator privileges; r cost By Vulnerability Level V rank 、Vulnerability Availability exploitble , the time when the vulnerability is exposed T vu ln is composed of three factors, whose values come from the basic score of the Common Vulnerability Scoring System CVSS, the usability score, and the length of time the vulnerability was released. cost It is expressed by the following formula: r cost =10sigmoid(T vuln )-V exploitble V rank / 10 where r cost The minimum value is 0; The reward for the permission acquisition action is the permission gain ΔR and the vulnerability consumption r cost Difference: R PA =ΔR-r cost Step 3-2-4, construct state transition P: Based on the actual penetration testing tool interacting with the real environment, explore the state transition through the interaction of the experimental environment; when reinforcement learning selects an action and executes it through the penetration testing platform, the platform will receive state change information fed back by the environment and explore the state transition accordingly; Step 4: Use the penetration testing platform to automatically call the attack module to attack, and perform reinforcement learning training and state simplification based on the returned results to continuously optimize the optimal attack path.
2. According to claim 1, an automated Windows domain penetration method based on reinforcement learning is characterized in that: In step 1, the following sub-steps are broken down: Step 1-1: The information collection tool is based on the existing tool nmap, the vulnerability scanning tool uses the commercial tool AWVS, and the vulnerability framework is based on the open source penetration testing tool set metasploit-framework (MSF); Step 1-2, expand MSF through vulnerability exploit code, standardize the codes in different programming languages, and rewrite them in ruby; Steps 1-3, use Python scripts to connect the reinforcement learning module with information collection, vulnerability scanning, and vulnerability attacks. When reinforcement learning decides to take an action, directly call the corresponding tool's related functions to implement it; Steps 1-4 complete the establishment of the entire penetration testing platform, which has three tools: information collection, vulnerability scanning, and vulnerability exploitation. Reinforcement learning is connected to the above three types of tools through the Python interface. When reinforcement learning completes training, automated penetration testing is achieved.
3. According to claim 1, an automated Windows domain penetration method based on reinforcement learning is characterized in that: In step 2, the following sub-steps are broken down: Step 2-1, use the real domain network as the experimental environment, access the penetration test platform built in step 1, and give the permission of a host in the target domain network. Through this host as a springboard, the platform will first use the tool nmap and its extended functions, and automatically call the python script to collect information and scan vulnerabilities of all hosts and domain users in the target domain network, and obtain the IP, operating system, port: service, and vulnerability information of each host; Step 2-2, for the vulnerability scanning results of each host in step 2-1, use the corresponding vulnerability exploitation conditions to match according to the host's operating system, port: service, and determine whether the scanning tool has a false positive. Only the vulnerabilities that meet the vulnerability exploitation conditions are selected as real attack actions by reinforcement learning.
4. According to claim 1, an automated Windows domain penetration method based on reinforcement learning is characterized in that: In step 4, the following sub-steps are broken down: Step 4-1, the reinforcement learning model constructed in steps 1 to 3 is subjected to reinforcement learning training using the penetration testing platform, that is, the strategy for this environment is optimized until the optimal strategy is found; at the beginning of automated penetration, the reinforcement learning algorithm cannot effectively generate the optimal path. At this time, the action is randomly selected, and the vulnerability attack module connected by the python script is used to attack the host in the target network, while recording its reinforcement learning reward and state transfer information, and updating the reinforcement learning strategy; When the target state is reached and the penetration test is completed, the platform will initialize the environment information and start the next round of penetration; Step 4-2 will simplify the penetration problem. In a domain environment, the purpose of penetrating a host is to penetrate the intranet or steal unknown domain account credentials. The actual benefit of penetrating any host is the same as the benefit of penetrating all hosts. Therefore, it is considered that this type of host meets the state simplification condition, and all its related states and vulnerability exploitation actions will be deleted. Step 4-3, after a certain number of rounds, the reinforcement learning algorithm will select actions. Similarly, the rewards and state transition information of the environment feedback after the attack will be used for reinforcement learning training. Finally, after a certain number of iterative cycles, the algorithm converges, and the platform outputs the optimal path that has been verified and optimized multiple times.
Citation Information
Patent Citations
Automatic penetration testing method based on deep reinforcement learning
CN113660241A
Network honeypot deployment method for penetration attacks
CN113783881A