Windows domain information scanning and penetration test path planning method based on reinforcement learning
By abstracting the Windows domain penetration testing process into a Markov decision-making process and training reinforcement learning agents, the problem of inefficient information scanning and penetration testing path planning in the Windows domain in the prior art is solved, and efficient and dynamic penetration testing path planning and security assessment are achieved.
Patent Information
- Application Number
- CN202510157447.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The prior art has problems such as low path planning efficiency, incomplete attack surface coverage, and weak environmental adaptability in information scanning and penetration testing path planning within the Windows domain.
Using a reinforcement learning-based method, the Windows domain penetration testing process is abstracted into a Markov decision-making process, information is collected through the LDAP protocol, state space, action space and reward functions are defined, and reinforcement learning agents are trained to generate the optimal penetration testing path.
It realizes the acquisition of Windows domain environment information with minimal movement, dynamically adapting to complex network environments, improving the efficiency and coverage of penetration testing, discovering weak links in the domain environment and evaluating security status.
Smart Images

Figure CN120012112A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning and information security, and specifically is a method for information scanning and penetration testing path planning within a Windows domain based on reinforcement learning. Background Art
[0002] With the continuous improvement of enterprise informatization, Windows domain environment has become the mainstream choice for enterprise network architecture due to its convenient centralized management characteristics. However, the complex trust relationship within the domain, multi-level permission structure and diversified service configuration make the Windows domain environment face severe security challenges. As the core means of actively discovering security vulnerabilities, the efficiency and intelligence of penetration testing directly affect the construction effect of the security protection system.
[0003] Traditional Windows domain information scanning tools often need to perform specific attack actions to obtain information, which is very noisy and can be easily discovered by firewalls when facing high-version operating systems. Traditional penetration testing methods mainly rely on manual detection and experience judgment by security experts, which have problems such as long testing cycles, high labor costs, and difficulty in ensuring test coverage. Although existing automated tools such as Metasploit and Cobalt Strike can achieve partial automated testing through pre-set attack modules, they essentially still adopt linear scanning strategies based on rule bases and lack the ability to dynamically adapt to complex network environments. Specifically, they are: (1) prone to local optimality when encountering multi-hop attack paths and unable to effectively identify cross-node combined attack vectors; (2) lack of autonomous decision-making capabilities when facing dynamically changing network topologies and real-time updated vulnerability information; (3) lack of in-depth correlation analysis of key elements such as group policy objects and active directory hierarchical relationships that are unique to Windows domains.
[0004] In recent years, some studies have attempted to apply machine learning technology to the field of penetration testing, but there are still significant limitations in practical applications: supervised learning methods rely on a large amount of labeled data for model training, while the acquisition of real attack scenario data is subject to legal and ethical restrictions; traditional reinforcement learning algorithms based on Q-learning face the dimensionality curse problem when dealing with high-dimensional state spaces, especially in scenarios involving hundreds of domain nodes, where the algorithm convergence rate drops sharply. In addition, existing methods generally lack targeted design for Windows domain environments, specifically: (1) they do not fully consider the state representation of Windows-specific attack surfaces such as Kerberos authentication and NTLM hash transfer; (2) they do not adequately model the authority transfer relationship between domain controllers, member servers, and client workstations; and (3) they lack a dynamic weight evaluation mechanism for key risk points such as group policy preferences and lateral movement paths.
[0005] Currently, there is an urgent need for a penetration testing system that can deeply integrate the characteristics of the Windows domain environment and realize intelligent path planning through autonomous reinforcement learning, so as to solve key problems existing in existing technologies such as low path planning efficiency, incomplete attack surface coverage, and weak environmental adaptability. Summary of the invention
[0006] In order to solve the above problems, the purpose of the present invention is to provide a Windows domain information scanning and penetration test path planning method based on reinforcement learning, which obtains Windows domain environment information with minimal movement, abstracts the domain penetration test into a Markov decision process, trains reinforcement learning agents, and constructs a domain penetration test model. By achieving the penetration goal and satisfying the maximization of the cumulative reward, the optimal penetration test path is found, thereby finding the weak links of the domain environment and evaluating the security status of the domain environment.
[0007] In order to achieve the above object, the present invention is achieved through the following technical solutions:
[0008] A Windows domain information scanning and penetration test path planning method based on reinforcement learning includes the following steps:
[0009] Step 1: Detect and scan the target Windows domain, establish a connection with the target host using the LDAP protocol, query the target domain environment information by setting filtering conditions, detect vulnerabilities related to ADCS, delegation, domain controllers, and Exchange servers, and scan and collect information including host information within the domain, login user information, related service configuration information, survival IP information, and group policy information;
[0010] Step 2: Based on the relevant information in the Windows domain collected in step 1, the collected environment information and potential vulnerability information are classified, sorted and stored. Combined with the detected environment information and potential vulnerability information, valuable information is retained and worthless information is filtered out.
[0011] Step 3: Based on step 2, the Windows domain penetration test process is abstracted into a Markov decision process, and the state space, action space, and reward function are defined. The state space is a combination of the states of the hosts within the domain and the entire Windows domain. The states of the hosts within the domain include login user permissions, host operating system, current control type, and special service type, etc. The states of the domain include whether it is a root domain, current control type, whether there are unconstrained delegation users, whether there are dangerous group policies, whether there are users who do not require Kerberos pre-authentication, and whether there are users with SPNs, etc. The action space includes a series of attack actions against the host, including privilege escalation, vulnerability exploitation, and group policy attacks, etc. The reward function is to take actions based on the current policy, and obtain the immediate reward for the current action according to the reward function after executing any action;
[0012] Step 4: Based on step 3, build a reinforcement learning agent, and generate a penetration test path by interactive training with the simulation environment. Adjust it according to the reward strategy, find the optimal penetration test path by accumulating rewards, build a Windows domain penetration test model, and finally generate the optimal penetration test path.
[0013] Furthermore, in step 1, an LDAP connection is established through the LDAP protocol to detect the target domain environment and collect information within the domain, specifically:
[0014] On the premise of obtaining the credentials of any host in the target Windows domain, in order to reduce the noise of the scanning action and avoid being detected by the target host firewall, an LDAP connection is established through the LDAP protocol, and filters are set to collect relevant information and exploitable vulnerability information in the domain. The collected relevant information in the domain includes the domain name, the name of the common host in the domain, the name of the domain controller, the name of the server in the domain, the domain user information, the user group information, the domain trust relationship information, etc. The detected vulnerability information in the domain includes the vulnerabilities related to ADCS, domain controller, delegation, and Exchange server.
[0015] Furthermore, in step 2, based on step 1, the collected environmental information and potential vulnerability information are classified, sorted and stored, valuable information is retained, and worthless information is screened out, specifically:
[0016] The obtained relevant information is processed, combined with the detected possible exploitable vulnerability information and the obtained domain host and other related information, to analyze whether the potential vulnerability has exploitation conditions, the vulnerability exploitation success rate, etc., and to analyze whether the user information, service information and other conditions required for vulnerability exploitation exist. If so, the corresponding exploitable vulnerability is stored in the dictionary, and if not, the vulnerability information is deleted, thereby retaining valuable vulnerability information and domain host and other related information. Finally, a yaml file is generated based on the processed information for subsequent intelligent agent training.
[0017] Furthermore, in step 3, based on step 2, the Windows domain penetration test process is abstracted into a Markov decision process, and the state space, action space and reward function are defined, specifically:
[0018] Step 3-1: Abstract the Windows domain penetration test process into a Markov decision process. The Markov decision process is a framework for solving multi-stage sequential decision problems. It is a mathematical model based on state transition probability and reward, which is used to describe the situation where an intelligent agent takes a series of decisions in a random environment. The Markov decision process consists of four tuples, which represent the action space, state space, reward function and state transition probability function.
[0019] Step 3-2: Define the action space. Each action in the action space consists of a triple (type, target, action). Type 0 means that the action target object is a host in the domain, and type 1 means that the action target object is a Windows domain. Target represents the serial number of the target host or Windows domain. Action is the action serial number corresponding to each attack action. The attack actions against the host include privilege escalation, unconstrained delegation, EternalBlue, MS14-068, MS08-067, Exchange Relay, Exchange Proxy, MSSQL CLR, Hash Dump, and DCSYNC. The attack actions against the domain include CVE-2020-1472, CVE-2021-42287, Group Policy Attack, AS_REP Roasting, Kerberoasting, and SIDHistory.
[0020] Step 3-3: Define the state space. The state space is composed of the state space describing the hosts in the domain and the state space of the Windows domain.
[0021] The state space describing the host is composed of a two-dimensional array, which represents the host object in the domain as the test target and a corresponding feature describing the host. Features include identity, current control type, whether to obtain operating system information, special service type, whether the current logged-in user includes a member of the local administrator group, whether there is a high-privilege process in the domain, and whether the high-privilege credentials in the domain are obtained. The value of identity is 0, which means that the host is an ordinary host in the domain, 1 means that the host is an ordinary server in the domain, 2 means that the host is an ordinary controller in the domain, and 3 means that the host is a domain controller in the root domain; the value of the current control type is 0, which means that the control of the host has not been obtained, 1 means that the ordinary domain user rights of the host are obtained, and 2 means that the SYSTEM control rights of the host are obtained; the value of the special service type is 0, which means that the host does not have special services, 1 means that the host has MSSQL database services, and 2 means that the host has Exchange mail services.
[0022] The state space describing the domain is also composed of a two-dimensional array, which represents the domain object as the test target and a corresponding feature describing the domain. Features include whether it is a root domain, the current control type, whether any high-privilege credentials of the domain have been obtained, whether there are unconstrained delegation accounts, whether there are dangerous group policy files, whether there are users who do not need Kerberos pre-authentication, and whether there are users with SPNs. The state space describing the domain uses 0 / 1 encoding to indicate whether the value of the item is found.
[0023] Step 3-4: Define the reward function. The reinforcement learning agent observes the state in the environment and takes actions based on the current strategy. After performing any action, it will obtain the immediate reward for the current action according to the reward function. The goal of the agent is to maximize the future reward accumulation, so the setting of the reward function can be used to guide the agent to learn the optimal strategy. The reward is divided into several parts, including the reward setting for changes in host control, the reward setting for changes in domain control, and the reward setting for credential collection.
[0024] Furthermore, in step 4, based on step 3, a reinforcement learning agent is constructed to find the optimal penetration test path, specifically:
[0025] Based on the already defined state space, action space and reward function, the deep Q network algorithm is used to define and initialize the Q network, and the environment state is encoded as a numerical vector, including the current host information, network topology, domain control status and other information. The neural network structure and its parameters are set, including the input layer, hidden layer and output layer of the core architecture, the action mask is set, and the Q value of illegal actions is set to negative infinity to filter invalid actions. A dual network mechanism is used to update parameters in real time in the training network for selecting actions and calculating the current Q value. In the target network, parameters are synchronized from the training network regularly to calculate the target Q value and stabilize the training process. And the synchronization frequency is set. During training, actions are selected according to the greedy algorithm, and actions with the largest Q value are selected during deployment.
[0026] During the training process, we first build a simulated Windows domain environment and initialize the training network and target network. Then we set hyperparameters, including discount factor, experience replay buffer capacity, batch size, initial exploration rate, decay rate, learning rate, and other parameters. In a single training cycle, we interact with the environment, select actions to execute and store experience, update the network, calculate the target Q value, and finally terminate the training according to the set termination conditions.
[0027] After sufficient training, the agent can select vulnerability exploitation actions based on the current state and action space, obtain more environmental information and permissions by successfully exploiting the vulnerability, and further improve the state space information. After selecting the vulnerability exploitation action, the agent calculates and accepts the reward according to the reward function. The agent updates the Q value of each action in the current state based on the feedback reward value and optimizes its behavior strategy. In the process of continuous iteration of the Q value, the agent's strategy is continuously optimized and gradually approaches the optimal penetration test path.
[0028] Through the above steps, the intelligent agent can eventually generate a penetration test path that meets the penetration test objectives and maximizes the cumulative reward, realize a comprehensive penetration test assessment of the target Windows domain environment with minimal movement, and ultimately achieve the purpose of evaluating and maintaining the security of the Windows domain environment.
[0029] The present invention obtains the environmental information in the Windows domain with minimal movement, abstracts the domain penetration test into a Markov decision process, trains reinforcement learning agents, constructs a domain penetration test model, and finds the optimal penetration test path by achieving the penetration goal and satisfying the maximization of the cumulative reward, thereby finding the weak links of the domain environment and evaluating the security status of the domain environment.
[0030] Compared with the prior art, the present invention has the following advantages:
[0031] (1) Use the LDAP protocol to establish an LDAP connection with the target host, and collect information within the domain by setting filters. The current common method of collecting information in a domain penetration test mainly relies on gaining control of a host within the domain and confirming the surviving host within the domain by communicating with other hosts within the domain. This method makes a lot of noise within the domain and is easily discovered by defense measures. The method using the LDAP protocol only requires the communication credentials of a host within the domain without requiring full control. In this way, information within the domain can be obtained with minimal noise. The collected information is then used to train the agent. The entire process can be carried out without being discovered by defense measures, and has good concealment.
[0032] (2) Compared with the traditional static scanning method based on rule base, this dynamic scanning method realized by deep reinforcement learning framework can collect information more timely to obtain key information of the domain environment, reducing the efficiency loss caused by manual method. Deep reinforcement Q network can fully consider various complex domain environments and penetration test scenarios, and can be compatible with new Windows network forms such as Azure AD hybrid architecture and multi-domain trust forest, and respond to changes in network topology in a timely manner, greatly improving efficiency while ensuring wide applicability.
[0033] (3) When defining the state space, this method fully considers the modeling method for Windows domain characteristics. Starting from the two dimensions of host state and domain state, it encodes elements such as Kerberos tickets, GPO policies, and AD topology into reinforcement learning state vectors, which fully reflects the status of the entire domain. This solves the problem of insufficient extraction of domain environment features in existing technologies, and greatly improves the penetration success rate in a real domain environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 The present invention provides a flow chart from scanning the domain environment to training the intelligent agent to building a domain penetration test model and then to penetration path planning.
[0035] Figure 2 This is a flow chart of collecting target domain information in the present invention.
[0036] Figure 3 This is a structural diagram for information collection, analysis and screening in the present invention.
[0037] Figure 4 A structural diagram of the parameters required to construct the intelligent agent in the present invention.
[0038] Figure 5 A flow chart for planning the penetration path of the intelligent agent in the present invention. DETAILED DESCRIPTION
[0039] The present invention is described below in conjunction with the accompanying drawings. For the purpose of clear description, many practical details will be described together in the following description. However, it should be understood that these practical details should not be used to limit the present invention. That is, in some embodiments of the present invention, these practical details are not necessary.
[0040] A Windows domain information scanning and penetration test path planning method based on reinforcement learning is designed to collect information and plan penetration paths for the target Windows domain with minimal noise. The implementation system of this method includes three modules, namely, the domain environment detection module, the domain information processing module, and the penetration path planning module. The main steps include: performing information detection on the target Windows domain environment, mapping its network topology, establishing a connection with the target machine by using protocols such as LDAP to reduce the noise during information scanning and avoid being discovered by the firewall of the target host, collecting relevant information in the domain, including host information in the domain, relevant service configuration information in the domain, surviving IP information in the domain, and information on possible exploitable vulnerabilities in the domain; collecting, organizing and abstracting the detected relevant information, using the Markov decision process to solve the multi-stage sequential decision problem, abstracting the Windows domain penetration test process into a Markov decision process, and realizing environmental simulation, and defining the state space, action space, reward function and state transition probability required for training the deep reinforcement learning DQN network according to the abstracted information. A domain penetration test model based on reinforcement learning is designed and trained; according to the domain information collected for the target Windows domain environment, the model evaluates the behavior through interaction with the simulation environment, and adjusts the strategy according to the reward to gradually reach the optimal strategy, and finally achieves the goal of completing the penetration test and gives the optimal penetration test path. The present invention targets the Windows domain environment and its changes, adopts a deep reinforcement learning Q network, models the Windows domain penetration test scenario as a Markov decision process, collects and analyzes the domain information, constructs a Windows domain penetration test model based on reinforcement learning, realizes the collection of Windows domain information and the planning of the penetration test path, and finally achieves the purpose of evaluating and maintaining the security of the Windows domain environment.
[0041] Specifically, a Windows domain information scanning and penetration test path planning method based on reinforcement learning, such as Figure 1 As shown, the method comprises the following steps:
[0042] Step 1: Detect and scan the target Windows domain, establish a connection with the target host using the LDAP protocol, query the target domain environment information by setting filtering conditions, detect vulnerabilities related to ADCS, delegation, domain controllers, and Exchange servers, and scan and collect information including host information within the domain, login user information, related service configuration information, survival IP information, and group policy information;
[0043] Step 2: Based on the relevant information in the Windows domain collected in step 1, the collected environment information and potential vulnerability information are classified, sorted and stored. Combined with the detected environment information and potential vulnerability information, valuable information is retained and worthless information is filtered out.
[0044] Step 3: Based on step 2, the Windows domain penetration test process is abstracted into a Markov decision process, and the state space, action space, and reward function are defined. The state space is a combination of the states of the hosts within the domain and the entire Windows domain. The states of the hosts within the domain include login user permissions, host operating system, current control type, and special service type, etc. The states of the domain include whether it is a root domain, current control type, whether there are unconstrained delegation users, whether there are dangerous group policies, whether there are users who do not require Kerberos pre-authentication, and whether there are users with SPNs, etc. The action space includes a series of attack actions against the host, including privilege escalation, vulnerability exploitation, and group policy attacks, etc. The reward function is to take actions based on the current policy, and obtain the immediate reward for the current action according to the reward function after executing any action;
[0045] Step 4: Based on step 3, build a reinforcement learning agent, and generate a penetration test path by interactive training with the simulation environment. Adjust it according to the reward strategy, find the optimal penetration test path by accumulating rewards, build a Windows domain penetration test model, and finally generate the optimal penetration test path.
[0046] like Figure 2 As shown, step 1 is specifically as follows:
[0047] On the premise of obtaining the credentials of any host in the target Windows domain, in order to reduce the noise of the scanning action and avoid being detected by the target host firewall, an LDAP connection is established through the LDAP protocol, and filters are set to collect relevant information and exploitable vulnerability information in the domain. The collected relevant information in the domain includes the domain name, the name of the common host in the domain, the name of the domain controller, the name of the server in the domain, the domain user information, the user group information, the domain trust relationship information, etc. The detected vulnerability information in the domain includes the vulnerabilities related to ADCS, domain controller, delegation, and Exchange server.
[0048] like Figure 3 As shown, step 2 is specifically as follows:
[0049] The obtained relevant information is processed, combined with the detected possible exploitable vulnerability information and the obtained domain host and other related information, to analyze whether the potential vulnerability has exploitation conditions, the vulnerability exploitation success rate, etc., and to analyze whether the user information, service information and other conditions required for vulnerability exploitation exist. If so, the corresponding exploitable vulnerability is stored in the dictionary, and if not, the vulnerability information is deleted, thereby retaining valuable vulnerability information and domain host and other related information. Finally, a yaml file is generated based on the processed information for subsequent intelligent agent training.
[0050] like Figure 4 As shown, step 3 is specifically as follows:
[0051] Step 3-1: Abstract the Windows domain penetration test process into a Markov decision process. The Markov decision process is a framework for solving multi-stage sequential decision problems. It is a mathematical model based on state transition probability and reward, which is used to describe the situation where an agent takes a series of decisions in a random environment. The Markov decision process consists of four tuples, which represent the action space, state space, reward function and state transition probability function.
[0052] Step 3-2: Define the action space. Each action in the action space consists of a triple (type, target, action). Type 0 means that the action target object is a host in the domain, and type 1 means that the action target object is a Windows domain. Target represents the serial number of the target host or Windows domain. Action is the action serial number corresponding to each attack action. The attack actions against the host include privilege escalation, unconstrained delegation, EternalBlue, MS14-068, MS08-067, Exchange Relay, Exchange Proxy, MSSQL CLR, Hash Dump, and DCSYNC. The attack actions against the domain include CVE-2020-1472, CVE-2021-42287, Group Policy Attack, AS_REP Roasting, Kerberoasting, and SIDHistory.
[0053] Step 3-3: Define the state space. The state space is composed of the state space describing the hosts in the domain and the state space of the Windows domain.
[0054] The state space describing the host is composed of a two-dimensional array, which represents the host object in the domain as the test target and a corresponding feature describing the host. Features include identity, current control type, whether to obtain operating system information, special service type, whether the current logged-in user includes a member of the local administrator group, whether there is a high-privilege process in the domain, and whether the high-privilege credentials in the domain are obtained. The value of identity is 0, which means that the host is an ordinary host in the domain, 1 means that the host is an ordinary server in the domain, 2 means that the host is an ordinary controller in the domain, and 3 means that the host is a domain controller in the root domain; the value of the current control type is 0, which means that the control of the host has not been obtained, 1 means that the ordinary domain user rights of the host are obtained, and 2 means that the SYSTEM control rights of the host are obtained; the value of the special service type is 0, which means that the host does not have special services, 1 means that the host has MSSQL database services, and 2 means that the host has Exchange mail services.
[0055] The state space describing the domain is also composed of a two-dimensional array, which represents the domain object as the test target and a corresponding feature describing the domain. Features include whether it is a root domain, the current control type, whether any high-privilege credentials of the domain have been obtained, whether there are unconstrained delegation accounts, whether there are dangerous group policy files, whether there are users who do not need Kerberos pre-authentication, and whether there are users with SPNs. The state space describing the domain uses 0 / 1 encoding to indicate whether the value of the item is found.
[0056] Step 3-4: Define the reward function. The reinforcement learning agent observes the state in the environment and takes actions based on the current strategy. After performing any action, it will obtain the immediate reward for the current action according to the reward function. The goal of the agent is to maximize the future reward accumulation, so the setting of the reward function can be used to guide the agent to learn the optimal strategy. The reward is divided into several parts, including the reward setting for changes in host control, the reward setting for changes in domain control, and the reward setting for credential collection.
[0057] like Figure 5 As shown, step 4 is specifically as follows:
[0058] Based on the already defined state space, action space and reward function, the deep Q network algorithm is used to define and initialize the Q network, and the environment state is encoded as a numerical vector, including the current host information, network topology, domain control status and other information. The neural network structure and its parameters are set, including the input layer, hidden layer and output layer of the core architecture, the action mask is set, and the Q value of illegal actions is set to negative infinity to filter invalid actions. A dual network mechanism is used to update parameters in real time in the training network for selecting actions and calculating the current Q value. In the target network, parameters are synchronized from the training network regularly to calculate the target Q value and stabilize the training process. And the synchronization frequency is set. During training, actions are selected according to the greedy algorithm, and actions with the largest Q value are selected during deployment.
[0059] During the training process, we first build a simulated Windows domain environment and initialize the training network and target network. Then we set hyperparameters, including discount factor, experience replay buffer capacity, batch size, initial exploration rate, decay rate, learning rate, and other parameters. In a single training cycle, we interact with the environment, select actions to execute and store experience, update the network, calculate the target Q value, and finally terminate the training according to the set termination conditions.
[0060] After sufficient training, the agent can select vulnerability exploitation actions based on the current state and action space, obtain more environmental information and permissions by successfully exploiting the vulnerability, and further improve the state space information. After selecting the vulnerability exploitation action, the agent calculates and accepts the reward according to the reward function. The agent updates the Q value of each action in the current state based on the feedback reward value and optimizes its behavior strategy. In the process of continuous iteration of the Q value, the agent's strategy is continuously optimized and gradually approaches the optimal penetration test path.
[0061] Through the above steps, the intelligent agent can eventually generate a penetration test path that meets the penetration test objectives and maximizes the cumulative reward, realize a comprehensive penetration test assessment of the target Windows domain environment with minimal movement, and ultimately achieve the purpose of evaluating and maintaining the security of the Windows domain environment.
[0062] The above description is only an embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of the claims of the present invention.
Claims
1. A Windows domain information scanning and penetration test path planning method based on reinforcement learning, characterized in that: The steps include: Step 1: Detect and scan the target Windows domain, establish a connection with the target host using the LDAP protocol, query the target domain environment information by setting filtering conditions, detect vulnerability information in the domain, and scan and collect information including host information in the domain, login user information, related service configuration information, survival IP information, and group policy information; Step 2: Based on the collected relevant information in the Windows domain, the collected environmental information and potential vulnerability information are classified, sorted and stored. Combined with the detected environmental information and potential vulnerability information, valuable information is retained and worthless information is filtered out. Step 3: Abstract the Windows domain penetration test process into a Markov decision process, and define the state space, action space, and reward function; the state space is a combination of the states of the hosts in the domain and the entire Windows domain. The states of the hosts in the domain include the login user permissions, host operating system, current control type, and special service type. The states of the domain include whether it is a root domain, the current control type, whether there are unconstrained delegation users, whether there are dangerous group policies, whether there are users who do not need Kerberos pre-authentication, and whether there are users with SPNs; The action space includes a range of attack actions against the host, including privilege escalation, vulnerability exploitation, and group policy attacks; The reward function is to take action based on the current strategy, and obtain the immediate reward of the current action according to the reward function after executing any action; Step 4: Build a reinforcement learning agent and generate a penetration test path by interactive training with the simulation environment. Adjust it according to the reward strategy, find the optimal penetration test path by accumulating rewards, build a Windows domain penetration test model, and finally generate the optimal penetration test path.
2. According to claim 1, a Windows domain information scanning and penetration test path planning method based on reinforcement learning is characterized in that: The step 1 is specifically as follows: After obtaining the credentials of any host in the target Windows domain, establish an LDAP connection through the LDAP protocol, and set filters to collect relevant information and exploitable vulnerability information in the domain; The collected domain-related information includes domain name, common host names in the domain, domain controller name, server names in the domain, domain user information, user group information, and domain trust relationship information. The detected domain vulnerability information includes vulnerabilities related to ADCS, domain controllers, delegation, and Exchange servers.
3. According to claim 1, a Windows domain information scanning and penetration test path planning method based on reinforcement learning is characterized in that: The step 2 is specifically as follows: The obtained relevant information is processed, and combined with the detected possible exploitable vulnerability information and the obtained information about the hosts in the domain, it is analyzed whether the potential vulnerability has the conditions for exploitation, the success rate of vulnerability exploitation, and whether the user information and service information conditions required for vulnerability exploitation exist. If so, the corresponding exploitable vulnerability is stored in the dictionary. If not, the vulnerability information is deleted, thereby retaining valuable vulnerability information and information about the hosts in the domain. Finally, a yaml file is generated based on the processed information for subsequent intelligent agent training.
4. According to claim 1, a Windows domain information scanning and penetration test path planning method based on reinforcement learning is characterized in that: Step 3 is as follows: Step 3-1: Abstract the Windows domain penetration test process into a Markov decision process. The Markov decision process is a framework for solving multi-stage sequential decision problems. It is a mathematical model based on state transition probability and reward, which is used to describe the situation in which an agent takes a series of decisions in a random environment. The Markov decision process consists of four tuples, representing the action space, state space, reward function and state transition probability function respectively; Step 3-2: Define the action space; each action in the action space consists of a triplet of type, target, and action. Type 0 means that the target object of the action is a host in the domain, and type 1 means that the target object of the action is a Windows domain. Target represents the serial number of the target host or Windows domain. Action is the action serial number corresponding to each attack action. The attack actions against the host include privilege escalation, unconstrained delegation, EternalBlue, MS14-068, MS08-067, Exchange Relay, Exchange Proxy, MSSQL CLR, Hash Dump, and DCSYNC. The attack actions against the domain include CVE-2020-1472, CVE-2021-42287, Group Policy Attack, AS_REP Roasting, Kerberoasting, and SIDHistory. Step 3-3: Define the state space; the state space is composed of the state space describing the hosts in the domain and the state space of the Windows domain; The state space describing the host is composed of a two-dimensional array, which respectively represents the host object in the domain as the test target and a corresponding feature describing the host. The features include identity, current control type, whether the operating system information is obtained, special service type, whether the current logged-in user includes a member of the local administrator group, whether there is a high-privilege process in the domain, and whether the high-privilege credentials in the domain are obtained; the identity value of 0 means that the host is an ordinary host in the domain, 1 means that the host is an ordinary server in the domain, 2 means that the host is an ordinary controller in the domain, and 3 means that the host is a domain controller in the root domain; the current control type value of 0 means that the control right of the host has not been obtained, 1 means that the ordinary domain user rights of the host are obtained, and 2 means that the SYSTEM control right of the host is obtained; the special service type value of 0 means that the host does not have a special service, 1 means that the host has an MSSQL database service, and 2 means that the host has an Exchange mail service; The state space describing the domain consists of a two-dimensional array, which respectively represents the domain object as the test target and a corresponding feature describing the domain. The features include whether it is a root domain, the current control type, whether any high-privilege credentials of the domain have been obtained, whether there are unconstrained delegation accounts, whether there are dangerous group policy files, whether there are users who do not need Kerberos pre-authentication, and whether there are users with SPNs; the state space describing the domain uses 0 / 1 encoding to indicate whether the value of the item is found; Step 3-4: Define the reward function; The reinforcement learning agent observes the state in the environment and takes actions based on the current strategy. After performing any action, it obtains the immediate reward for the current action according to the reward function. The goal of the agent is to maximize the future reward accumulation. Therefore, the setting of the reward function is used to guide the agent to learn the optimal strategy. The reward is divided into several parts, including the reward setting for changes in host control, the reward setting for changes in domain control, and the reward setting for credential collection.
5. According to claim 1, a Windows domain information scanning and penetration test path planning method based on reinforcement learning is characterized in that: Step 4 is as follows: Based on the already defined state space, action space and reward function, a deep Q network algorithm is used to define and initialize the Q network, encode the environment state into a numerical vector, including the current host information, network topology, and domain control state information; the neural network structure and its parameters are set, including the input layer, hidden layer, and output layer of the core architecture, and the action mask is set. The Q value of illegal actions is set to negative infinity to filter invalid actions; a dual network mechanism is used to update parameters in the training network in real time for selecting actions and calculating the current Q value; parameters are regularly synchronized from the training network in the target network for calculating the target Q value, stabilizing the training process, and setting the synchronization frequency. Actions are selected according to the greedy algorithm during training, and actions with the largest Q value are selected during deployment; During the training process, firstly, a simulated Windows domain environment is built to initialize the training network and the target network. Secondly, hyperparameters are set, including discount factor, experience replay buffer capacity, batch size, initial exploration rate, decay rate, and learning rate. In a single training cycle, the network is interactively cycled with the environment, actions are selected to be executed and experiences are stored, and the network is updated to calculate the target Q value. Finally, the training is terminated according to the set termination conditions. After sufficient training, the agent selects the vulnerability exploitation action based on the current state and action space, and obtains more environmental information and permissions by successfully exploiting the vulnerability, further improving the state space information; after selecting the vulnerability exploitation action, the agent calculates and accepts the reward according to the reward function. The agent updates the Q value of each action in the current state according to the feedback reward value, and optimizes its behavior strategy; In the process of continuous iteration of Q value, the agent's strategy is constantly optimized and gradually approaches the optimal penetration test path. The agent finally generates a penetration test path that meets the penetration test objectives and maximizes the cumulative reward.
Citation Information
Patent Citations
Penetration testing method and device
CN106874768A
Automatic Windows domain penetration method based on reinforcement learning
CN114444086A
Method and device for defending penetration attack based on reinforcement learning, and electronic equipment
CN115473677A
Deep reinforcement learning intelligent penetration testing method and device based on imitation learning
CN115473706A
Intelligent penetration testing method and system based on deep reinforcement learning
CN117235742A
Cited By
Domain penetration testing method based on adaptive vulnerability utilization
CN121485996A
A domain penetration testing method based on adaptive exploit
CN121485996B
Automatic penetration testing method and system based on cognitive decision model
CN121615150A
An automated penetration testing method and system based on a cognitive decision model
CN121615150B
Power system network intelligent penetration test method based on deep reinforcement learning
CN121690739A