Windows domain penetration test system and method based on deep reinforcement learning
Through the Windows domain penetration testing system based on deep reinforcement learning, the problem that traditional penetration testing cannot cope with complex domain controller attacks is solved, and automated penetration testing and risk assessment are realized, which significantly improves testing efficiency and accuracy.
Patent Information
- Application Number
- CN202510143704.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional manual penetration testing and passive security protection measures cannot effectively deal with complex and diverse domain controller attacks, and it is difficult to meet modern network security needs.
Using a Windows domain penetration testing system based on deep reinforcement learning, automated penetration testing and risk assessment are realized through information collection modules, vulnerability detection modules, agent generation modules and agent training modules. The agent dynamically analyzes the domain environment, uses reward function optimization strategies, discovers security issues and generates test reports.
It realizes rapid learning of effective attack paths in complex and variable domain environments, covers a wider range of vulnerabilities, and generates more efficient penetration testing solutions, significantly improving the efficiency and accuracy of penetration testing.
Smart Images

Figure CN120050074A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning and automated penetration testing, and in particular to a Windows domain penetration testing system and method based on deep reinforcement learning. Background Art
[0002] In the field of modern network security, Windows domains have been widely used in various enterprise intranet systems for their efficient resource sharing and centralized management capabilities. However, this convenience also hides huge security risks. As the core of the Windows domain environment, the domain controller may cause the collapse of the security protection system of the entire network once it is controlled by an attacker. In recent years, with the continuous upgrading of attack techniques, the attack methods against domain controllers have become more complex and diverse. By mining and exploiting vulnerabilities and weaknesses in the domain, attackers gradually penetrate the target system, steal sensitive data, and even completely take over the network. Faced with such severe threats, traditional manual penetration testing and passive security protection measures can no longer meet the response needs. Their efficiency and coverage are far from enough to cope with increasingly complex attack methods and dynamically changing network environments.
[0003] In this context, a Windows domain penetration testing system based on deep reinforcement learning came into being. By introducing reinforcement learning technology, the system can demonstrate excellent performance and adaptability in automated penetration testing. Reinforcement learning is a machine learning method, the core of which is to allow the agent to gradually optimize the strategy based on trial and error experience through continuous interaction with the environment until a specific goal is achieved. In the scenario of domain penetration testing, the agent is designed to explore and identify potential attack paths within the domain. By dynamically analyzing the information structure and vulnerability distribution in the domain environment, the agent can adjust its strategy to maximize the future cumulative reward based on the immediate feedback of the reward function after each step of operation. This mechanism enables the agent to quickly learn effective attack paths, cover a wider range of vulnerability combinations, and generate more efficient penetration testing solutions in a limited time when facing complex and changing domain environments.
[0004] The penetration testing system driven by reinforcement learning is not only significantly superior to traditional methods in terms of efficiency and accuracy, but can also dynamically adapt to changes in the domain environment and promptly discover new security risks. Through comprehensive analysis of attack paths and threat assessment, the system can provide network security personnel with more targeted defense suggestions, effectively reduce the risk of core resources in the domain being breached, and ultimately improve the security protection capabilities of the entire network environment. Summary of the invention
[0005] In order to solve the problems existing in the background technology, the purpose of the present invention is to provide a Windows domain penetration testing system and method based on deep reinforcement learning, which utilizes automated penetration testing technology to realize fully automatic penetration testing and risk assessment of the domain environment, and combines deep reinforcement learning to more accurately, quickly and specifically discover various security issues in the Windows domain, assist security personnel in efficiently allocating resources, and concentrate on defending high-risk path areas, thereby enhancing overall network security.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] A Windows domain penetration testing system based on deep reinforcement learning includes an information collection module, a vulnerability detection module, an intelligent agent generation module, and an intelligent agent training module.
[0008] Among them, the information collection module is used to collect information of each network node in the target domain environment, including domain information and host information; the vulnerability detection module encapsulates various Exps for Windows domain environment vulnerabilities, which will be used for testing when the intelligent agent conducts penetration testing on the target domain environment, and scores each vulnerability based on the score of each attack and the actual penetration testing experience for interactive learning of the intelligent agent; the intelligent agent generation module uses all the domain information collected by the information collection module to generate a configuration file representing the domain, and uses the configuration file to build a deep reinforcement learning intelligent agent; the intelligent agent training module includes a vulnerability detection module, whose function is to enable the intelligent agent to interact with the target domain environment. The intelligent agent uses the attack method of the vulnerability detection module in the domain environment to conduct penetration testing on the entire domain environment. In the test, the next strategy is determined based on the constructed reward function, and finally a comprehensive penetration test is conducted on the entire domain environment and a test report is generated.
[0009] A Windows domain penetration testing method based on deep reinforcement learning includes the following steps:
[0010] 1) Select a device from the domain as the initial device, deploy the information collection module, conduct in-depth information collection on all network nodes in the domain, and fully obtain the target domain environment information, including host information and domain information;
[0011] 2) Sort out the collected domain information, extract sensitive information from the domain controller and host, generate a configuration file, and use the configuration file to initialize the deep reinforcement learning agent and state space, action space, and reward function;
[0012] 3) Use the constructed intelligent agent to perform automated penetration testing on the entire domain environment, use the vulnerability detection module to detect potential problems that may exist in the domain, and combine the reward mechanism to score the results of each action of the intelligent agent to mark the threat level of the problems in the domain environment. Finally, a test report for the entire domain environment is formed to help security personnel use limited resources for key protection.
[0013] Furthermore, the step 1) is specifically as follows:
[0014] (11) Select a controlled host in the subdomain as the starting point for the operation. After logging in to the host, deploy the information collection tool to ensure that the tool can run normally and establish a connection with the resources in the domain.
[0015] (12) Using the two-way trust relationship of the Windows domain forest, establish an interactive connection with the Active Directory (AD) database. Through this connection, prepare to fully capture basic information related to the domain.
[0016] (13) Run the information collection tool to extract the following information from the Active Directory: domain name, domain identity information, list of common hosts in the domain, list of domain controllers, list of domain servers, list of members of high-privileged and special user groups, delegation information in the domain, and service principal name (SPN).
[0017] (14) Use information collection tools to send SMB probe packets and query the remote registry to collect more details about the hosts in the domain, such as the operating system version, the list of currently logged-in users, and the activity information of high-privilege users. This data will provide support for subsequent analysis and operations.
[0018] Further, the step 2) is specifically as follows:
[0019] (21) Conduct a detailed analysis of the information exposed in each host and each domain, and combine penetration testing experience and the severity of vulnerability threats to find potential vulnerability exploits, focusing on possible breakthroughs.
[0020] (22) Integrate the possible vulnerabilities analyzed with the detailed information collected previously to generate a configuration file containing vulnerability exploitation points and target environment information, providing a basis for the subsequent construction of reinforcement learning agents.
[0021] (23) Using the generated configuration file, the core components required by the deep reinforcement learning agent are defined, including: state space: abstract representation of all host information in the domain and the current permissions; action space: defines possible attack actions, such as exploiting vulnerabilities to escalate privileges or move laterally; reward function: used to evaluate the effect of each action and give rewards or penalties based on the degree to which the attack goal is achieved.
[0022] Furthermore, the state space construction in step (23) is specifically as follows:
[0023] (31) The agent’s authority over the target host can be divided into the following stages, which are used to measure the current level of mastery of the host information and authority within the domain:
[0024] N: The target host has not been authorized.
[0025] U: Get the common user rights of the target host (including domain common users and local common users).
[0026] L: Obtain local administrator privileges on the target host.
[0027] These stages reflect a gradual escalation of authority and provide a basis for partitioning the state space.
[0028] (32) The key information of each host in the domain is stored in a structured form to facilitate subsequent operations. The storage format is: "IP-operating system-port: service-vulnerability-acquired permissions". For example: "172.16.253.17-Windows7-22:SSH-CVE-2020-15778-L". This record shows that the host with IP 172.16.253.17 runs the Windows 7 operating system, provides SSH service on port 22, has the vulnerability CVE-2020-15778, and the current permission is the local administrator permission (L). This ensures that the status of each host is accurately recorded in the database.
[0029] (33) The penetration state of the intelligent agent mainly depends on the change of permissions, so the state space can be described by the number of each host and its corresponding permissions. State representation format: state number-acquired permissions Example: 1: L-2: U-3: N-4: N-5: N means: Host 1: has obtained local administrator permissions (L); Host 2: has obtained ordinary user permissions (U); Host 3, Host 4, Host 5: has not obtained any permissions (N). This format is convenient for dynamically updating the status of the host and adapting to different attack stages.
[0030] (34) In order to measure the final target state of domain penetration, “successful penetration of the domain controller” is defined as the target state G. If host n is a domain controller, its state can be:
[0031] N: Domain controller permissions are not obtained.
[0032] G: Successfully obtained domain controller permissions.
[0033] The expanded form of the state space is: 1:X-2:X-3:X-4:X-5:X-…n-1:Xn:Y, where:
[0034] X indicates the permission status of a normal host (N, U, L).
[0035] Y indicates the authority status of the domain controller (N or G).
[0036] For example, 1:L-2:N-3:L-4:U-5:N-…n-1:Nn:G means that the attacker's status before successfully infiltrating the domain controller is: Host 1: obtained local administrator privileges (L); Host 2: has not yet obtained privileges (N); Host 3: obtained local administrator privileges (L); Host 4: obtained ordinary user privileges (U); Host 5 and n-1: have not yet obtained privileges (N); Host n (domain controller): has successfully obtained target privileges (G).
[0037] (35) The state space consists of various combinations of N, U, L, and G, and its specific size is determined by the following factors:
[0038] Number of hosts (n).
[0039] The permission status of each host.
[0040] For example, for n hosts, the maximum size of the state space is 3^n (permutations of the three states N, U, and L). If host n is a domain controller, the additional target state G is added, and the state combination is further expanded.
[0041] (36) During the penetration process, the attacker's permissions will change dynamically as the attack path progresses, mainly reflected in: permissions gradually increase from N to U or L; the ultimate goal is to obtain the permissions of the domain controller (G). Through the combination of state number and permission description, the dynamic change trajectory of the attacker's control over resources in the domain is fully described, laying the foundation for decision-making in subsequent reinforcement learning.
[0042] Furthermore, the action space construction in step (23) is specifically as follows:
[0043] (41) Before the penetration test begins, clearly define each attack action to ensure that each action is an independent operation and is associated with a specific target entity. When designing actions, fully consider the frequency of vulnerability exploitation and the difficulty of actual execution to improve the targeting and effectiveness of the attack.
[0044] (42) During the infiltration process, in order to avoid repeatedly executing the same action on the same entity, the system marks the action for each target entity. Once an action has been executed, the status will be recorded to ensure that subsequent action selections no longer waste resources on repeated attempts, thereby shortening the running time.
[0045] (43) The action space is trimmed based on prior knowledge. For example, multiple vulnerabilities that meet the same prerequisites and produce the same results are uniformly classified into one action type. In this way, the types of actions are significantly reduced, optimizing the efficiency and time cost of training.
[0046] (44) In the operation, the common continuous operation steps in the actual penetration process (such as multiple dependent attack behaviors) are merged into a coherent action. This can reduce unnecessary intermediate processes, improve the efficiency of operation execution, and be closer to the real penetration test scenario.
[0047] Furthermore, the reward function in step (23) is constructed as follows:
[0048] (51) Define an immediate reward mechanism for the agent in reinforcement Q learning to ensure that it can get feedback after performing an action. The design goal of the reward is to guide the agent to gradually optimize its behavior and ultimately achieve the effect of maximizing the cumulative reward.
[0049] (52) During the infiltration process, when the intelligent agent causes the target host's permissions to be upgraded through specific actions (for example, from no permissions to ordinary user permissions, or from ordinary user permissions to administrator permissions), the system will assign different reward values according to the degree of permission change. The more significant the permission upgrade, the higher the reward.
[0050] (53) Reward rules are designed for penetration and permission changes of domain controllers. If the agent successfully obtains domain controller permissions or upgrades to domain administrator permissions through actions, the system will assign a higher reward value to encourage the agent to prioritize exploring such high-value targets.
[0051] (54) When the agent successfully obtains the credentials of a user in the domain (such as a plaintext password or password hash) through some attack method, the system will assign a reward to this action. This rule is designed to encourage the agent to try to obtain more user information to assist in subsequent penetration operations.
[0052] (55) By integrating the above reward conditions and adjusting the reward priority and score ratio according to the requirements of the infiltration task, the intelligent agent can find the best balance between exploration and utilization, quickly learn efficient strategies, and cover more potential paths.
[0053] Furthermore, the step 3) is specifically as follows:
[0054] (61) After configuring the agent and defining the state space, action space, and reward function, the agent starts with the host where the information collector is located and begins to perform automated penetration testing on the domain environment.
[0055] (62) The agent reads the configuration file and performs the penetration operation step by step according to the action space defined therein. After each successful action, the system updates the current state space and assigns a corresponding reward value according to the reward function to guide the agent to optimize subsequent penetration behavior.
[0056] (63) In a complete penetration, if the agent successfully achieves the final goal (such as obtaining the root domain controller authority of the target domain), the system will record this attack path and the total reward value obtained by completing this path. The agent will repeat this process and continue training until all penetration tasks are completed.
[0057] (64) After the training is completed, the system sorts all the vulnerabilities found in the domain and sorts them according to the threat reward value. In this way, the threat level of each vulnerability can be intuitively displayed, which is convenient for further analysis and exploitation.
[0058] (65) The intelligent agent has dynamic adaptability and can adjust strategies in real time to deal with emergencies or the emergence of new vulnerabilities. In different domain environments, the intelligent agent maintains efficient penetration performance to ensure the comprehensiveness and effectiveness of penetration testing.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] It realizes efficient automation of information collection, vulnerability exploitation and path generation, greatly improving penetration efficiency; it has dynamic adaptability and can adjust strategies in real time to cope with environmental changes; it comprehensively covers complex vulnerability combinations and potential paths to discover hidden threats; it gives priority to high-threat paths through a precise reward mechanism; it simulates real attack behaviors to be closer to actual scenarios; it supports path threat sorting to help prioritize the protection of high-risk areas; it has strong scalability and sustainable update capabilities; it reduces false alarms and false detections to provide more accurate test results; it supports multi-stage testing to gradually simulate the escalation of permissions. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 The flowchart of the present invention from information collection to building an intelligent agent to automatic penetration of the domain environment is shown;
[0062] Figure 2 It is a flow chart of information collection by the information collector in the present invention;
[0063] Figure 3 A structural diagram of the parameters required to construct an intelligent agent in the present invention;
[0064] Figure 4 Generate an attack path flow chart for the present invention; DETAILED DESCRIPTION
[0065] The following will disclose the embodiments of the present invention with drawings. For the purpose of clear description, many physical details will be described together in the following description. However, it should be understood that these physical details should not be used to limit the present invention. In other words, in some embodiments of the present invention, these practical details are not necessary.
[0066] A Windows domain penetration testing system based on deep reinforcement learning includes an information collection module, a vulnerability detection module, an intelligent agent generation module, and an intelligent agent training module.
[0067] Among them, the information collection module is used to collect information of each network node in the target domain environment, including domain information and host information; the vulnerability detection module encapsulates various Exps for Windows domain environment vulnerabilities, which will be used for testing when the intelligent agent conducts penetration testing on the target domain environment, and scores each vulnerability based on the score of each attack and the actual penetration testing experience for interactive learning of the intelligent agent; the intelligent agent generation module uses all the domain information collected by the information collection module to generate a configuration file representing the domain, and uses the configuration file to build a deep reinforcement learning intelligent agent; the intelligent agent training module includes a vulnerability detection module, whose function is to enable the intelligent agent to interact with the target domain environment. The intelligent agent uses the attack method of the vulnerability detection module in the domain environment to conduct penetration testing on the entire domain environment. In the test, the next strategy is determined based on the constructed reward function, and finally a comprehensive penetration test is conducted on the entire domain environment and a test report is generated.
[0068] A Windows domain penetration testing method based on deep reinforcement learning, such as Figure 1 As shown, the following steps are included:
[0069] 1) Select a device from the domain as the initial device, deploy the information collection module, conduct in-depth information collection on all network nodes in the domain, and fully obtain the target domain environment information, including host information and domain information;
[0070] 2) Sort out the collected domain information, extract sensitive information from the domain controller and host, generate a configuration file, and use the configuration file to initialize the deep reinforcement learning agent and state space, action space, and reward function;
[0071] 3) Use the constructed intelligent agent to perform automated penetration testing on the entire domain environment, use the vulnerability detection module to detect potential problems that may exist in the domain, and combine the reward mechanism to score the results of each action of the intelligent agent to mark the threat level of the problems in the domain environment. Finally, a test report for the entire domain environment is formed to help security personnel use limited resources for key protection.
[0072] like Figure 2As shown, step 1) is specifically as follows:
[0073] (11) Select a controlled host in the subdomain as the starting point for the operation. After logging in to the host, deploy the information collection tool to ensure that the tool can run normally and establish a connection with the resources in the domain.
[0074] (12) Using the two-way trust relationship of the Windows domain forest, establish an interactive connection with the Active Directory (AD) database. Through this connection, prepare to fully capture basic information related to the domain.
[0075] (13) Run the information collection tool to extract the following information from the Active Directory: domain name, domain identity information, list of common hosts in the domain, list of domain controllers, list of domain servers, list of members of high-privileged and special user groups, delegation information in the domain, and service principal name (SPN).
[0076] (14) Use information collection tools to send SMB probe packets and query the remote registry to collect more details about the hosts in the domain, such as the operating system version, the list of currently logged-in users, and the activity information of high-privilege users. This data will provide support for subsequent analysis and operations.
[0077] like Figure 3 As shown, step 2) is specifically as follows:
[0078] (21) Conduct a detailed analysis of the information exposed in each host and each domain, and combine penetration testing experience and the severity of vulnerability threats to find potential vulnerability exploits, focusing on possible breakthroughs.
[0079] (22) Integrate the possible vulnerabilities analyzed with the detailed information collected previously to generate a configuration file containing vulnerability exploitation points and target environment information, providing a basis for the subsequent construction of reinforcement learning agents.
[0080] (23) Using the generated configuration file, the core components required by the deep reinforcement learning agent are defined, including: state space: abstract representation of all host information in the domain and the current permissions; action space: defines possible attack actions, such as exploiting vulnerabilities to escalate privileges or move laterally; reward function: used to evaluate the effect of each action and give rewards or penalties based on the degree to which the attack goal is achieved.
[0081] Step (23) defines the various core components required for a deep reinforcement learning agent.
[0082] The state space is constructed as follows:
[0083] (31) The agent’s authority over the target host can be divided into the following stages, which are used to measure the current level of mastery of the host information and authority within the domain:
[0084] N: The target host has not been authorized.
[0085] U: Get the common user rights of the target host (including domain common users and local common users).
[0086] L: Obtain local administrator privileges on the target host.
[0087] These stages reflect a gradual escalation of authority and provide a basis for partitioning the state space.
[0088] (32) The key information of each host in the domain is stored in a structured form to facilitate subsequent operations. The storage format is: "IP-operating system-port: service-vulnerability-acquired permissions". For example: "172.16.253.17-Windows7-22:SSH-CVE-2020-15778-L". This record shows that the host with IP 172.16.253.17 runs the Windows 7 operating system, provides SSH service on port 22, has the vulnerability CVE-2020-15778, and the current permission is the local administrator permission (L). This ensures that the status of each host is accurately recorded in the database.
[0089] (33) The penetration state of the intelligent agent mainly depends on the change of permissions, so the state space can be described by the number of each host and its corresponding permissions. State representation format: state number-acquired permissions Example: 1: L-2: U-3: N-4: N-5: N means: Host 1: has obtained local administrator permissions (L); Host 2: has obtained ordinary user permissions (U); Host 3, Host 4, Host 5: has not obtained any permissions (N). This format is convenient for dynamically updating the status of the host and adapting to different attack stages.
[0090] (34) In order to measure the final target state of domain penetration, “successful penetration of the domain controller” is defined as the target state G. If host n is a domain controller, its state can be:
[0091] N: Domain controller permissions are not obtained.
[0092] G: Successfully obtained domain controller permissions.
[0093] The expanded form of the state space is: 1:X-2:X-3:X-4:X-5:X-…n-1:Xn:Y, where:
[0094] X indicates the permission status of a normal host (N, U, L).
[0095] Y indicates the authority status of the domain controller (N or G).
[0096] For example, 1:L-2:N-3:L-4:U-5:N-…n-1:Nn:G means that the attacker's status before successfully infiltrating the domain controller is: Host 1: obtained local administrator privileges (L); Host 2: has not yet obtained privileges (N); Host 3: obtained local administrator privileges (L); Host 4: obtained ordinary user privileges (U); Host 5 and n-1: have not yet obtained privileges (N); Host n (domain controller): has successfully obtained target privileges (G).
[0097] (35) The state space consists of various combinations of N, U, L, and G, and its specific size is determined by the following factors:
[0098] Number of hosts (n).
[0099] The permission status of each host.
[0100] For example, for n hosts, the maximum size of the state space is 3^n (permutations of the three states N, U, and L). If host n is a domain controller, the additional target state G is added, and the state combination is further expanded.
[0101] (36) During the penetration process, the attacker's permissions will change dynamically as the attack path progresses, mainly reflected in: permissions gradually increase from N to U or L; the ultimate goal is to obtain the permissions of the domain controller (G). Through the combination of state number and permission description, the dynamic change trajectory of the attacker's control over resources in the domain is fully described, laying the foundation for decision-making in subsequent reinforcement learning.
[0102] The action space is constructed as follows:
[0103] (41) Before the penetration test begins, clearly define each attack action to ensure that each action is an independent operation and is associated with a specific target entity. When designing actions, fully consider the frequency of vulnerability exploitation and the difficulty of actual execution to improve the targeting and effectiveness of the attack.
[0104] (42) During the infiltration process, in order to avoid repeatedly executing the same action on the same entity, the system marks the action for each target entity. Once an action has been executed, the status will be recorded to ensure that subsequent action selections no longer waste resources on repeated attempts, thereby shortening the running time.
[0105] (43) The action space is trimmed based on prior knowledge. For example, multiple vulnerabilities that meet the same prerequisites and produce the same results are uniformly classified into one action type. In this way, the types of actions are significantly reduced, optimizing the efficiency and time cost of training.
[0106] (44) In the operation, the common continuous operation steps in the actual penetration process (such as multiple dependent attack behaviors) are merged into a coherent action. This can reduce unnecessary intermediate processes, improve the efficiency of operation execution, and be closer to the real penetration test scenario.
[0107] The reward function is constructed as follows:
[0108] (51) Define an immediate reward mechanism for the agent in reinforcement Q learning to ensure that it can get feedback after performing an action. The design goal of the reward is to guide the agent to gradually optimize its behavior and ultimately achieve the effect of maximizing the cumulative reward.
[0109] (52) During the infiltration process, when the intelligent agent causes the target host's permissions to be upgraded through specific actions (for example, from no permissions to ordinary user permissions, or from ordinary user permissions to administrator permissions), the system will assign different reward values according to the degree of permission change. The more significant the permission upgrade, the higher the reward.
[0110] (53) Reward rules are designed for penetration and permission changes of domain controllers. If the agent successfully obtains domain controller permissions or upgrades to domain administrator permissions through actions, the system will assign a higher reward value to encourage the agent to prioritize exploring such high-value targets.
[0111] (54) When the agent successfully obtains the credentials of a user in the domain (such as a plaintext password or password hash) through some attack method, the system will assign a reward to this action. This rule is designed to encourage the agent to try to obtain more user information to assist in subsequent penetration operations.
[0112] (55) By integrating the above reward conditions and adjusting the reward priority and score ratio according to the requirements of the infiltration task, the intelligent agent can find the best balance between exploration and utilization, quickly learn efficient strategies, and cover more potential paths.
[0113] like Figure 4 As shown, step 3) is specifically as follows:
[0114] (61) After configuring the agent and defining the state space, action space, and reward function, the agent starts with the host where the information collector is located and begins to perform automated penetration testing on the domain environment.
[0115] (62) The agent reads the configuration file and performs the penetration operation step by step according to the action space defined therein. After each successful action, the system updates the current state space and assigns a corresponding reward value according to the reward function to guide the agent to optimize subsequent penetration behavior.
[0116] (63) In a complete penetration, if the agent successfully achieves the final goal (such as obtaining the root domain controller authority of the target domain), the system will record this attack path and the total reward value obtained by completing this path. The agent will repeat this process and continue training until all penetration tasks are completed.
[0117] (64) After the training is completed, the system sorts all the vulnerabilities found in the domain and sorts them according to the threat reward value. In this way, the threat level of each vulnerability can be intuitively displayed, which is convenient for further analysis and exploitation.
[0118] (65) The intelligent agent has dynamic adaptability and can adjust strategies in real time to deal with emergencies or the emergence of new vulnerabilities. In different domain environments, the intelligent agent maintains efficient penetration performance to ensure the comprehensiveness and effectiveness of penetration testing.
Claims
1. A Windows domain penetration testing system based on deep reinforcement learning, characterized in that: It includes information collection module, vulnerability detection module, agent generation module and agent training module; Among them, the information collection module is used to collect information of each network node in the target domain environment, including domain information and host information; the vulnerability detection module encapsulates various Exps for Windows domain environment vulnerabilities, which are used for testing when the intelligent agent conducts penetration testing on the target domain environment, and scores each vulnerability based on the score of each attack and the actual penetration testing experience for interactive learning of the intelligent agent; the intelligent agent generation module uses all the domain information collected by the information collection module to generate a configuration file representing the domain, and uses the configuration file to build a deep reinforcement learning intelligent agent; the intelligent agent training module includes a vulnerability detection module, which enables the intelligent agent to interact with the target domain environment. The intelligent agent uses the attack method of the vulnerability detection module in the domain environment to conduct penetration testing on the entire domain environment. In the test, the next strategy is determined based on the constructed reward function, and finally a comprehensive penetration test is conducted on the entire domain environment and a test report is generated.
2. A Windows domain penetration testing method based on deep reinforcement learning, characterized in that: The following steps are involved: 1) Select a device from the domain as the initial device, deploy the information collection module, conduct in-depth information collection on all network nodes in the domain, and fully obtain the target domain environment information, including host information and domain information; 2) Sort out the collected domain information, extract sensitive information from the domain controller and host, generate a configuration file, and use the configuration file to initialize the deep reinforcement learning agent and state space, action space, and reward function; 3) Use the constructed intelligent agent to perform automated penetration testing on the entire domain environment, use the vulnerability detection module to detect potential problems that may exist in the domain, and combine the reward mechanism to score the results of each action of the intelligent agent to mark the threat level of the problems in the domain environment. Finally, a test report for the entire domain environment is formed to help security personnel use limited resources for key protection.
3. According to a Windows domain penetration testing method based on deep reinforcement learning according to claim 2, it is characterized in that: The step 1) is specifically as follows: (11) Select a controlled host in the subdomain as the operation starting point. After logging into the host, deploy the information collection tool to ensure that the tool can run normally and establish a connection with the resources in the domain. (12) Using the two-way trust relationship of the Windows domain forest, establish an interactive connection with the Active Directory database, and through this connection, prepare to fully capture basic information related to the domain; (13) Run the information collection tool to extract the following from the Active Directory: domain name, domain identity information, list of common hosts in the domain, list of domain controllers, list of domain servers, list of members of high-privileged and special user groups, delegation information in the domain, and service principal names; (14) Use information collection tools to send SMB probe packets and query the remote registry to collect more details about the hosts in the domain to provide support for subsequent analysis and operations.
4. According to a Windows domain penetration testing method based on deep reinforcement learning according to claim 2, it is characterized in that: The step 2) is specifically as follows: (21) Conduct a detailed analysis of the information exposed in each host and each domain, and combine penetration testing experience and the severity of vulnerability threats to find potential vulnerability exploits, focusing on possible breakthroughs; (22) Integrate the analyzed possible vulnerabilities with the detailed information collected previously to generate a configuration file containing vulnerability exploitation points and target environment information, providing a basis for the subsequent construction of reinforcement learning agents; (23) Using the generated configuration file, the core components required by the deep reinforcement learning agent are defined, including: state space: abstract representation of all host information in the domain and the current authority status; action space: defines possible attack actions; reward function: used to evaluate the effect of each action and give rewards or penalties based on the degree of achievement of the attack goal.
5. According to a Windows domain penetration testing method based on deep reinforcement learning according to claim 4, it is characterized in that: The state space construction in step (23) is specifically as follows: (31) The agent’s authority over the target host is divided into the following stages, which are used to measure the current level of mastery of host information and authority within the domain: N: The target host has not been authorized. U: Get the common user rights of the target host, including domain common users and local common users. L: Get local administrator privileges on the target host. These stages reflect the gradual increase in authority and provide a basis for the division of the state space; (32) The key information of each host in the domain is stored in a structured form to facilitate subsequent operations; the storage format is: "IP-operating system-port: service-existing vulnerability-acquired permissions", ensuring that the status of each host is accurately recorded in the database; (33) The penetration status of the intelligent agent mainly depends on the change of permissions. Therefore, the state space is described by the number of each host and its corresponding permissions. The state representation format is: state number-acquired permissions. This format is convenient for dynamically updating the status of the host and adapting to different attack stages. (34) In order to measure the final target state of domain penetration, "successful penetration of the domain controller" is defined as the target state G. If the host is a domain controller, its state is: N: Domain controller permission is not obtained; G: Successfully obtained domain controller permissions; The expanded form of the state space is: 1:X-2:X-3:X-4:X-5:X-…n-1:Xn:Y, where: X indicates the permission status of a common host: N, U, L; Y indicates the permission status of the domain controller: N or G; (35) The state space consists of various combinations of N, U, L, and G, and its specific size is determined by the following factors: The number of hosts n, The permission status of each host; (36) During the penetration process, the attacker's permissions change dynamically as the attack path progresses, which is reflected in: the permissions are gradually increased from N to U or L; the ultimate goal is to obtain the permission G of the domain controller; through the combination of state number and permission description, the dynamic change trajectory of the attacker's control over the resources in the domain is fully described, laying the foundation for decision-making in subsequent reinforcement learning.
6. A Windows domain penetration testing method based on deep reinforcement learning according to claim 4, characterized in that: The action space construction in step (23) is specifically as follows: (41) Before the penetration test begins, clearly define each attack action to ensure that each action is an independent operation and is associated with a specific target entity; (42) During the infiltration process, in order to avoid repeatedly executing the same action on the same entity, the system marks each target entity with an action. Once an action has been executed, the status will be recorded to ensure that subsequent action selections no longer waste resources on repeated attempts, thereby shortening the runtime; (43) The action space is pruned based on prior knowledge. In this way, the types of actions are significantly reduced, optimizing the efficiency and time cost of training. (44) In operation, the common continuous operation steps in the actual penetration process are merged into a coherent action, which reduces unnecessary intermediate processes, improves the efficiency of operation execution, and is closer to the actual penetration test scenario.
7. A Windows domain penetration testing method based on deep reinforcement learning according to claim 4, characterized in that: The reward function of step (23) is specifically constructed as follows: (51) Define an immediate reward mechanism for the agent in reinforcement Q learning to ensure that it receives feedback after performing an action. The design goal of the reward is to guide the agent to gradually optimize its behavior and ultimately achieve the effect of maximizing the cumulative reward. (52) During the infiltration process, when the agent causes the target host's permissions to increase through specific actions, the system assigns different reward values according to the degree of permission change. The more significant the permission increase, the higher the reward. (53) Reward rules are designed for the penetration and permission change of domain controllers. If the agent successfully obtains domain controller permissions or upgrades to domain administrator permissions through actions, the system assigns a higher reward value to encourage the agent to prioritize exploring such high-value targets. (54) When the agent successfully obtains the credentials of a user in the domain through some attack method, the system assigns a reward to this action. This rule is intended to encourage the agent to try to obtain more user information to assist in subsequent penetration operations; (55) By integrating the above reward conditions and adjusting the reward priority and score ratio according to the requirements of the infiltration task, the intelligent agent can find the best balance between exploration and utilization, quickly learn efficient strategies, and cover more potential paths.
8. A Windows domain penetration testing method based on deep reinforcement learning according to claim 2, characterized in that: The step 3) is specifically as follows: (61) After configuring the agent and defining the state space, action space, and reward function, the agent starts with the host where the information collector is located and begins to perform automated penetration testing on the domain environment; (62) The agent reads the configuration file and performs the penetration operation step by step according to the action space defined therein. After each successful action, the system updates the current state space and assigns a corresponding reward value according to the reward function to guide the agent to optimize subsequent penetration behavior. (63) In a complete penetration, if the agent successfully achieves the final goal, the system records this attack path and the total reward value obtained by completing this path. The agent will repeat this process and continue training until all penetration tasks are completed; (64) After the training is completed, the system organizes all the vulnerabilities found in the domain and sorts them according to the threat reward value, intuitively displaying the threat level of each vulnerability to facilitate further analysis and exploitation; (65) The intelligent agent has the ability to dynamically adapt and adjust its strategies in real time to deal with emergencies or the emergence of new vulnerabilities. In different domain environments, the intelligent agent maintains efficient penetration performance to ensure the comprehensiveness and effectiveness of the penetration test.
Citation Information
Cited By
Domain penetration attack path generation method based on graph structure
CN120750674A
Domain penetration testing method based on adaptive vulnerability utilization
CN121485996A