Intelligent network vulnerability repairing method and system based on penetration testing technology
By applying reinforcement learning, domain randomization and meta-reinforcement learning technologies in penetration testing agents, the problems of high cost, long time and error-prone traditional penetration testing methods are solved, and efficient vulnerability identification and repair in the new environment is achieved, and network security is improved.
Patent Information
- Application Number
- CN202510539660.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Traditional penetration testing methods rely on manual operations, which are costly, long-term and prone to errors, making it difficult to effectively deal with large-scale and dynamically changing network environments.
Adopting an autonomous penetration testing agent based on reinforcement learning, training agents through domain randomization and meta-reinforcement learning techniques allows them to generalize and automatically identify and repair network vulnerabilities in unseen real environments.
It significantly improves the adaptability and performance of the agent in the new environment, improves the efficiency of vulnerability identification and repair, reduces cost and time, and enhances the security and reliability of the network system.
Smart Images

Figure CN120090872A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and more specifically, to a network vulnerability intelligent repair method and system based on penetration testing technology. Background Art
[0002] With the rapid development of network technology, network security issues have become increasingly prominent. Network security vulnerabilities are usually identified and repaired by penetration testing. Traditional penetration testing methods mainly rely on manual operations, which are costly, time-consuming, and error-prone. Manual testing requires professional security personnel, and has high requirements for the technical level and experience of personnel. In addition, the strategies trained in the simulation environment by traditional methods are often difficult to be directly applied to the real environment, there is a "reality gap", resulting in poor performance of the strategies in the new environment. These limitations make it difficult for traditional penetration testing methods to effectively cope with large-scale and dynamically changing network environments. Summary of the Invention
[0003] The object of the present invention is to propose a network vulnerability intelligent repair method and system based on penetration testing technology, and train an autonomous penetration testing agent that can be generalized to unseen real environments through domain randomization and meta-reinforcement learning, improve the adaptability and performance of the agent in the new environment, so as to improve the efficiency of vulnerability identification and repair.
[0004] To achieve the above object, the present invention proposes a network vulnerability intelligent repair method based on penetration testing technology, including:
[0005] S1: Construct an agent for penetration testing based on a reinforcement learning model;
[0006] S2: Make the agent perform end-to-end policy learning in the target network environment, learn how to explore and exploit the vulnerabilities of the target host through interaction with the original training environment, and at the same time the agent uses a reinforcement learning algorithm to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment;
[0007] S3: During the interaction between the agent and the target network environment, use the collected host configuration data to construct an original simulation environment, and the original simulation environment is a digital mapping of the original training environment;
[0008] S4: Based on the original simulation environment and the vulnerability description, the agent uses domain randomization technology to randomize the host configuration parameters and the vulnerability description of the original simulation environment, and generate diverse synthetic environments;
[0009] S5: In the synthesis environment, the agent uses the Proximal Policy Optimization algorithm and the model - agnostic meta - learning algorithm for in - depth training to optimize the agent's policy parameters and initial parameters, enabling the agent to extract generalizable policies and biases from diverse meta - training environments to generalize the learned policies to new or similar network environments;
[0010] S6: Test the generalization ability of the trained agent in the test environment;
[0011] S7: Apply the agent that has completed training and passed the test to the real target network environment for vulnerability identification. The agent automatically identifies network vulnerabilities based on the actions performed and the rewards received in the network environment;
[0012] S8: Select corresponding repair strategies for repair based on the identified network vulnerabilities.
[0013] Optionally, in step S1, constructing an agent for penetration testing based on a reinforcement learning model includes:
[0014] S101: Define the reinforcement learning framework of the agent, including:
[0015] State space S: The environmental state observed by the agent, including host configuration, ports, services, operating system, and website fingerprint information;
[0016] Action space A: The actions that the agent can perform, such as information collection, system detection, and vulnerability exploitation;
[0017] Reward function R: The learning objective of the agent, that is, while minimizing the action cost, collecting key information from the target host and exploiting vulnerabilities;
[0018] S102: Define the policy learning of the agent, including:
[0019] Policy network: The agent uses a neural network to represent its policy. This network maps states to actions, and the parameters of the policy network are updated through reinforcement learning algorithms;
[0020] Reinforcement learning algorithm: The agent uses the Proximal Policy Optimization algorithm to update the policy parameters;
[0021] S103: Define the environmental interaction of the agent, including:
[0022] Observe the environment: The agent uses scanning tools to observe the environmental state of the target host and collect information;
[0023] Decision - making: Based on the collected information, the agent uses its policy to decide the next actions, which include information collection, system detection, and vulnerability exploitation;
[0024] Execute action: The agent executes the selected action by calling and executing commands in the penetration testing toolkit;
[0025] Receive reward: The agent receives a reward based on the result of the executed action. The reward function gives a positive or negative reward according to whether the vulnerability is successfully exploited and the information collected;
[0026] S103: Define the policy update of the agent:
[0027] Experience replay: The agent stores the experience of each interaction in the experience replay pool. The experience includes state, action, reward, and next state;
[0028] Policy optimization: The agent samples experiences from the experience replay pool and uses the proximal policy optimization algorithm to update the parameters of the policy network to maximize the expected cumulative reward.
[0029] Optionally, the reward function R is:
[0030]
[0031] where value(h) is the reward for successfully identifying the vulnerability of the target host h, and cost(a) is the cost of action a.
[0032] Optionally, the proximal policy optimization algorithm is a reinforcement learning algorithm based on policy gradient. It minimizes the objective function through multi-step stochastic gradient descent to optimize the policy. The objective function of the proximal policy optimization algorithm is:
[0033]
[0034] where ϕ are the parameters of the policy network; π is the current policy; is the trajectory under the policy π, that is, a sequence of states, actions, and rewards; is the probability ratio, which represents the ratio of the probability of taking action at under the current policy ϕ to the probability of taking the same action under the old policy ϕ old That is, ; is the estimate of the advantage function, which represents the additional benefit of taking action at in state st compared to the average case; represents clipping the probability ratio
[0035] Optionally, in step S4, the domain randomization technique is used to randomize the host configuration parameters and vulnerability descriptions of the original simulation environment, and a diversified synthetic environment is generated, including:
[0036] Use the original simulation environment and known vulnerability descriptions as the input to a pre-trained large language model. Generate diverse environment configurations and vulnerability descriptions through the large language model, and output a synthetic environment.
[0037] Optionally, in step S5, the meta-learning algorithm includes inner-loop training and outer-loop training;
[0038] In inner-loop training, the agent samples trajectories in each synthetic environment and updates the policy parameters;
[0039] In outer-loop training, the agent updates the initial parameters, optimizing the initial parameters of the model so that the agent can quickly adapt to new tasks.
[0040] Optionally, the calculation formula for the inner-loop training is:
[0041]
[0042] where ϕ is the parameter of the meta-policy; ϕ i ′ is the adapted policy parameter under task Mi; α is the learning rate of the inner loop; is the estimate of the expected discounted reward under task M i ; is the policy gradient estimate under task M i ; D i is the trajectory data collected under task Mi;
[0043] The calculation formula for the outer-loop training is:
[0044]
[0045] where ϕ is the parameter of the meta-policy; β is the learning rate of the outer loop; is the task distribution, containing all tasks used for meta-training; is the adapted policy under task M i ; D i ′ is the trajectory data collected under task M i using the adapted policy ; represents the estimate of the expected discounted reward under task using the adapted policy ; represents the meta-policy gradient estimate, i.e., the sum of the policy gradient estimates over all tasks M i ;
[0046] Optionally, in step S6, testing the generalization ability of the trained agent in the test environment includes:
[0047] The agent conducts a zero-shot policy transfer test in a test environment similar to the training environment and evaluates the generalization gap. The formula for calculating the generalization gap is:
[0048]
[0049] Where is the generalization gap; π is the policy, which is a function representing the probability distribution that maps the state s of the environment to the action a, that is, π(a∣s) represents the probability of taking action a in state s; τ is the trajectory under the policy π, that is, a sequence of states, actions, and rewards; P π is the trajectory distribution under the policy π; G(τ) is the expected cumulative reward of the trajectory τ, , where r t is the reward at time step t, γ is the discount factor used to reduce the weight of future rewards, and T is the total number of time steps; is the simulated environment or the synthetic environment generated through domain randomization; is the test environment, which is the target environment where the agent conducts zero-shot policy transfer; , respectively represent the expected cumulative rewards of the trajectory τ generated according to the policy π in the environments , , which is the average performance metric of the policy π in the environments , .
[0050] Optionally, in step S6, testing the generalization ability of the trained agent in the test environment also includes:
[0051] The agent conducts a fast policy adaptation test in a test environment not similar to the training environment and improves the agent's performance through a small amount of fine-tuning. The formula for calculating the adapted policy parameters is:
[0052]
[0053] Where ϕ is the parameter of the initial policy; ϕ adapted is the parameter of the adapted policy; α is the learning rate; π ϕ represents the policy with parameter ϕ; D test is the trajectory data collected under the test environment M test ; is the estimate of the expected discounted reward under the test environment M test ; It is the policy gradient estimation in the test environment M test underneath.
[0054] The present invention also proposes an intelligent network vulnerability repair system based on penetration testing technology, including:
[0055] An agent construction module, configured to construct an agent for penetration testing based on a reinforcement learning model;
[0056] A training module, configured to execute:
[0057] Enable the agent to perform end-to-end policy learning in the target network environment, learn how to explore and exploit the vulnerabilities of the target host through interaction with the original training environment, and at the same time the agent uses a reinforcement learning algorithm to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment;
[0058] During the interaction between the agent and the target network environment, the agent constructs an original simulation environment using the collected host configuration data, and the original simulation environment is a digital mapping of the original training environment;
[0059] Based on the original simulation environment and vulnerability description, the agent uses domain randomization technology to randomize the host configuration parameters and vulnerability description of the original simulation environment, and generates a diverse synthetic environment;
[0060] In the synthetic environment, the agent uses the proximal policy optimization algorithm and the model-agnostic meta-learning algorithm for in-depth training to optimize the agent's policy parameters and initial parameters, enabling the agent to extract generalizable policies and biases from diverse meta-training environments to generalize the learned policies to new or similar network environments;
[0061] Test the generalization ability of the trained agent in the test environment;
[0062] A vulnerability identification module, configured to apply the trained and tested agent to a real target network environment for vulnerability identification, and the agent automatically identifies network vulnerabilities based on the actions performed and the rewards received in the network environment;
[0063] A vulnerability repair module, configured to select a corresponding repair strategy for repair based on the identified network vulnerabilities.
[0064] The beneficial effects of the present invention are as follows:
[0065] The present invention trains an agent by combining domain randomization and meta-reinforcement learning techniques, achieving zero-shot policy transfer and rapid policy adaptation, significantly improving the agent's policy learning and adaptation capabilities in unknown environments. By enabling the agent to automatically identify network vulnerabilities and repair them automatically, the efficiency of penetration testing and vulnerability repair is effectively improved, enhancing the security and reliability of the network system, and providing an efficient and automated solution for network security testing and intelligent vulnerability repair.
[0066] The system of the present invention has other characteristics and advantages that will be apparent from or will be described in detail in the accompanying drawings and the subsequent detailed description incorporated herein. These accompanying drawings and detailed description are used together to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] By describing the exemplary embodiments of the present invention in more detail with reference to the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more apparent. In the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0068] Figure 1 The flowchart of a method for intelligent repair of network vulnerabilities based on penetration testing technology according to an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] The present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0070] As Figure 1 shown, an embodiment of the present invention provides a method for intelligent repair of network vulnerabilities based on penetration testing technology, including:
[0071] S1: Construct an agent for penetration testing based on a reinforcement learning model;
[0072] This step specifically includes:
[0073] S101: Define the reinforcement learning framework of the agent (model the penetration testing process as a Markov decision process), including:
[0074] State space S: The environmental state observed by the agent, including host configuration, ports, services, operating system, and website fingerprint information;
[0075] Action space A: The actions that the agent can perform, such as information collection, system detection, and vulnerability exploitation;
[0076] Reward function R: The learning objective of the agent, that is, while minimizing the action cost, collecting key information from the target host and exploiting vulnerabilities; the reward function R is as follows:
[0077]
[0078] Among them, value(h) is the reward for successfully identifying the vulnerability of the target host h, and cost(a) is the cost of action a.
[0079] S102: Define the policy learning of the agent, including:
[0080] Policy network: The agent uses a neural network to represent its policy. This network maps states to actions, and the parameters of the policy network are updated through a reinforcement learning algorithm;
[0081] Reinforcement learning algorithm: The agent uses the Proximal Policy Optimization algorithm to update the policy parameters; the Proximal Policy Optimization algorithm is a policy gradient-based reinforcement learning algorithm that minimizes the objective function through multi-step stochastic gradient descent to optimize the policy. The objective function of the Proximal Policy Optimization algorithm is as follows:
[0082]
[0083] Among them, ϕ is the parameter of the policy network; π is the current policy; is the trajectory under the policy π, that is, a sequence of states, actions, and rewards; is the probability ratio, indicating the ratio of the probability of taking action at under the current policy ϕ to the probability of taking the same action under the old policy ϕ old That is, ; is the estimate of the advantage function, indicating the additional benefit of taking action at in state st compared to the average situation; represents clipping the probability ratio
[0084] S103: Define the environment interaction of the agent, including:
[0085] Observe the environment: The agent uses a scanning tool to observe the environmental status of the target host and collect information;
[0086] Decision-making: Based on the collected information, the agent uses its policy to decide the next action, and these actions include information collection, system detection, and vulnerability exploitation;
[0087] Execute actions: The agent executes the selected actions by calling and executing commands in the penetration testing toolkit;
[0088] Receive rewards: The agent receives rewards based on the results of the executed actions. The reward function gives positive or negative rewards according to whether the vulnerability is successfully exploited and the information collected;
[0089] S104: Define the policy update of the agent:
[0090] Experience replay: The agent stores the experience of each interaction in the experience replay pool. The experience includes the state, action, reward, and next state;
[0091] Policy optimization: The agent samples experiences from the experience replay pool and uses the Proximal Policy Optimization algorithm to update the parameters of the policy network to maximize the expected cumulative reward.
[0092] Furthermore, the functions implemented by the agent mainly include:
[0093] (1) Information collection:
[0094] Scanning tools: The agent uses various scanning tools (such as Nmap, whatweb, dirb, etc.) to collect detailed information about the target host, including open ports, running services, operating system type, website fingerprints, etc.
[0095] (2) Information analysis: The agent uses the pre-trained Sentence-BERT model to embed the collected text information into the state vector so that the neural network can understand and process this information.
[0096] (3) Vulnerability exploitation:
[0097] Select vulnerabilities: Based on the information collected, the agent selects appropriate vulnerability exploitation tools or payloads to attack the target host.
[0098] Execute attacks: The agent executes the selected vulnerability exploitation actions by calling and executing commands in the penetration testing toolkit (such as MSF) and attempts to compromise the target host.
[0099] (4) Policy optimization:
[0100] Reward function: The goal of the agent is to maximize the expected cumulative reward. The reward function defines the positive rewards obtained by the agent when successfully exploiting vulnerabilities and collecting key information, as well as the negative rewards received when executing invalid or incorrect actions.
[0101] (5) Policy update: The agent uses the Proximal Policy Optimization algorithm to update its policy parameters and minimizes the objective function through multi-step Stochastic Gradient Descent (SGD) to optimize the policy.
[0102] This embodiment does not limit the specific network structure of the agent. The agent can adopt any suitable reinforcement learning AI model or an intelligent agent integrating an AI model. In one example, the network structure of the agent can be a deep neural network, and its specific structure includes the following parts:
[0103] Input layer:
[0104] State input: The input layer receives the state information of the environment, such as host configuration, ports, services, operating systems, website fingerprints, etc. These state information are usually encoded as high-dimensional vectors.
[0105] Hidden layer:
[0106] Multilayer perceptron (MLP): Use a multilayer perceptron as the hidden layer to extract state features through multiple fully connected layers; usually, an activation function, such as ReLU, is connected after each fully connected layer.
[0107] Convolutional layer: If the state information contains image or grid data, a convolutional layer can be used to extract local features.
[0108] LSTM layer: Use an LSTM layer to process sequential data and capture temporal dependencies.
[0109] Output layer:
[0110] Action output: The output layer generates the action probability distribution or deterministic action of the agent. For a discrete action space, the output layer usually uses the softmax function to convert the output into a probability distribution; for a continuous action space, the output layer can use the tanh or sigmoid function to limit the output within a specific range.
[0111] Value output: The output layer can also generate the state value or action value for evaluating the expected return (reward) of the current state or action.
[0112] S2: Enable the agent to perform end-to-end policy learning in the target network environment, learn how to explore and exploit the vulnerabilities of the target host through interaction with the original training environment, and at the same time, the agent uses a reinforcement learning algorithm to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment;
[0113] Specifically, the agent performs policy learning in the target network environment, which is a real and unknown network environment. The agent needs to learn how to explore and exploit the vulnerabilities of the target host through interaction with the environment. This process is end-to-end, meaning that the agent directly learns from the original feedback of the environment without the need for manually labeled data or a predefined environment model.
[0114] S3: During the interaction with the target network environment, the agent constructs an original simulation environment using the collected host configuration data. The original simulation environment is a digital mapping of the original training environment.
[0115] Specifically, based on the host configuration data collected in the target network environment, the agent constructs a simulation environment in JSON format. This simulation environment is a digital mapping of the original training environment, which can truly reflect the characteristics of the original environment, enabling the agent's interaction in the simulation environment to be equivalent to that in the real environment. This simulation environment is not only used for efficient policy training and verification but also serves as an example of the real-world environment for subsequent environment enhancement.
[0116] S4: Based on the original simulation environment and the vulnerability description, the agent randomizes the host configuration parameters and vulnerability description of the original simulation environment using domain randomization techniques and generates diverse synthetic environments.
[0117] In this step, the original simulation environment and the known vulnerability description are used as inputs to a pre-trained large language model. The large language model generates diverse environment configurations and vulnerability descriptions and outputs synthetic environments.
[0118] Specifically, a large language model (LLM) is used to generate synthetic environments. The LLM generates variants of the original simulation environment based on the official vulnerability description and the original simulation environment. In these variants, for example, a web application (such as Drupal) may be exposed on a non-default port, and there will also be random changes in the operating system version, Apache HTTP server version, and Drupal version, etc. These changes achieve domain randomization, making the synthetic environment better replicate the diversity of the real world, thus preventing the agent from relying on fixed host configuration details.
[0119] S5: In the synthetic environment, the agent uses the proximal policy optimization algorithm and the model-agnostic meta-learning algorithm for in-depth training to optimize the agent's policy parameters and initial parameters, enabling the agent to extract generalizable policies and biases from diverse meta-training environments to generalize the learned policies to new or similar network environments.
[0120] In this step, the meta-learning algorithm includes inner-loop training and outer-loop training.
[0121] In inner-loop training, the agent samples trajectories in each synthetic environment and updates the policy parameters.
[0122] The calculation formula for the inner-loop training is:
[0123]
[0124] where ϕ is the parameter of the meta-policy; ϕ i′ The adapted policy parameters under task Mi; α is the learning rate of the inner loop; is the estimate of the expected discounted reward under task M i ; is the estimate of the expected discounted reward under task M i ; D i is the trajectory data collected under task Mi;
[0125] In the outer loop training, the agent updates the initial parameters and optimizes the initial parameters of the model, enabling the agent to quickly adapt to new tasks.
[0126] The calculation formula for the outer loop training is:
[0127]
[0128] where ϕ is the parameter of the meta-policy; β is the learning rate of the outer loop; is the task distribution, including all tasks used for meta-training; is the policy adapted under task M i ; D i ′ is the trajectory data collected under task M i using the adapted policy ; represents the estimate of the expected discounted reward under task using the adapted policy ; represents the meta-policy gradient estimate, i.e., the sum of the policy gradient estimates over all tasks M i .
[0129] Specifically, based on the synthetic environment set generated by the LLM, model-agnostic meta-learning (MAML) algorithm is used for meta-reinforcement learning training. The MAML algorithm optimizes the initial parameters of the model through the learning processes of the inner loop and the outer loop, enabling the agent to quickly adapt to new tasks. In the inner loop, the agent samples trajectories in each synthetic environment using the initial policy and updates the policy parameters using the policy gradient algorithm. In the outer loop, the agent samples trajectories again using the adapted policy and updates the initial parameters. This process enables the agent to extract generalizable policies and biases from diverse meta-training environments, thus effectively generalizing the learned policies to new and similar test environments.
[0130] S6: Test the generalization ability of the trained agent in the test environment;
[0131] In this step, testing the generalization ability of the trained agent in the test environment includes:
[0132] The agent conducts zero-shot policy transfer testing in a test environment similar to the training environment and evaluates the generalization gap. The formula for calculating the generalization gap is as follows:
[0133]
[0134] where is the generalization gap; π is the policy, which is a function representing the probability distribution that maps the state s of the environment to the action a, i.e., π(a∣s) represents the probability of taking action a in state s; τ is the trajectory under the policy π, i.e., a sequence of states, actions, and rewards; P π is the trajectory distribution under the policy π; G(τ) is the expected cumulative reward of the trajectory τ, , where r t is the reward at time step t, γ is the discount factor used to reduce the weight of future rewards, and T is the total number of time steps; is the simulated environment or the synthetic environment generated through domain randomization; is the test environment, which is the target environment where the agent conducts zero-shot policy transfer; 、 respectively represent the expected cumulative rewards of the trajectory τ generated according to the policy π in the environments 、 , which are the average performance metrics of the policy π in the environments 、 .
[0135] Testing the generalization ability of the trained agent in the test environment also includes:
[0136] The agent conducts fast policy adaptation testing in a test environment that is not similar to the training environment, and improves the agent's performance through a small amount of fine-tuning. The formula for calculating the adapted policy parameters is as follows:
[0137]
[0138] where ϕ is the parameter of the initial policy; ϕ adapted is the parameter of the adapted policy; α is the learning rate; π ϕ represents the policy with parameter ϕ; D test is the trajectory data collected under the test environment M test ; is the estimate of the expected discounted reward under the test environment M test ; is the policy gradient estimate under the test environment M test .
[0139] Specifically, tests are conducted in a test environment that is similar or dissimilar to the training environment to verify the zero-shot policy transfer ability and fast policy adaptation ability of the agent. And metrics such as learning curves, training time, average success rate, and generalization gap are used to evaluate the performance of the agent.
[0140] Through the above training and testing processes, the agent can autonomously learn and adapt to different penetration testing environments, improving its policy learning and generalization abilities in unknown real environments. The agent obtains autonomous recognition of vulnerabilities by interacting with the environment in a real network environment to conduct penetration testing on the vulnerabilities.
[0141] S7: Apply the agent that has completed training and passed the test to a real target network environment for vulnerability identification. The agent automatically identifies network vulnerabilities based on the actions performed and the rewards received in the network environment;
[0142] S8: Select corresponding repair strategies based on the identified network vulnerabilities for repair.
[0143] In this step, analyze the network vulnerabilities identified by the agent to determine the specific types and scopes of influence of the vulnerabilities, analyze their causes and potential repair methods. According to the vulnerability analysis results, generate specific repair steps and commands, and complete the vulnerability repair by updating software, patching, configuration changes, etc. The process of vulnerability repair can be directly completed by the agent or by using existing advanced automated repair tools.
[0144] In this method, the agent realizes penetration testing of network vulnerabilities through technologies such as interaction with the network environment, policy learning, environment simulation, domain randomization, and meta-reinforcement learning. The agent can autonomously learn and adapt to different penetration testing environments, improving its policy learning and generalization abilities in unknown real environments. The trained agent can achieve zero-shot policy transfer in similar environments and fast policy adaptation in dissimilar environments, thus effectively conducting penetration testing of network vulnerabilities and then realizing intelligent repair of network vulnerabilities.
[0145] Embodiment 2
[0146] The embodiment of the present invention also proposes a network vulnerability intelligent repair system based on penetration testing technology, including:
[0147] An agent construction module for constructing an agent for penetration testing based on a reinforcement learning model;
[0148] A training module for performing:
[0149] Enable the agent to perform end-to-end policy learning in the target network environment, learn how to explore and exploit the vulnerabilities of the target host through interaction with the original training environment, and at the same time the agent uses reinforcement learning algorithms to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment;
[0150] During the interaction between the agent and the target network environment, the agent uses the collected host configuration data to construct an original simulation environment, which is a digital mapping of the original training environment;
[0151] Based on the original simulation environment and vulnerability descriptions, the agent uses domain randomization techniques to randomize the host configuration parameters and vulnerability descriptions of the original simulation environment and generate diverse synthetic environments;
[0152] In the synthetic environment, the agent uses the Proximal Policy Optimization algorithm and model-agnostic meta-learning algorithms for in-depth training to optimize the agent's policy parameters and initial parameters, enabling the agent to extract generalizable policies and biases from diverse meta-training environments to generalize the learned policies to new or similar network environments;
[0153] Test the generalization ability of the trained agent in the test environment;
[0154] A vulnerability identification module for applying the trained and tested agent to a real target network environment for vulnerability identification, where the agent automatically identifies network vulnerabilities based on the actions performed and the rewards received in the network environment;
[0155] A vulnerability repair module for selecting corresponding repair strategies based on the identified network vulnerabilities for repair.
[0156] In summary, a network vulnerability intelligent repair method and system according to the present invention combines domain randomization and meta-reinforcement learning techniques to achieve zero-shot policy transfer and fast policy adaptation, significantly improving the agent's adaptability and performance in new environments, not only enhancing the generalization ability and robustness of the policy, but also enabling the agent to be directly deployed in the target network environment for vulnerability identification and repair, reducing deployment time and costs, and providing an efficient and automated solution for the network security field.
[0157] The various embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
Claims
1. A network vulnerability intelligent repair method based on penetration testing technology, characterized in that: include: S1: Building an agent for penetration testing based on a reinforcement learning model; S2: Enable the agent to perform end-to-end policy learning in the target network environment, learn how to explore and exploit the vulnerabilities of the target host by interacting with the original training environment, and use the reinforcement learning algorithm to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment; S3: During the process of interacting with the target network environment, the agent uses the collected host configuration data to build an original simulation environment, where the original simulation environment is a digital mapping of the original training environment; S4: Based on the original simulation environment and vulnerability description, the agent uses domain randomization technology to randomize the host configuration parameters and vulnerability description of the original simulation environment and generate a diverse synthetic environment; S5: In the synthetic environment, the agent uses a proximal policy optimization algorithm and a model-independent meta-learning algorithm for in-depth training to optimize the agent's policy parameters and initial parameters, so that the agent can extract generalizable policies and deviations from a diverse meta-training environment to generalize its learned policies to new or similar network environments; S6: Test the generalization ability of the trained agent in the test environment; S7: Apply the trained and tested agents to the real target network environment for vulnerability identification. The agents automatically identify network vulnerabilities based on the actions performed in the network environment and the rewards received. S8: Select a corresponding repair strategy based on the identified network vulnerability for repair.
2. The method according to claim 1, characterized in that In step S1, building an agent for penetration testing based on a reinforcement learning model includes: S101: Define a reinforcement learning framework for your agent, including: State space S: the state of the environment observed by the agent, including host configuration, ports, services, operating system, and website fingerprint information; Action space A: actions that the agent can perform, such as information collection, system detection, and vulnerability exploitation; Reward function R: the learning goal of the agent, which is to collect key information from the target host and exploit vulnerabilities while minimizing the cost of actions; S102: Define the agent's strategy learning, including: Policy network: The agent uses a neural network to represent its policy, which maps states to actions. The parameters of the policy network are updated via a reinforcement learning algorithm. Reinforcement Learning Algorithm: The agent uses a proximal policy optimization algorithm to update policy parameters; S103: Define the agent's environmental interactions, including: Observe the environment: The agent uses scanning tools to observe the environment status of the target host and collect information; Decision-making: Based on the information collected, the agent uses its strategy to decide the next action, which includes information gathering, system probing, and vulnerability exploitation; Execute Action: The agent performs the selected action by calling and executing commands in the penetration testing toolkit; Receive rewards: The agent receives rewards based on the results of the actions performed. The reward function gives positive or negative rewards based on whether the vulnerability is successfully exploited and the information collected; S104: Define the policy update of the agent: Experience replay: The agent stores the experience of each interaction in the experience replay pool, which includes state, action, reward, and next state; Policy Optimization: The agent samples experiences from the experience replay pool and uses a proximal policy optimization algorithm to update the parameters of the policy network to maximize the expected cumulative reward.
3. The method according to claim 2, characterized in that The reward function R is: ; Where value(h) is the reward for successfully identifying the vulnerability of target host h, and cost(a) is the cost of action a.
4. The method according to claim 2, characterized in that: The proximal policy optimization algorithm is a reinforcement learning algorithm based on policy gradient, which minimizes the objective function through multi-step stochastic gradient descent to optimize the policy. The objective function of the proximal policy optimization algorithm is: ; Among them, ϕ is the parameter of the policy network; π is the current policy; is the trajectory under strategy π, i.e., a sequence of states, actions, and rewards; is the probability ratio, which represents the probability of taking action at under the current policy ϕ and the probability of taking action at under the old policy ϕ old The ratio of the probabilities of taking the same action under ; is an estimate of the advantage function, which represents the additional benefit of taking action at in state st compared to the average case; Represents the probability ratio Clipping is performed and limited to between 1−ϵ and 1+ϵ to prevent the policy update from being too large. ϵ is a hyperparameter used to control the range of clipping.
5. The method according to claim 1, characterized in that In step S4, the domain randomization technique is used to randomize the host configuration parameters and vulnerability descriptions of the original simulation environment, and a diverse synthetic environment is generated, including: The original simulation environment and known vulnerability descriptions are used as inputs to a pre-trained large language model, and a diverse set of environment configurations and vulnerability descriptions are generated through the large language model, and a synthetic environment is output.
6. The method according to claim 1, characterized in that In step S5, the meta-learning algorithm includes inner loop training and outer loop training; In the inner loop training, the agent samples trajectories in each synthetic environment and updates the policy parameters; In the outer loop training, the agent updates the initial parameters and optimizes the initial parameters of the model, enabling the agent to quickly adapt on new tasks.
7. The method according to claim 6, characterized in that The calculation formula of the inner loop training is: ; Among them, ϕ is the parameter of the meta-strategy; i ′ is the policy parameter after adaptation under task Mi; α is the learning rate of the inner loop; In task M i Estimates of expected discounted rewards under ; In task M i Policy gradient estimation under i is the trajectory data collected under task Mi; The calculation formula of the outer loop training is: ; Among them, ϕ is the parameter of the meta-strategy; β is the learning rate of the outer loop; is the task distribution, which contains all tasks used for meta-training; In task M i The strategy after adaptation; D i ′ is the i Next, use the adapted strategy The collected trajectory data; Indicates that in the task Next, use the adapted strategy An estimate of the expected discounted reward; represents the meta-policy gradient estimate, that is, in all tasks M i The sum of the policy gradient estimates on .
8. The method according to claim 1, characterized in that In step S6, testing the generalization ability of the trained agent in the test environment includes: The agent is tested for zero-shot policy transfer in a test environment similar to the training environment and the generalization gap is evaluated, which is calculated as: ; in, is the generalization gap; π is the strategy, which is a function that represents the probability distribution of mapping the state s of the environment to the action a, that is, π(a|s) represents the probability of taking action a in state s; τ is the trajectory under the strategy π, that is, a sequence of states, actions and rewards; P π is the trajectory distribution under strategy π; G(τ) is the expected cumulative reward of trajectory τ, , where r t is the reward at time step t, γ is the discount factor used to reduce the weight of future rewards, and T is the total number of time steps; It is a simulated environment or a synthetic environment generated by domain randomization; The test environment is the target environment in which the agent performs zero-shot policy transfer; , Respectively in the environment , The expected cumulative reward of the trajectory τ generated by the strategy π under the environment , , the average performance indicator of strategy π.
9. The method according to claim 1, characterized in that: In step S6, testing the generalization ability of the trained agent in the test environment also includes: The agent performs a quick policy adaptation test in a test environment that is dissimilar to the training environment, and improves the agent performance through a small amount of fine-tuning. The calculation formula of the adapted policy parameters is: ; Among them, ϕ is the parameter of the initial strategy; ϕ adapted is the strategy parameter after adaptation; α is the learning rate; π ϕ represents the strategy with parameter φ; D test For the test environment M test The trajectory data collected below; In the test environment M test Estimates of expected discounted rewards under ; In the test environment M test Policy gradient estimation under .
10. A network vulnerability intelligent repair system based on penetration testing technology, characterized in that: include: An agent building module for building agents for penetration testing based on reinforcement learning models; Training module, which performs: The agent performs end-to-end policy learning in the target network environment, learns how to explore and exploit the vulnerabilities of the target host by interacting with the original training environment, and uses a reinforcement learning algorithm to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment; The agent, in the process of interacting with the target network environment, uses the collected host configuration data to build an original simulation environment, wherein the original simulation environment is a digital mapping of the original training environment; Based on the original simulation environment and vulnerability description, the agent uses domain randomization technology to randomize the host configuration parameters and vulnerability description of the original simulation environment and generate a diverse synthetic environment; In the synthetic environment, the agent uses a proximal policy optimization algorithm and a model-independent meta-learning algorithm for in-depth training to optimize the agent's policy parameters and initial parameters, so that the agent can extract generalizable policies and deviations from a diverse meta-training environment to generalize its learned policies to new or similar network environments; Test the generalization ability of the trained agent in the test environment; The vulnerability identification module is used to apply the trained and tested agents to the real target network environment for vulnerability identification. The agents automatically identify network vulnerabilities based on the actions performed in the network environment and the rewards received; The vulnerability repair module is used to select corresponding repair strategies based on the identified network vulnerabilities for repair.
Citation Information
Patent Citations
Automatic Windows domain penetration method based on reinforcement learning
CN114444086A
Persistent feature expression method and system based on domain randomization and meta-learning
CN118015335A
Penetration test attack path discovery method based on PPO deep reinforcement learning model
CN118118202A
Intranet penetration environment method and device based on reinforcement learning and virtualization container
CN118944923A
Energy management method combining meta-learning and near-end strategy optimization algorithm
CN119359080A
Cited By
Vulnerability POC generation method based on AI agent
CN121841858A