Intelligent Network Vulnerability Repair Method and System Based on Penetration Testing Technology

By building a penetration testing agent based on reinforcement learning model, combining domain randomization and meta-learning algorithms, the problem of traditional penetration testing methods performing poorly in the new environment is solved, and efficient vulnerability identification and repair is achieved.

CN120090872BActive Publication Date: 2025-07-04YUNNAN BLUE TEAM CLOUD COMPUTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510539660.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-04
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Traditional penetration testing methods rely on manual operations, are costly and long-term, and are difficult to effectively respond in a new environment. There are realistic gaps, resulting in poor strategy performance.

Method used

By building agents based on reinforcement learning models, training agents using end-to-end strategy learning, domain randomization and meta-learning algorithms, it enables them to generalize policies in real environments and identify vulnerabilities and fix them.

Benefits of technology

It realizes zero-sample policy transfer and fast policy adaptation, improves vulnerability identification and repair efficiency, and enhances the security and reliability of the network system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090872B_ABST
    Figure CN120090872B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of network security technology, and discloses an intelligent network vulnerability repair method and system based on penetration testing technology. The method includes: constructing an agent for penetration testing; enabling the agent to perform end-to-end policy learning in the target network environment; during the interaction between the agent and the target network environment, using the collected host configuration data to construct an original simulation environment; using domain randomization technology to randomize the host configuration parameters and vulnerability descriptions of the original simulation environment to generate diverse synthetic environments; in the synthetic environment, the agent uses the proximal policy optimization algorithm and the model-agnostic meta-learning algorithm for in-depth training; testing the generalization ability of the trained agent in the test environment; applying the agent to the real target network environment for vulnerability identification; and selecting corresponding repair strategies based on the identified network vulnerabilities for repair. The present invention can effectively improve the efficiency of vulnerability identification and repair.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and more specifically, to a network vulnerability intelligent repair method and system based on penetration testing technology. Background Art

[0002] With the rapid development of network technology, network security issues have become increasingly prominent. Network security vulnerabilities are usually identified and repaired through penetration testing. Traditional penetration testing methods mainly rely on manual operations, which are costly, time-consuming, and error-prone. Manual testing requires professional security personnel, and has high requirements for the technical level and experience of personnel. In addition, the strategies trained in the simulation environment by traditional methods are often difficult to be directly applied to the real environment, there is a "reality gap", resulting in poor performance of the strategies in the new environment. These limitations make it difficult for traditional penetration testing methods to effectively cope with large-scale and dynamically changing network environments. Summary of the Invention

[0003] The object of the present invention is to propose a network vulnerability intelligent repair method and system based on penetration testing technology, and train an autonomous penetration testing agent that can be generalized to unseen real environments through domain randomization and meta-reinforcement learning, improve the adaptability and performance of the agent in the new environment, so as to improve the efficiency of vulnerability identification and repair.

[0004] To achieve the above object, the present invention proposes a network vulnerability intelligent repair method based on penetration testing technology, including:

[0005] S1: Construct an agent for penetration testing based on a reinforcement learning model;

[0006] S2: Make the agent perform end-to-end policy learning in the target network environment, learn how to explore and exploit the vulnerabilities of the target host through interaction with the original training environment, and at the same time the agent uses a reinforcement learning algorithm to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment;

[0007] S3: During the interaction between the agent and the target network environment, use the collected host configuration data to construct an original simulation environment, and the original simulation environment is a digital mapping of the original training environment;

[0008] S4: Based on the original simulation environment and the vulnerability description, the agent uses domain randomization technology to randomize the host configuration parameters and the vulnerability description of the original simulation environment, and generate diverse synthetic environments;

[0009] S5: In the synthesis environment, the agent uses the Proximal Policy Optimization algorithm and the model - agnostic meta - learning algorithm for in - depth training to optimize the agent's policy parameters and initial parameters, enabling the agent to extract generalizable policies and biases from diverse meta - training environments to generalize the learned policies to new or similar network environments;

[0010] S6: Test the generalization ability of the trained agent in the test environment;

[0011] S7: Apply the agent that has completed training and passed the test to the real target network environment for vulnerability identification. The agent automatically identifies network vulnerabilities based on the actions performed and the rewards received in the network environment;

[0012] S8: Select the corresponding repair strategy based on the identified network vulnerabilities for repair.

[0013] Optionally, in step S1, constructing an agent for penetration testing based on a reinforcement learning model includes:

[0014] S101: Define the reinforcement learning framework of the agent, including:

[0015] State space S: The environmental state observed by the agent, including host configuration, ports, services, operating system, and website fingerprint information;

[0016] Action space A: The actions that the agent can perform, such as information collection, system detection, and vulnerability exploitation;

[0017] Reward function R: The learning objective of the agent, that is, while minimizing the action cost, collecting key information from the target host and exploiting vulnerabilities;

[0018] S102: Define the policy learning of the agent, including:

[0019] Policy network: The agent uses a neural network to represent its policy. This network maps the state to actions, and the parameters of the policy network are updated through the reinforcement learning algorithm;

[0020] Reinforcement learning algorithm: The agent uses the Proximal Policy Optimization algorithm to update the policy parameters;

[0021] S103: Define the environmental interaction of the agent, including:

[0022] Observe the environment: The agent uses scanning tools to observe the environmental state of the target host and collect information;

[0023] Decision - making: Based on the collected information, the agent uses its policy to decide the next actions, which include information collection, system detection, and vulnerability exploitation;

[0024] Execute Action: The agent executes the selected action by calling and executing commands in the penetration testing toolkit;

[0025] Receive Reward: The agent receives a reward based on the result of the executed action. The reward function gives a positive or negative reward according to whether the vulnerability is successfully exploited and the information collected;

[0026] S103: Define the policy update of the agent:

[0027] Experience Replay: The agent stores the experience of each interaction in the experience replay pool. The experience includes state, action, reward, and next state;

[0028] Policy Optimization: The agent samples experiences from the experience replay pool and uses the Proximal Policy Optimization algorithm to update the parameters of the policy network to maximize the expected cumulative reward.

[0029] Optionally, the reward function R is:

[0030]

[0031] where value(h) is the reward for successfully identifying the vulnerability of the target host h, and cost(a) is the cost of action a.

[0032] Optionally, the Proximal Policy Optimization algorithm is a reinforcement learning algorithm based on policy gradients. It minimizes the objective function through multi-step stochastic gradient descent to optimize the policy. The objective function of the Proximal Policy Optimization algorithm is:

[0033]

[0034] where ϕ are the parameters of the policy network; π is the current policy; is the trajectory under the policy π, that is, a sequence of states, actions, and rewards; is the probability ratio, which represents the ratio of the probability of taking action at under the current policy ϕ to the probability of taking the same action under the old policy ϕ old That is ; is an estimate of the advantage function, which represents the additional benefit of taking action at in state st compared to the average case; represents the probability ratio is clipped to be between 1−ϵ and 1+ϵ to prevent the policy update from being too large. ϵ is a hyperparameter used to control the clipping range.

[0035] Optionally, in step S4, the domain randomization technique is used to randomize the host configuration parameters and vulnerability descriptions of the original simulation environment, and a diverse synthetic environment is generated, including:

[0036] Use the original simulation environment and known vulnerability descriptions as the input to a pre-trained large language model. Generate diverse environment configurations and vulnerability descriptions through the large language model, and output a synthetic environment.

[0037] Optionally, in step S5, the meta-learning algorithm includes inner-loop training and outer-loop training;

[0038] In inner-loop training, the agent samples trajectories in each synthetic environment and updates the policy parameters;

[0039] In outer-loop training, the agent updates the initial parameters, optimizing the initial parameters of the model so that the agent can quickly adapt to new tasks.

[0040] Optionally, the calculation formula for the inner-loop training is:

[0041]

[0042] where ϕ is the parameter of the meta-policy; ϕ i ′ is the adapted policy parameter under task Mi; α is the learning rate of the inner loop; is the estimate of the expected discounted reward under task M i ; is the policy gradient estimate under task M i ; D i is the trajectory data collected under task Mi;

[0043] The calculation formula for the outer-loop training is:

[0044]

[0045] where ϕ is the parameter of the meta-policy; β is the learning rate of the outer loop; is the task distribution, containing all tasks used for meta-training; is the adapted policy under task M i ; D i ′ is the trajectory data collected under task M i using the adapted policy ; represents the estimate of the expected discounted reward under task using the adapted policy ; represents the meta-policy gradient estimate, i.e., the sum of the policy gradient estimates over all tasks M i .

[0046] Optionally, in step S6, testing the generalization ability of the trained agent in a test environment includes:

[0047] The agent conducts a zero-shot policy transfer test in a test environment similar to the training environment and evaluates the generalization gap. The formula for calculating the generalization gap is:

[0048]

[0049] Where, is the generalization gap; π is the policy, which is a function representing the probability distribution that maps the state s of the environment to the action a, that is, π(a∣s) represents the probability of taking action a in state s; τ is the trajectory under the policy π, that is, a sequence of states, actions, and rewards; P π is the trajectory distribution under the policy π; G(τ) is the expected cumulative reward of the trajectory τ, , where r t is the reward at time step t, γ is the discount factor used to reduce the weight of future rewards, and T is the total number of time steps; is the simulated environment or the synthetic environment generated through domain randomization; is the test environment, which is the target environment where the agent conducts zero-shot policy transfer; , respectively represent the expected cumulative rewards of the trajectory τ generated according to the policy π in the environments , , which is the average performance metric of the policy π in the environments , .

[0050] Optionally, in step S6, testing the generalization ability of the trained agent in a test environment also includes:

[0051] The agent conducts a fast policy adaptation test in a test environment not similar to the training environment, and improves the agent's performance through a small amount of fine-tuning. The formula for calculating the adapted policy parameters is:

[0052]

[0053] Where, ϕ is the parameter of the initial policy; ϕ adapted is the parameter of the adapted policy; α is the learning rate; π ϕ represents the policy with parameter ϕ; D test is the trajectory data collected under the test environment M test ; is the estimate of the expected discounted reward under the test environment M test ; It is the policy gradient estimation in the test environment M test underneath.

[0054] The present invention also proposes an intelligent network vulnerability repair system based on penetration testing technology, including:

[0055] An agent construction module for constructing an agent for penetration testing based on a reinforcement learning model;

[0056] A training module for performing:

[0057] Enabling the agent to perform end-to-end policy learning in the target network environment, learning how to explore and exploit the vulnerabilities of the target host through interaction with the original training environment. Meanwhile, the agent uses a reinforcement learning algorithm to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment;

[0058] During the process of the agent interacting with the target network environment, the agent constructs an original simulation environment using the collected host configuration data, and the original simulation environment is a digital mapping of the original training environment;

[0059] Based on the original simulation environment and vulnerability description, the agent uses domain randomization technology to randomize the host configuration parameters and vulnerability description of the original simulation environment and generate diverse synthetic environments;

[0060] In the synthetic environment, the agent uses the proximal policy optimization algorithm and model-agnostic meta-learning algorithm for in-depth training to optimize the agent's policy parameters and initial parameters, enabling the agent to extract generalizable policies and biases from diverse meta-training environments to generalize the learned policies to new or similar network environments;

[0061] Testing the generalization ability of the trained agent in the test environment;

[0062] A vulnerability identification module for applying the trained and tested agent to the real target network environment for vulnerability identification, and the agent automatically identifies network vulnerabilities according to the actions performed in the network environment and the received rewards;

[0063] A vulnerability repair module for selecting corresponding repair strategies for repair based on the identified network vulnerabilities.

[0064] The beneficial effects of the present invention are as follows:

[0065] The present invention trains an agent by combining domain randomization and meta-reinforcement learning techniques, achieving zero-shot policy transfer and fast policy adaptation, significantly improving the agent's policy learning and adaptation capabilities in unknown environments. By enabling the agent to automatically identify network vulnerabilities and automatically repair them, the efficiency of penetration testing and vulnerability repair is effectively improved, enhancing the security and reliability of the network system, and providing an efficient and automated solution for network security testing and intelligent vulnerability repair.

[0066] The system of the present invention has other characteristics and advantages that will be apparent from or will be described in detail in the accompanying drawings and the subsequent detailed description incorporated herein. These accompanying drawings and detailed description are used together to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] By describing the exemplary embodiments of the present invention in more detail with reference to the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more apparent. In the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.

[0068] Figure 1 The flowchart showing the steps of a network vulnerability intelligent repair method based on penetration testing technology according to an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] The present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. Instead, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.

[0070] As Figure 1 shown, an embodiment of the present invention provides a network vulnerability intelligent repair method based on penetration testing technology, including:

[0071] S1: Construct an agent for penetration testing based on a reinforcement learning model;

[0072] This step specifically includes:

[0073] S101: Define the reinforcement learning framework of the agent (model the penetration testing process as a Markov decision process), including:

[0074] State space S: The environmental state observed by the agent, including host configuration, ports, services, operating system, and website fingerprint information;

[0075] Action space A: The actions that the agent can perform, such as information collection, system detection, and vulnerability exploitation;

[0076] Reward function R: The learning objective of the agent, that is, while minimizing the action cost, collect key information from the target host and exploit vulnerabilities; the reward function R is as follows:

[0077]

[0078] Among them, value(h) is the reward for successfully identifying the vulnerability of the target host h, and cost(a) is the cost of action a.

[0079] S102: Define the policy learning of the agent, including:

[0080] Policy network: The agent uses a neural network to represent its policy, which maps states to actions, and the parameters of the policy network are updated through a reinforcement learning algorithm;

[0081] Reinforcement learning algorithm: The agent uses the Proximal Policy Optimization algorithm to update the policy parameters; the Proximal Policy Optimization algorithm is a policy gradient-based reinforcement learning algorithm, which minimizes the objective function through multi-step stochastic gradient descent to optimize the policy. The objective function of the Proximal Policy Optimization algorithm is as follows:

[0082]

[0083] Among them, ϕ is the parameter of the policy network; π is the current policy; is the trajectory under the policy π, that is, a sequence of states, actions, and rewards; is the probability ratio, indicating the ratio of the probability of taking action at under the current policy ϕ to the probability of taking the same action under the old policy ϕ old That is ; is the estimate of the advantage function, indicating the additional benefit of taking action at compared to the average situation in state st; represents clip the probability ratio

[0084] S103: Define the environmental interaction of the agent, including:

[0085] Observe the environment: The agent uses a scanning tool to observe the environmental status of the target host and collect information;

[0086] Decision-making: Based on the collected information, the agent uses its policy to decide the next action, and these actions include information collection, system detection, and vulnerability exploitation;

[0087] Execute actions: The agent executes the selected actions by calling and executing commands in the penetration testing toolkit;

[0088] Receive rewards: The agent receives rewards based on the results of the executed actions. The reward function gives positive or negative rewards according to whether the vulnerability is successfully exploited and the information collected;

[0089] S104: Define the policy update of the agent:

[0090] Experience replay: The agent stores the experience of each interaction in the experience replay pool. The experience includes the state, action, reward, and next state;

[0091] Policy optimization: The agent samples experiences from the experience replay pool and uses the Proximal Policy Optimization algorithm to update the parameters of the policy network to maximize the expected cumulative reward.

[0092] Furthermore, the functions implemented by the agent mainly include:

[0093] (1) Information collection:

[0094] Scanning tools: The agent uses various scanning tools (such as Nmap, whatweb, dirb, etc.) to collect detailed information about the target host, including open ports, running services, operating system type, website fingerprints, etc.

[0095] (2) Information analysis: The agent uses the pre-trained Sentence-BERT model to embed the collected text information into the state vector so that the neural network can understand and process this information.

[0096] (3) Vulnerability exploitation:

[0097] Select vulnerabilities: Based on the information collected, the agent selects appropriate vulnerability exploitation tools or payloads to attack the target host.

[0098] Execute attacks: The agent executes the selected vulnerability exploitation actions by calling and executing commands in the penetration testing toolkit (such as MSF) and tries to compromise the target host.

[0099] (4) Policy optimization:

[0100] Reward function: The goal of the agent is to maximize the expected cumulative reward. The reward function defines the positive rewards obtained by the agent when successfully exploiting vulnerabilities and collecting key information, as well as the negative rewards received when executing invalid or wrong actions.

[0101] (5) Policy update: The agent uses the Proximal Policy Optimization algorithm to update its policy parameters and minimizes the objective function through multi-step Stochastic Gradient Descent (SGD) to optimize the policy.

[0102] This embodiment does not limit the specific network structure of the agent. The agent can adopt any suitable reinforcement learning AI model or an agent integrating an AI model. In one example, the network structure of the agent can be a deep neural network, and its specific structure includes the following parts:

[0103] Input layer:

[0104] State input: The input layer receives the state information of the environment, such as host configuration, ports, services, operating systems, website fingerprints, etc. These state information are usually encoded as high-dimensional vectors.

[0105] Hidden layer:

[0106] Multi-Layer Perceptron (MLP): Use a multi-layer perceptron as the hidden layer to extract state features through multiple fully connected layers; usually, an activation function, such as ReLU, is connected after each fully connected layer.

[0107] Convolutional layer: If the state information contains image or grid data, a convolutional layer can be used to extract local features.

[0108] LSTM layer: Use an LSTM layer to process sequential data and capture temporal dependencies.

[0109] Output layer:

[0110] Action output: The output layer generates the action probability distribution or deterministic action of the agent. For a discrete action space, the output layer usually uses the softmax function to convert the output into a probability distribution; for a continuous action space, the output layer can use the tanh or sigmoid function to limit the output within a specific range.

[0111] Value output: The output layer can also generate the state value or action value for evaluating the expected return (reward) of the current state or action.

[0112] S2: Enable the agent to perform end-to-end policy learning in the target network environment. By interacting with the original training environment, the agent learns how to explore and exploit the vulnerabilities of the target host. Meanwhile, the agent uses a reinforcement learning algorithm to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment;

[0113] Specifically, the agent performs policy learning in the target network environment, which is a real and unknown network environment. The agent needs to learn how to explore and exploit the vulnerabilities of the target host through interaction with the environment. This process is end-to-end, meaning that the agent directly learns from the original feedback of the environment without the need for manually labeled data or a predefined environment model.

[0114] S3: During the interaction with the target network environment, the agent constructs an original simulation environment using the collected host configuration data, and the original simulation environment is a digital mapping of the original training environment;

[0115] Specifically, based on the host configuration data collected in the target network environment, the agent constructs a simulation environment in JSON format. This simulation environment is a digital mapping of the original training environment, which can truly reflect the characteristics of the original environment, enabling the agent's interaction in the simulation environment to be equivalent to that in the real environment. This simulation environment is not only used for efficient policy training and verification but also serves as an example of the real-world environment for subsequent environment enhancement.

[0116] S4: Based on the original simulation environment and the vulnerability description, the agent randomizes the host configuration parameters and vulnerability description of the original simulation environment using domain randomization techniques and generates diverse synthetic environments;

[0117] In this step, the original simulation environment and the known vulnerability description are used as inputs to a pre-trained large language model. The large language model generates diverse environment configurations and vulnerability descriptions and outputs synthetic environments.

[0118] Specifically, a large language model (LLM) is used to generate synthetic environments. The LLM generates variants of the original simulation environment based on the official vulnerability description and the original simulation environment. In these variants, for example, a web application (such as Drupal) may be exposed on a non-default port, and there will also be random changes in the operating system version, Apache HTTP server version, and Drupal version, etc. These changes achieve domain randomization, making the synthetic environment better replicate the diversity of the real world and thus preventing the agent from relying on fixed host configuration details.

[0119] S5: In the synthetic environment, the agent uses the proximal policy optimization algorithm and the model-agnostic meta-learning algorithm for in-depth training to optimize the agent's policy parameters and initial parameters, enabling the agent to extract generalizable policies and biases from diverse meta-training environments to generalize the policies it has learned to new or similar network environments;

[0120] In this step, the meta-learning algorithm includes inner-loop training and outer-loop training;

[0121] In inner-loop training, the agent samples trajectories in each synthetic environment and updates the policy parameters;

[0122] The calculation formula for the inner-loop training is:

[0123]

[0124] where ϕ is the parameter of the meta-policy; ϕ i′ The adapted policy parameters under task Mi; α is the learning rate of the inner loop; is the estimate of the expected discounted reward under task M i ; is the estimate of the expected discounted reward under task M i ; D i is the trajectory data collected under task Mi;

[0125] In the outer loop training, the agent updates the initial parameters and optimizes the initial parameters of the model, enabling the agent to quickly adapt to new tasks.

[0126] The calculation formula for the outer loop training is:

[0127]

[0128] where ϕ is the parameter of the meta-policy; β is the learning rate of the outer loop; is the task distribution, including all tasks used for meta-training; is the adapted policy under task M i ; D i ′ is the trajectory data collected under task M i using the adapted policy ; represents the estimate of the expected discounted reward under task using the adapted policy ; represents the meta-policy gradient estimate, i.e., the sum of the policy gradient estimates over all tasks M i ;

[0129] Specifically, based on the synthetic environment set generated by the LLM, model-agnostic meta-learning (MAML) algorithm is used for meta-reinforcement learning training. The MAML algorithm optimizes the initial parameters of the model through the learning processes of the inner loop and the outer loop, enabling the agent to quickly adapt to new tasks. In the inner loop, the agent samples trajectories in each synthetic environment using the initial policy and updates the policy parameters using the policy gradient algorithm. In the outer loop, the agent samples trajectories again using the adapted policy and updates the initial parameters. This process enables the agent to extract generalizable policies and biases from diverse meta-training environments, thus effectively generalizing the learned policies to new and similar test environments.

[0130] S6: Test the generalization ability of the trained agent in the test environment;

[0131] In this step, testing the generalization ability of the trained agent in the test environment includes:

[0132] The agent conducts zero-shot policy transfer testing in a test environment similar to the training environment and evaluates the generalization gap. The formula for calculating the generalization gap is:

[0133]

[0134] where is the generalization gap; π is the policy, which is a function representing the probability distribution that maps the state s of the environment to the action a, i.e., π(a∣s) represents the probability of taking action a in state s; τ is the trajectory under the policy π, i.e., a sequence of states, actions, and rewards; P π is the trajectory distribution under the policy π; G(τ) is the expected cumulative reward of the trajectory τ, , where r t is the reward at time step t, γ is the discount factor used to reduce the weight of future rewards, and T is the total number of time steps; is the simulated environment or the synthetic environment generated through domain randomization; is the test environment, which is the target environment where the agent conducts zero-shot policy transfer; 、 respectively represent the expected cumulative rewards of the trajectory τ generated according to the policy π in the environments 、 , which are the average performance metrics of the policy π in the environments 、 .

[0135] Testing the generalization ability of the trained agent in the test environment also includes:

[0136] The agent conducts fast policy adaptation testing in a test environment dissimilar to the training environment, and improves the agent's performance through a small amount of fine-tuning. The formula for calculating the adapted policy parameters is:

[0137]

[0138] where ϕ is the parameter of the initial policy; ϕ adapted is the parameter of the adapted policy; α is the learning rate; π ϕ represents the policy with parameter ϕ; D test is the trajectory data collected under the test environment M test ; is the estimate of the expected discounted reward under the test environment M test ; is the policy gradient estimate under the test environment M test .

[0139] Specifically, tests are conducted in a test environment that is similar or dissimilar to the training environment to verify the zero-shot policy transfer ability and fast policy adaptation ability of the agent. And metrics such as learning curves, training time, average success rate, and generalization gap are used to evaluate the performance of the agent.

[0140] Through the above training and testing processes, the agent can autonomously learn and adapt to different penetration testing environments, improve its policy learning and generalization abilities in unknown real environments, and the agent completes the autonomous identification of vulnerabilities through interaction with the environment in the real network environment for penetration testing of vulnerabilities.

[0141] S7: Apply the agent that has completed training and passed the test to the real target network environment for vulnerability identification. The agent automatically identifies network vulnerabilities based on the actions performed and the rewards received in the network environment;

[0142] S8: Select the corresponding repair strategy based on the identified network vulnerabilities for repair.

[0143] In this step, analyze the network vulnerabilities identified by the agent to determine the specific types and scopes of influence of the vulnerabilities, analyze their causes and potential repair methods, and generate specific repair steps and commands based on the vulnerability analysis results. Complete vulnerability repair by updating software, patching, configuration changes, etc.; the process of vulnerability repair can be directly completed by the agent or by using existing advanced automated repair tools.

[0144] In this method, the agent realizes the penetration testing of network vulnerabilities through technologies such as interaction with the network environment, policy learning, environment simulation, domain randomization, and meta-reinforcement learning. The agent can autonomously learn and adapt to different penetration testing environments, improve its policy learning and generalization abilities in unknown real environments. The trained agent can achieve zero-shot policy transfer in similar environments and fast policy adaptation in dissimilar environments, thus effectively conducting penetration testing of network vulnerabilities and then realizing the intelligent repair of network vulnerabilities.

[0145] Embodiment 2

[0146] The embodiment of the present invention also proposes a network vulnerability intelligent repair system based on penetration testing technology, including:

[0147] An agent construction module for constructing an agent for penetration testing based on a reinforcement learning model;

[0148] A training module for performing:

[0149] Enable the agent to perform end-to-end policy learning in the target network environment, learn how to explore and exploit the vulnerabilities of the target host through interaction with the original training environment, and at the same time the agent uses reinforcement learning algorithms to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment;

[0150] During the interaction between the agent and the target network environment, the agent uses the collected host configuration data to construct an original simulation environment, which is a digital mapping of the original training environment;

[0151] Based on the original simulation environment and vulnerability descriptions, the agent uses domain randomization techniques to randomize the host configuration parameters and vulnerability descriptions of the original simulation environment and generate diverse synthetic environments;

[0152] In the synthetic environment, the agent uses the Proximal Policy Optimization algorithm and model-agnostic meta-learning algorithms for in-depth training to optimize the agent's policy parameters and initial parameters, enabling the agent to extract generalizable policies and biases from diverse meta-training environments to generalize the learned policies to new or similar network environments;

[0153] Test the generalization ability of the trained agent in the test environment;

[0154] A vulnerability identification module for applying the trained and tested agent to a real target network environment for vulnerability identification. The agent automatically identifies network vulnerabilities based on the actions performed and the rewards received in the network environment;

[0155] A vulnerability repair module for selecting corresponding repair strategies for repair based on the identified network vulnerabilities.

[0156] In summary, a network vulnerability intelligent repair method and system of the present invention combines domain randomization and meta-reinforcement learning techniques to achieve zero-shot policy transfer and fast policy adaptation, significantly improving the agent's adaptability and performance in new environments, not only enhancing the generalization ability and robustness of the policy, but also enabling the agent to be directly deployed in the target network environment for vulnerability identification and repair, reducing deployment time and cost, and providing an efficient and automated solution for the network security field.

[0157] The various embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments.

Claims

1. An intelligent network vulnerability repair method based on penetration testing technology, characterized in that, Including: S1: Construct an agent for penetration testing based on a reinforcement learning model; S2: Enable the agent to perform end-to-end policy learning in the target network environment, learn how to explore and exploit vulnerabilities of target hosts through interaction with the original training environment, and at the same time the agent uses a reinforcement learning algorithm to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment; S3: During the interaction between the agent and the target network environment, the agent constructs an original simulation environment using the collected host configuration data, and the original simulation environment is a digital mapping of the original training environment; S4: Based on the original simulation environment and vulnerability descriptions, the agent uses domain randomization techniques to randomize the host configuration parameters and vulnerability descriptions of the original simulation environment and generate diverse synthetic environments; S5: In the synthetic environment, the agent uses the proximal policy optimization algorithm and model-agnostic meta-learning algorithm for in-depth training to optimize the agent's policy parameters and initial parameters, enabling the agent to extract generalizable policies and biases from diverse meta-training environments to generalize the learned policies to new or similar network environments; S6: Test the generalization ability of the trained agent in a test environment; S7: Apply the trained and tested agent to a real target network environment for vulnerability identification, and the agent automatically identifies network vulnerabilities based on the actions performed and the rewards received in the network environment; S8: Select corresponding repair strategies for repair based on the identified network vulnerabilities.

2. The method according to claim 1, characterized in that, In step S1, constructing an agent for penetration testing based on a reinforcement learning model includes: S101: Define the reinforcement learning framework of the agent, including: State space S: The environmental state observed by the agent, including host configuration, ports, services, operating systems, and website fingerprint information; Action space A: The actions that the agent can perform, including information collection, system detection, and vulnerability exploitation; Reward function R: The learning objective of the agent, that is, while minimizing the action cost, collect key information from the target host and exploit vulnerabilities; S102: Define the policy learning of the agent, including: Policy network: The agent uses a neural network to represent its policy, which maps states to actions, and the parameters of the policy network are updated through a reinforcement learning algorithm; Reinforcement learning algorithm: The agent uses the proximal policy optimization algorithm to update the policy parameters; S103: Define the environmental interaction of the agent, including: Observe the environment: The agent uses scanning tools to observe the environmental state of the target host and collect information; Decision-making: Based on the collected information, the agent uses its policy to decide the next actions, which include information collection, system detection, and vulnerability exploitation; Execute actions: The agent executes the selected actions by calling and executing commands in the penetration testing toolkit; Receive rewards: The agent receives rewards according to the results of the executed actions, and the reward function gives positive or negative rewards based on whether the vulnerability is successfully exploited and the information collected; S104: Define the policy update of the agent: Experience replay: The agent stores the experience of each interaction in the experience replay pool, and the experience includes states, actions, rewards, and next states; Policy Optimization: The agent samples experiences from the experience replay pool and uses the Proximal Policy Optimization (PPO) algorithm to update the parameters of the policy network to maximize the expected cumulative reward.

3. The method according to claim 2, wherein The reward function R is as follows: ; where value(h) is the reward for successfully identifying the vulnerabilities of target host h, and cost(a) is the cost of action a.

4. The method according to claim 2, wherein The Proximal Policy Optimization algorithm is a policy gradient-based reinforcement learning algorithm that optimizes the policy by minimizing the objective function through multi-step stochastic gradient descent. The objective function of the Proximal Policy Optimization algorithm is: ; where $\phi$ are the parameters of the policy network; $\pi$ is the current policy; $\tau$ is a trajectory under policy $\pi$, i.e., a sequence of states, actions, and rewards; $\frac{\pi_{\phi}(a_t|s_t)}{\pi_{\phi'}(a_t|s_t)}$ is the probability ratio, representing the ratio of the probability of taking action $a_t$ under the current policy $\phi$ to the probability of taking the same action under the old policy $\phi'$, i.e., old ; ; $A^{\pi}(s_t,a_t)$ is an estimate of the advantage function, representing the additional benefit of taking action $a_t$ at state $s_t$ compared to the average case; $\text{clip}(\frac{\pi_{\phi}(a_t|s_t)}{\pi_{\phi'}(a_t|s_t)},1-\epsilon,1+\epsilon)$ represents clipping the probability ratio $\frac{\pi_{\phi}(a_t|s_t)}{\pi_{\phi'}(a_t|s_t)}$ to be between $1-\epsilon$ and $1+\epsilon$ to prevent the policy update from being too large, where $\epsilon$ is a hyperparameter used to control the clipping range.

5. The method according to claim 1, characterized in that, In step S4, the domain randomization technique is used to randomize the host configuration parameters and vulnerability descriptions of the original simulation environment, and diverse synthetic environments are generated, including: Using the original simulation environment and known vulnerability descriptions as inputs to a pre-trained large language model, diverse environment configurations and vulnerability descriptions are generated by the large language model, and the synthetic environment is output.

6. The method according to claim 1, wherein In step S5, the meta-learning algorithm includes inner-loop training and outer-loop training; In inner-loop training, the agent samples trajectories in each synthetic environment and updates the policy parameters; In outer-loop training, the agent updates the initial parameters to optimize the initial parameters of the model, enabling the agent to quickly adapt to new tasks.

7. The method according to claim 6, wherein The calculation formula for the inner-loop training is: ; Among them, ϕ is the parameter of the meta-policy; ϕ i ′ is the adapted policy parameter under task Mi; α is the learning rate of the inner loop; is the estimate of the expected discounted reward under task M i ; is the policy gradient estimate under task M i ; D i is the trajectory data collected under task Mi; The calculation formula for the outer-loop training is: ; Among them, ϕ is the parameter of the meta-policy; β is the learning rate of the outer loop; is the task distribution, including all tasks used for meta-training; is the policy adapted under task M i ; D i ′ is the trajectory data collected under task M i using the adapted policy ; represents the estimate of the expected discounted reward using the adapted policy under task ; represents the meta-policy gradient estimate, that is, the sum of the policy gradient estimates over all tasks M i .

8. The method according to claim 1, characterized in that In step S6, testing the generalization ability of the trained agent in the test environment includes: The agent conducts zero-shot policy transfer testing in a test environment similar to the training environment and evaluates the generalization gap. The calculation formula for the generalization gap is: ; where is the generalization gap; π is the policy, which is a function representing the probability distribution that maps the state s of the environment to the action a, i.e., π(a∣s) represents the probability of taking action a in state s; τ is the trajectory under the policy π, i.e., a sequence of states, actions, and rewards; P π is the trajectory distribution under the policy π; G(τ) is the expected cumulative reward of the trajectory τ, , where r t is the reward at time step t, γ is the discount factor used to reduce the weight of future rewards, and T is the total number of time steps; is the simulated environment or the synthetic environment generated through domain randomization; is the test environment, which is the target environment where the agent performs zero-shot policy transfer; 、 respectively represent the expected cumulative rewards of the trajectory τ generated according to the policy π in the environment 、 , which are the average performance metrics of the policy π in the environments 、 .

9. The method according to claim 1, wherein In step S6, testing the generalization ability of the trained agent in the test environment also includes: The agent conducts fast policy adaptation testing in a test environment dissimilar to the training environment, improving the agent's performance through a small amount of fine-tuning. The calculation formula for the adapted policy parameters is: ; Among them, ϕ is the parameter of the initial policy; ϕ adapted is the parameter of the adapted policy; α is the learning rate; π ϕ represents the policy with parameter ϕ; D test is the trajectory data collected under the test environment M test ; is the estimate of the expected discounted reward under the test environment M test ; is the policy gradient estimate under the test environment M test .

10. A network vulnerability intelligent repair system based on penetration testing technology, characterized in that, including: Agent construction module, used to construct an agent for penetration testing based on the reinforcement learning model; Training module, used to execute: Enable the agent to perform end-to-end policy learning in the target network environment, learning how to explore and exploit the vulnerabilities of target hosts through interaction with the original training environment. At the same time, the agent uses the reinforcement learning algorithm to update its policy parameters to maximize the expected cumulative reward; the target network environment is a real network environment; During the interaction between the agent and the target network environment, the agent constructs the original simulation environment using the collected host configuration data. The original simulation environment is a digital mapping of the original training environment; Based on the original simulation environment and vulnerability descriptions, the agent uses the domain randomization technique to randomize the host configuration parameters and vulnerability descriptions of the original simulation environment and generates diverse synthetic environments; In the synthetic environment, the agent uses the Proximal Policy Optimization algorithm and the model-agnostic meta-learning algorithm for in-depth training to optimize the agent's policy parameters and initial parameters, enabling the agent to extract generalizable policies and biases from diverse meta-training environments to generalize the learned policies to new or similar network environments; Test the generalization ability of the trained agent in the test environment; Vulnerability identification module, which is used to apply the agent that has completed training and passed the test to the real target network environment for vulnerability identification. The agent automatically identifies network vulnerabilities according to the actions executed in the network environment and the received rewards; Vulnerability repair module, which is used to select corresponding repair strategies for repair based on the identified network vulnerabilities.

Citation Information

Patent Citations

  • Automatic Windows domain penetration method based on reinforcement learning

    CN114444086A

  • Penetration test attack path discovery method based on PPO deep reinforcement learning model

    CN118118202A