Transferable automatic penetration testing method and medium integrating reinforcement learning and HER algorithm

By integrating reinforcement learning with HER algorithms in the penetration testing field, a transferable automatic penetration testing method is solved, and the problem of slow training of reinforcement learning models and poor adaptability across scenarios is realized, and the efficient use of the model and optimal strategy learning in different scenarios are achieved.

CN119377971BActive Publication Date: 2025-05-16ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411944204.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-16
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

In the field of penetration testing, reinforcement learning models are slow to train due to a large number of low feedback behaviors, and when the network scenario changes, the performance of the trained models in the original scenario in the new scenario is greatly reduced and the efficiency is not high.

Method used

A transferable automatic penetration testing method using a fusion reinforcement learning and HER algorithm is used to model the penetration testing process as a Markov decision-making process, a basic model is built using the DQN model, and the target state reconstruction experience pool data format is added during the training process to generate new empirical samples. A new penetration test model is built for new scenarios, the parameters of the optimal model in the original training scenario are migrated to the new model, and fine-tuned training is performed in the new scenario.

Benefits of technology

Accelerate the convergence of the model, improve penetration testing efficiency, and realize the low-cost and efficient cross-scene use of the model in different scenarios, avoiding the problem of reduced performance of the training model in the original scenario in the new scenario.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119377971B_ABST
    Figure CN119377971B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of network security and machine learning, and discloses a transferable automatic penetration testing method and medium integrating reinforcement learning and HER algorithm, including modeling the penetration testing process as a Markov decision process, using a DQN model to build a basic model of automatic penetration testing; adding a target state to the MDP task structure of the basic model, and reconstructing the data format of the experience samples in the experience pool according to the target state to obtain a penetration testing model; in the training process of the penetration testing model, generating new experience samples according to the experience samples in the experience pool until the training is completed and the optimal penetration testing model in the training scenario is output; constructing a new penetration testing model for a new scenario, and migrating the parameters of the optimal penetration testing model in the training scenario to the new penetration testing model; freezing the migrated parameters in the new penetration testing model to train the new penetration testing model. The present invention accelerates model convergence and more efficiently obtains the optimal penetration strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security and machine learning, and specifically relates to a transferable automatic penetration testing method and medium integrating reinforcement learning and HER algorithm. Background Art

[0002] Penetration testing (PT) is an important means of evaluating the security of computer network systems. By simulating the malicious behavior of hackers and using various means to attack the security vulnerabilities of the system, it helps relevant organizations discover the shortcomings of their own systems, thereby improving security measures and protecting sensitive data. However, today's network equipment is huge and complex, and the cost of traditional manual penetration testing is too high and inefficient. Many automated penetration testing methods have been proposed, such as describing the PT process as an attack graph to solve the penetration path planning problem, but they still rely on attack rules set by experts and are therefore unable to handle dynamic and unknown network environments.

[0003] On the other hand, with the development of artificial intelligence technology, automatic penetration testing methods based on machine learning have received widespread attention. For example, the PT method based on reinforcement learning (RL) is similar to humans walking through a maze, and finally obtains an optimal path after continuous attempts. Compared with the PT path planning technology based on attack graphs, the PT method based on RL can autonomously learn a new optimal penetration path and significantly improve the PT efficiency, but this method still has certain challenges. Due to the complexity of the network, attackers will generate a large number of low-feedback behaviors during the penetration process, which is more serious in penetration testing tasks based on reinforcement learning. The reinforcement learning agent can only obtain a small amount of successful experience and reward signals due to a large number of worthless behaviors, resulting in slow model training. And when the network scenario changes, the RL model trained in the original scenario has a significantly reduced PT performance in the new scenario, or even cannot be used. It is necessary to train a new RL model from scratch, which is inefficient.

[0004] In response to the above problems, how to deal with the difficulty of slow model training caused by a large number of low-feedback behaviors of agents in the penetration field of reinforcement learning models, and how to improve the efficiency of low-cost cross-scenario use of models in different scenarios is an issue that needs to be solved urgently. Summary of the invention

[0005] The purpose of the present invention is to provide a transferable automatic penetration testing method and medium that integrates reinforcement learning and HER algorithm, accelerate model convergence, and obtain the optimal penetration strategy more efficiently.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] A transferable automatic penetration testing method integrating reinforcement learning and HER algorithm, the transferable automatic penetration testing method integrating reinforcement learning and HER algorithm comprising:

[0008] The penetration test process is modeled as a Markov decision process to obtain the MDP task structure. Based on the MDP task structure, the DQN model is used to build the basic model of automatic penetration testing.

[0009] Add the target state to the MDP task structure of the basic model, and reconstruct the data format of the experience samples in the experience pool according to the target state to obtain the penetration test model;

[0010] During the training process of the penetration test model, new experience samples are generated based on the experience samples in the experience pool until the training is completed and the optimal penetration test model in the training scenario is output. The generation operation includes taking the target state corresponding to the next state in the experience sample as the target state of the new experience sample;

[0011] Build a new penetration test model for the new scenario, and migrate the parameters of the optimal penetration test model in the original training scenario to the new penetration test model;

[0012] Freeze the migrated parameters in the new penetration test model, train the new penetration test model in the new scenario, and generate new experience samples based on the experience samples in the experience pool during the training process until the training is completed and the optimal penetration test model in the new scenario is output.

[0013] Several optional methods are also provided below, but they are not intended to be additional limitations on the above-mentioned overall solution, but are merely further supplements or preferences. Under the premise that there are no technical or logical contradictions, each optional method can be combined with the above-mentioned overall solution separately, and multiple optional methods can also be combined.

[0014] Preferably, the penetration testing process is modeled as a Markov decision process, comprising:

[0015] Define the penetration test network environment as a state space, including network configuration information and host configuration information;

[0016] Define network scanning, vulnerability scanning, vulnerability exploitation, lateral movement, and privilege escalation operations as action spaces;

[0017] The probability value of the environment changing to the next state after taking an action in the current state is defined as the state transition probability;

[0018] The quality of each action performed during the penetration test is defined as a reward function. If the action is successfully executed, a positive reward value is obtained, otherwise a negative reward value is obtained.

[0019] Preferably, the step of adding a target state to the MDP task structure of the basic model includes:

[0020] The MDP task structure in the initially constructed DQN model is represented as a tuple , is the state space, is the action space, is the state transition probability, is the reward function, is the discount factor;

[0021] To tuple Add a new element to the target state , the target state represents the target state that the DQN model expects to reach in a state, then the updated MDP task structure is represented as a tuple ;

[0022] and and Add a mapping relationship between them, for all states There are corresponding target states , so the input of the DQN model is updated to , where the mapping relationship is as follows:

[0023]

[0024] in, Represents the mapping relationship function, represents the target state set, Indicates status To the target state The mapping relationship, For all states There are corresponding target states .

[0025] Preferably, the data format of the experience samples in the experience pool is reconstructed according to the target state, including:

[0026] The data format of the experience samples stored in the initial DQN model experience pool is , an experience sample is represented as a state To status The transformation express The state of the moment, express The action of the moment, express The rewards of the moment, express The state of the moment;

[0027] Add target state to the MDP task structure of the base model Afterwards, the data format of the experience sample is reconstructed into , after reconstruction, the environment state is changed from state The corresponding target state The joint statement The reward at the moment is associated with the state With target status The mapping relationship, that is, when the state With target status When the mapping relationship is established, the reward value The value of is positive; otherwise the reward value The value of is negative, express Always subject to target status The reward value of the impact.

[0028] Preferably, the generating operation further includes updating the reward value to a positive value, then the new experience sample is expressed as , express The state of the moment, Indicates status The corresponding target state, express The action of the moment, express Always subject to target status The reward value of the impact, express The state of the moment, Indicates status The corresponding target state is the target state of the new experience sample.

[0029] Preferably, the step of migrating the parameters of the optimal penetration test model in the original training scenario to the new penetration test model includes:

[0030] For parameter migration of the input layer in the DQN model, the input features of the original training scene and the new scene are compared, the corresponding parameters are directly copied for the overlapping parts, and the parameters of the corresponding features are randomly initialized for the new input features.

[0031] For the parameter migration of the hidden layer in the DQN model, the hidden layer structure of the optimal penetration test model in the original training scenario and the newly constructed penetration test model in the new scenario are compared. If the structures are the same, all the weights and bias parameters of the hidden layer are directly copied; if there are differences, the weights and bias parameters of the same part are copied, and the weights and bias parameters of different parts are randomly initialized.

[0032] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the transferable automatic penetration testing method integrating reinforcement learning and HER algorithm are implemented.

[0033] The present invention provides a transferable automatic penetration testing method and medium integrating reinforcement learning and HER algorithm, which has the following beneficial effects compared with the prior art:

[0034] 1. Combining the reinforcement learning model with the HER algorithm in the field of penetration testing can effectively solve the problem of slow model training caused by the lack of successful experience and positive rewards in the training process of the reinforcement learning model in the penetration testing scenario, thereby accelerating model convergence and obtaining the optimal penetration strategy more efficiently.

[0035] 2. Combined with the parameter migration method in transfer learning, the reinforcement learning model in the penetration testing task can quickly adapt to the new penetration testing scenario without retraining the model, so that the reinforcement learning model can complete the penetration testing tasks in different scenarios at low cost and high efficiency.

[0036] 3. Compared with the existing model-based task instance sample transfer method, the reinforcement learning model parameter transfer method has simpler and more efficient implementation steps. This method is based on fine-tuning training based on existing model parameters, does not require processing a large amount of complex task instance sample data, and can avoid the interference of irrelevant instance samples in the original scene and the problem of mismatch of task instance samples between the two scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A flow chart of a transferable automatic penetration testing method integrating reinforcement learning and HER algorithm of the present invention;

[0038] Figure 2 A flowchart of the HER algorithm incorporated into the basic model of the present invention;

[0039] Figure 3 A schematic diagram of a method for implementing parameter migration by applying transfer learning to the penetration testing model of the present invention. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0042] like Figure 1 As shown, this embodiment provides a transferable automatic penetration testing method integrating reinforcement learning and HER algorithm, including the following steps:

[0043] (1) Constructing the basic model of automatic penetration testing based on the reinforcement learning DQN algorithm: According to the definition of the reinforcement learning paradigm, the penetration testing process is modeled as a Markov decision process (MDP) to obtain the MDP task structure. Based on the MDP task structure, the DQN model is used to construct the basic model of automatic penetration testing.

[0044] (1-1) Modeling the penetration testing process as an MDP task structure: According to the reinforcement learning system paradigm, the penetration testing process is modeled as a Markov decision process. The MDP consists of a tuple Formalization, where represents the state space, represents the action space, represents the state transition probability, represents the reward function, Represents the discount factor parameter. Specifically, the specified penetration test network environment is defined as the state space , including network configuration information and host configuration information; defining available network scanning, vulnerability scanning, vulnerability exploitation, lateral movement, and privilege escalation operations as action spaces ; The current state Take action After the environment changes to the next state The probability value is defined as the state transition probability ; Define the degree of goodness of each action performed during the penetration test as a reward function , if the action If the execution is successful, a positive reward value will be obtained, otherwise the opposite; is the discount factor used to determine the importance of long-term rewards.

[0045] (1-2) Use the DQN model to build the basic model of automatic penetration testing: The DQN model contains two neural networks, namely the Q network and the target network. Use a convolutional neural network to build the Q network and randomly initialize the Q network parameters. The network structure includes an input layer, multiple hidden layers, and an output layer. The input layer is used to receive the feature vector of the environment state; the fully connected layer is used as the hidden layer to learn the abstract representation of the state characteristics, and the ReLU activation function is used to introduce nonlinear transformations to process the output of the hidden layer, so that the neural network can learn and represent more complex functional relationships; the value of the output layer is the Q value of the action selected in a given state. The target network is copied from the Q network. The goal of the DQN agent (the agent is the main body of the reinforcement learning model, represents the model itself, and is responsible for interacting with the environment) is to learn the optimal strategy , it observes the state of the environment, selects and executes actions according to the Q value, then updates the state of the environment and obtains the returned reward value, with the goal of maximizing the reward value.

[0046] The above steps are stored in the experience pool of the DQN model to update the Q value, and then the reward value and Q value are used to update the neural network parameters. Then the above steps are repeated, and actions are selected, executed, and Q values ​​and network parameters are updated continuously until the DQN model converges and obtains the optimal strategy, or the model reaches the specified limit number of iterations. The action selection function enables the DQN agent to correctly select the action in the state The optimal action to maximize the Q value is expressed as follows:

[0047]

[0048] Calculate the target Q value, that is, the expected return of taking a certain action in the current state. The formula is as follows:

[0049]

[0050] The update formula for defining the Q value is as follows:

[0051]

[0052] in, Indicates the status, Indicates action, Indicates the next state, Indicates the next action. Indicates in status The best action to choose is Indicates in status Next select action The optimal action value function value of are model parameters, Expressed as the target Q value, that is, the expected return of taking an action in the current state, Indicates reward, Indicates in status Next select action The action-value function value of Indicates in status Next select action The action-value function value of represents the learning rate, that is, the weight between the new and old Q values ​​when updating the Q value, Indicates in status Next select action The action-value function value of .

[0053] (2) Optimize the MDP task structure and reconstruct the experience sample based on the HER algorithm: Figure 2 As shown, the target state is added to the MDP task structure of the basic model As a new element, the data format of the experience samples generated during model training is reconstructed, and new experience samples are generated during the training process.

[0054] (2-1) Optimizing the MDP task structure: The MDP task structure in the constructed DQN model is represented as a tuple , add a new element target state to the tuple structure , which represents the target state that the DQN agent expects to achieve in a specific state, thus forming a new MDP task structure tuple ,and and There is a one-to-one mapping relationship between all states Each has its corresponding target state , so the input of the model is no longer the original state , but the state and the target state are used as common input The mapping relationship is as follows:

[0055]

[0056] in, Represents the mapping relationship function, represents the set of environmental states, represents the target state set, Indicates the status To the target state A mapping relationship of For all states Each has its corresponding target state , their mapping relationship always holds true. If it is equal to 0, it means that the mapping relationship does not hold. For state s, the target state g does not correspond to it.

[0057] (2-2) Reconstruct the experience sample data format in the experience pool: The format of the permeation experience data stored in the constructed DQN model experience pool is , one piece of experience data is represented as a state To status In increasing the target state Afterwards, the data format of the experience samples stored in the experience pool is reconstructed and expressed as , represents an integral element in the empirical sample, represented by and The two elements are jointly represented in the reconstructed empirical sample data format. The environmental state at this moment. At this time, the environmental state is determined by the state The corresponding target state Joint representation, simultaneous reward value It will also follow and The positive and negative rewards will fluctuate due to the mapping relationship of the target element. Specifically, when the state With target status When the mapping relationship is established (that is, the attack experience is a successful experience), If the value is positive, it is set to a negative value; otherwise, it is set to a negative value.

[0058] Steps (2-1) and (2-2) are performed before the interaction training between the model and the environment. Therefore, the data format of the experience samples generated after the interaction between the model and the environment is , the target state at this time According to the status The desired next state Generated, the transition process of the environment state in a single training is random, so if the next state in training Status The desired next state, then the state With target status The mapping relationship is established, then the experience sample is a successful experience sample; otherwise, it is a failed experience sample.

[0059] (2-3) Generate a new experience sample for each experience sample: During the training process of the penetration test model, an additional target state is set for each experience sample Specifically, select The target state corresponding to the state As the new target of the experience data. On the basis of retaining the original failed experience sample, a new experience sample is reconstructed and the reward value is re-determined is positive or negative, then the new experience sample is Since the target state in the new experience sample Based on confirmed status Generate, so the state With target status The mapping relationship must hold, and the newly generated experience sample must be a successful experience sample, so the reward value is set is positive. Finally, the new experience sample is stored in the experience replay buffer. This generation operation actually not only increases the attack experience sample data, but also increases the percentage of successful attack experience samples. Therefore, the DQN agent can get more positive rewards during the training process, thereby accelerating the model convergence speed.

[0060] (3) Accelerate the training of reinforcement learning models in new scenarios based on transfer learning: Figure 3 As shown, the transferable neural network parameters are selected in the penetration test model that has been trained in the original network scenario (which can be regarded as the training scenario), a new penetration test model is built in the new network scenario (which can be regarded as the new scenario), the model parameters specified in the original model are migrated to the corresponding parameters in the new model, and some neural network parameter values ​​in the new model are frozen. Finally, the training of the new model in the new scenario is completed.

[0061] (3-1) Screening parameters to be transferred: The parameters to be transferred in the DQN model are mainly the parameters of each layer of the neural network. The DQN model takes the scene state as input. The network structure and scale are different in different network scenes, so the corresponding input features are different. The model will output the Q value corresponding to the function action based on the learned state action. Therefore, the parameters to be transferred include the input layer parameters and hidden layer parameters of the neural network.

[0062] (3-2) Build a new penetration test model: Refer to step (1), step (2-1) and step (2-2) to build the same DQN algorithm-based penetration test model for the new network scenario (training has not started).

[0063] (3-3) Selective migration parameters: The neural network parameter values ​​of the original DQN model selected in step (3-1) are selectively assigned to the neural network layer corresponding to the new DQN model. Specifically, for the parameter migration of the input layer, first compare the input features of the corresponding scenes of the new and old models. The corresponding parameters can be directly copied for the overlapping parts, and the parameters of the corresponding features need to be randomly initialized for the new input features; for the parameter migration of the hidden layer, compare and analyze the hidden layer structures of the new and old models. If the structures are the same, directly copy all the weights and bias parameters of the hidden layer. If there are differences, copy the weights and bias parameters of the same parts. For the weights and bias parameters of different parts, perform random initialization in the new model.

[0064] (3-4) Selective freezing of parameters: Select the parameters of the feature extraction layer with excellent performance in the original DQN model, that is, the parameters of all the migrated input layers and hidden layers, and freeze the corresponding layer parameters in the new model. The parameters will not be updated during the back-propagation step of the new model training process. Since the feature extraction layer at the bottom of the DQN model neural network has strong versatility for different tasks, the parameters migrated from the original model already have good initial values, so freezing the parameters of this layer can help the new model converge faster.

[0065] (3-5) Train a new model to achieve automatic penetration testing: After completing parameter migration and parameter freezing adjustments, start training a new penetration testing model in the new network scenario. The DQN agent continuously interacts with the environment. At the same time, according to step (2-3), new experience samples are generated according to the experience samples in the experience pool during the training process, and finally the optimal penetration strategy is learned. After the training is completed, the optimal penetration testing model in the new scenario is output. The optimal penetration testing model is used for automatic penetration testing.

[0066] The model training experiment of the present invention is carried out based on the network attack simulator NaSim platform: 6 penetration test network scenarios with different network scales and structures are constructed, and the number of hosts and host configurations in different network scenarios are different, as shown in Table 1.

[0067] Table 1 Host quantity and host configuration

[0068]

[0069] In Table 1, ①, ②, ③, ④, ⑤ and ⑥ represent 6 different penetration test network scenarios respectively. The HTTP in the host configuration stands for HyperText Transfer Protocol; SSH stands for Secure Shell, a network protocol used to encrypt remote login sessions and other network services; FTP stands for File Transfer Protocol; SMTP stands for Simple Mail Transfer Protocol; SAMBA stands for Server Message Block, also known as the network file sharing protocol.

[0070] In network scenario ①, host 1 runs HTTP; host 2 runs SMTP; host 3 runs HTTP and SMTP.

[0071] In network scenario ②, host 1 runs HTTP and FTP; host 2 runs HTTP and SSH; host 3 runs HTTP; host 4 runs SSH; host 5 runs FTP and SSH; host 6 runs HTTP and SSH; host 7 runs HTTP, FTP, and SSH; host 8 runs HTTP, FTP, and SSH.

[0072] In network scenario ③, host 1 runs FTP; host 2 runs SSH; host 3 runs SMTP; host 4 runs FTP and SSH; host 5 runs FTP and SMTP; host 6 runs SSH and SMTP; host 7 runs FTP and SSH; host 8 runs FTP, SSH, and SMTP; host 9 runs FTP, SSH, and SMTP; host 10 runs FTP, SSH, and SMTP.

[0073] In network scenario ④, host 1 runs HTTP; host 2 runs FTP; host 3 runs SMTP; host 4 runs SAMBA; host 5 runs FTP and SMTP; host 6 runs HTTP and SMTP; host 7 runs FTP and HTTP; host 8 runs FTP, SAMBA, and SMTP; host 9 runs FTP, SAMBA, and SMTP; host 10 runs FTP, HTTP, and SMTP; host 11 runs FTP, HTTP, and SAMBA; host 12 runs FTP, HTTP, SMTP, and SAMBA; host 13 runs FTP, HTTP, SMTP, and SAMBA; host 14 runs FTP, HTTP, SMTP, and SAMBA.

[0074] In network scenario ⑤, host 1 runs HTTP; host 2 runs SSH and FTP; host 3 runs FTP; host 4 runs SMTP; host 5 runs FTP and SMTP; host 6 runs SSH and SMTP; host 7 runs HTTP, FTP, and SSH; host 8 runs HTTP, FTP, SSH, and SMTP; host 9 runs HTTP, FTP, SSH, and SMTP; host 10 runs FTP, SSH, SMTP, and SAMBA. Host 11 runs FTP, HTTP, and SAMBA; host 12 runs SSH, FTP, HTTP, SMTP, and SAMBA; host 13 runs SSH, FTP, HTTP, SMTP, and SAMBA; host 14 runs SSH, FTP, HTTP, SMTP, and SAMBA.

[0075] In network scenario ⑥, host 1 runs HTTP and SSH; host 2 runs FTP; host 3 runs FTP and SMTP; host 4 runs SMTP; host 5 runs FTP and SMTP; host 6 runs SSH and SMTP; host 7 runs SSH, FTP, and SMTP; host 8 runs HTTP, FTP, SSH, and SMTP; host 9 runs FTP, SSH, SMTP, and SAMBA; host 10 runs HTTP, FTP, SSH, SMTP, and SAMBA. Host 11 runs HTTP, FTP, HTTP, and SAMBA; host 12 runs SSH, FTP, HTTP, SMTP, and SAMBA; host 13 runs SSH, FTP, HTTP, and SMTP; host 14 runs HTTP, SSH, FTP, SMTP, and SAMBA; host 15 runs HTTP, SSH, FTP, SMTP, and SAMBA; host 16 runs HTTP, SSH, FTP, SMTP, and SAMBA.

[0076] The network scenario ④ is selected as the training scenario, and the training results of the model of the present invention are compared with the training results of other reinforcement learning models, including the DQN (Deep Q-Network) model, the Double DQN (Double Deep Q-Network) model, and the DPPO (Distributed Proximal Policy Optimization) model. Then, the parameter migration method based on transfer learning is used to migrate the trained model parameters to the new model under the new scenario (other network scenarios, i.e., network scenarios ①, ②, ③, ⑤, and ⑥) and fine-tune and retrain, and a comparison experiment of the training results under the new scenario is performed. The experimental results are shown in Table 2.

[0077] Table 2 Experimental results

[0078]

[0079] The results show that compared with other reinforcement learning models that do not incorporate the HER algorithm, the model training convergence speed of the present invention is faster, the model is more efficient in learning the optimal strategy, and the model of the present invention can spend less time and cost to complete the learning and acquisition of new optimal strategies in new scenarios. The experimental results strongly prove the effectiveness and usability of the present invention, and can enable the trained model to be applied across multiple scenarios.

[0080] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0081] The above-mentioned embodiments only express several implementation modes of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.

Claims

1. A transferable automatic penetration testing method integrating reinforcement learning and HER algorithm, characterized in that: The transferable automatic penetration testing method integrating reinforcement learning and HER algorithm includes: The penetration test process is modeled as a Markov decision process to obtain the MDP task structure. Based on the MDP task structure, the DQN model is used to build the basic model of automatic penetration testing. Add the target state to the MDP task structure of the basic model, and reconstruct the data format of the experience samples in the experience pool according to the target state to obtain the penetration test model; During the training process of the penetration test model, new experience samples are generated based on the experience samples in the experience pool until the training is completed and the optimal penetration test model in the training scenario is output. The generation operation includes taking the target state corresponding to the next state in the experience sample as the target state of the new experience sample; Build a new penetration test model for the new scenario, and migrate the parameters of the optimal penetration test model in the original training scenario to the new penetration test model; Freeze the migrated parameters in the new penetration test model, train the new penetration test model in the new scenario, and generate new experience samples based on the experience samples in the experience pool during the training process until the training is completed and the optimal penetration test model in the new scenario is output.

2. The transferable automatic penetration testing method integrating reinforcement learning and HER algorithm according to claim 1 is characterized in that: The penetration testing process is modeled as a Markov decision process, including: Define the penetration test network environment as a state space, including network configuration information and host configuration information; Define network scanning, vulnerability scanning, vulnerability exploitation, lateral movement, and privilege escalation operations as action spaces; The probability value of the environment changing to the next state after taking an action in the current state is defined as the state transition probability; The quality of each action performed during the penetration test is defined as a reward function. If the action is successfully executed, a positive reward value is obtained, otherwise a negative reward value is obtained.

3. The transferable automatic penetration testing method integrating reinforcement learning and HER algorithm according to claim 1 is characterized in that: The target state is added to the MDP task structure of the basic model, including: The MDP task structure in the initially constructed DQN model is represented as a tuple , is the state space, is the action space, is the state transition probability, is the reward function, is the discount factor; To tuple Add a new element to the target state , the target state represents the target state that the DQN model expects to reach in a state, then the updated MDP task structure is represented as a tuple ; and and Add a mapping relationship between them, for all states There are corresponding target states , so the input of the DQN model is updated to , where the mapping relationship is as follows: ; in, Represents the mapping relationship function, represents the target state set, Indicates status To the target state The mapping relationship, For all states There are corresponding target states .

4. The transferable automatic penetration testing method integrating reinforcement learning and HER algorithm according to claim 1 is characterized in that: The data format of the experience samples in the experience pool is reconstructed according to the target state, including: The data format of the experience samples stored in the initial DQN model experience pool is , an experience sample is represented as a state To status The transformation express The state of the moment, express The action of the moment, express The rewards of the moment, express The state of the moment; Add target state to the MDP task structure of the base model Afterwards, the data format of the experience sample is reconstructed into , after reconstruction, the environment state is changed from state The corresponding target state The joint statement The reward at the moment is associated with the state With target status The mapping relationship, that is, when the state With target status When the mapping relationship is established, the reward value The value of is positive; otherwise the reward value The value of is negative, express Always subject to target status The reward value of the impact.

5. The transferable automatic penetration testing method integrating reinforcement learning and HER algorithm according to claim 1 is characterized in that: The generation operation also includes updating the reward value to a positive value, then the new experience sample is expressed as , express The state of the moment, Indicates status The corresponding target state, express The action of the moment, express Always subject to target status The reward value of the impact, express The state of the moment, Indicates status The corresponding target state is the target state of the new experience sample.

6. The transferable automatic penetration testing method integrating reinforcement learning and HER algorithm according to claim 1 is characterized in that: The step of migrating the parameters of the optimal penetration test model in the original training scenario to the new penetration test model includes: For parameter migration of the input layer in the DQN model, the input features of the original training scene and the new scene are compared, the corresponding parameters are directly copied for the overlapping parts, and the parameters of the corresponding features are randomly initialized for the new input features. For the parameter migration of the hidden layer in the DQN model, the hidden layer structure of the optimal penetration test model in the original training scenario and the newly constructed penetration test model in the new scenario are compared. If the structures are the same, all the weights and bias parameters of the hidden layer are directly copied; if there are differences, the weights and bias parameters of the same part are copied, and the weights and bias parameters of different parts are randomly initialized.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the transferable automatic penetration testing method integrating reinforcement learning and HER algorithm described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method and device for defending penetration attack based on reinforcement learning, and electronic equipment

    CN115473677A

  • Deep reinforcement learning intelligent penetration testing method and device based on imitation learning

    CN115473706A