Intelligent penetration testing method and device based on deep reinforcement learning

Through the intelligent penetration testing method of deep reinforcement learning, Markov decision-making process and priority replay experience pool are built, and attack strategies are dynamically optimized, which solves the problems of insufficient processing capabilities of high-dimensional state space in the existing technology and weak generalization capabilities, and realizes efficient penetration testing in complex network environments.

CN120263476APending Publication Date: 2025-07-04BEIJING UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510411571.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing network penetration testing technology has insufficient processing capabilities in high-dimensional state space, poor real-time performance, weak strategy generalization capabilities, and is difficult to cope with active interference from dynamic network environments and defense parties, resulting in inefficient efficiency and high false alarm rate, and is unable to effectively deal with complex and changeable network attacks.

Method used

Using an intelligent penetration testing method based on deep reinforcement learning, Markov decision-making process and priority replay experience pool are built, and the attack strategy is dynamically optimized through the actor-critic structure of the policy network and the value network, and combined with the Soft Actor-Critic algorithm, the balance between strategy exploration and utilization is improved, and training efficiency and stability are improved.

Benefits of technology

Quickly identify potential attack paths in complex network environments, improve path discovery efficiency and accuracy, enhance system adaptability and stability, adapt to dynamic network changes, and reduce training complexity and overfitting risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263476A_ABST
    Figure CN120263476A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intelligent penetration testing method and device based on deep reinforcement learning, and the method comprises the steps: constructing a network information environment, setting basic parameters, initializing the network information environment, and constructing an act-critic structure with an independent strategy network and an independent value network; modeling the Markov decision process of the penetration test into a quaternion group lt; s, A, R, Pgt; a Markov model is constructed; constructing a priority playback experience pool; performing penetration testing based on a Markov model, calculating and outputting an attack action to a network information environment through a strategy network based on a current penetration state to obtain a next penetration state, rewarding the attack action through a value network, and storing data in a priority playback experience pool; and calculating and determining a mode of preferentially selecting sampling according to the priority of experience, repeating the penetration test operation of the Markov model until the function is converged, and outputting an optimal penetration strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of network security technology, and in particular, to an intelligent penetration testing method and device based on deep reinforcement learning. Background Art

[0002] The current mainstream network penetration testing technologies generally face dual bottlenecks of efficiency and intelligence. Traditional manual penetration testing relies on security experts to manually analyze network topologies, vulnerability databases (such as CVE) and firewall rules, and gradually explore attack paths through toolchains (such as Nmap, Metasploit). The cost of a single test is expensive and it takes up to several weeks, and it is easy to miss hidden paths (such as cross-domain jump attacks) due to personnel experience. Although automated scanning tools (such as Nessus, OpenVAS) have improved the speed of vulnerability detection, their rule-based matching mechanism results in a false alarm rate as high as 30%-40%, and they cannot generate action sequences for multi-step attacks, and the coverage rate of zero-day vulnerabilities is less than 5%. Systems based on predefined policies (such as Metasploit module chains) are rigid in dynamic network environments. For example, when encountering real-time policy adjustments of software-defined networks (SDN) or honeypot lures, the attack success rate drops by more than 70%. Early reinforcement learning methods (such as Q-learning) tried to model penetration as a discrete decision-making process, but it was difficult to handle network topologies with more than 50 nodes due to the curse of dimensionality, and the probability of policy failure in cross-scenario generalization tests reached 62.8% (experimental data). The core defects of the above methods are as follows: First, the processing ability of the high-dimensional state space (covering thousands of features such as device types, service versions, and vulnerability CVSS scores) is insufficient, resulting in poor real-time path planning (delay > 10 minutes); second, the policy generalization mechanism is missing, and the attack success rate drops sharply when the network configuration deviates from the training set; third, the adaptability to adversarial scenarios is weak, and it cannot effectively cope with the active interference behaviors of the defense party (such as traffic confusion, dynamic credential updates). These limitations seriously restrict the application effectiveness of penetration testing technology in advanced persistent threat (APT) defense.

[0003] With the continuous development and increasing complexity of network attack means, traditional penetration testing technologies no longer meet the current multi-stage attack methods and types. Therefore, discovering intelligent attack paths based on the deep reinforcement learning (DRL) algorithm has become a research hotspot. The following is a comparative analysis of similar domestic and foreign technical solutions to the present invention:

[0004] 1. Discovery of network attack paths based on traditional methods

[0005] In traditional network attack path discovery at home and abroad, many studies have adopted technical means such as rule-based security policies, attack graphs, and probability graphs. Traditional network attack path discovery methods generally manually formulate attack paths by scanning network devices, analyzing configuration files, and checking known vulnerabilities. For example, the research method based on the Attack Graph analyzes possible attack paths from the perspective of graph theory by constructing an attack graph model of the network. Although these methods can effectively display attack paths, their processing capabilities are limited in the face of complex and ever-changing network environments, especially in high-dimensional network scenarios and dynamic environments, and it is difficult to adapt to real-time changes.

[0006] 2. Network Attack Strategy Discovery Based on Deep Reinforcement Learning

[0007] With the rapid development of deep reinforcement learning (DRL) technology, more and more studies have begun to use DRL algorithms to optimize attack paths. Typical algorithms adopted in existing studies at home and abroad include the Deep Q-Network (DQN) and its variant algorithms to solve the problem of generating attack strategies. And some improved reinforcement learning algorithms. For example, NoisyNet-A3C enhances the policy exploration ability by introducing a noise mechanism in the policy network, thereby accelerating convergence and improving stability; while NHSC-PPO (Noise Hierarchical Strategy Clustering with PPO) optimizes the path planning process through hierarchical policy clustering, effectively solving the problem of penetration path selection in complex networks. Although DQN and NHSC-PPO can provide certain path optimization, they usually require a large amount of computing resources and a long training time, and in dealing with high-dimensional complex environments, overfitting problems often occur, and it is difficult to effectively expand to different-scale network environments.

[0008] Nevertheless, the existing network attack strategy discovery methods based on DRL still have limitations such as low training efficiency and difficulty in quickly converging in high-dimensional action space scenarios; insufficient robustness and generalization ability of strategies in complex network topologies; and limited adaptability to changes in dynamic networks, making it difficult to effectively respond to the strategy adjustments of the defense side. Therefore, there is an urgent need for an efficient penetration testing method based on deep reinforcement learning to solve the above problems. Summary of the Invention

[0009] The purpose of the present invention is to provide an intelligent penetration testing method and device based on deep reinforcement learning, aiming to solve the above problems in the prior art.

[0010] The present invention provides an intelligent penetration testing method based on deep reinforcement learning, including:

[0011] Build a network information environment, set basic parameters, initialize the network information environment, and build an actor-critic structure with independent policy network and value network;

[0012] Model the Markov decision process of penetration testing as a quadruple <S, A, R, P>, and build a Markov model. Among them, the goal of the Markov model is to select the optimal penetration strategy so as to obtain the optimal attack action under a given penetration state for maximizing the reward. Here, S represents a set of observed penetration states, A represents a set of possible attack actions corresponding to different penetration states, R represents a reward function for allocating rewards according to different penetration states, and P is the attack state transition probability, which represents the possibility of the attacker transferring between different penetration states;

[0013] Build a prioritized replay experience pool;

[0014] Conduct penetration testing based on the Markov model. Based on the current penetration state through the policy network, calculate and output an attack action to the network information environment to obtain the next penetration state. Reward the attack action through the value network, save all attack actions and environmental feedback data in the prioritized replay experience pool, calculate and determine the sampling method for priority selection according to the priority of the experience, train the samples with high priority more frequently, repeat the penetration testing operation of the above Markov model until the function converges, output the optimal penetration strategy, generate a test case report according to the optimal penetration strategy, and obtain a penetration test report according to the test case report.

[0015] The present invention provides an intelligent penetration testing device based on deep reinforcement learning, including:

[0016] A construction module, configured to build a network information environment, set basic parameters, initialize the network information environment, and build an actor-critic structure with independent policy network and value network; model the Markov decision process of penetration testing as a quadruple <S, A, R, P>, and build a Markov model. Among them, the goal of the Markov model is to select the optimal penetration strategy so as to obtain the optimal attack action under a given penetration state for maximizing the reward. Here, S represents a set of observed penetration states, A represents a set of possible attack actions corresponding to different penetration states, R represents a reward function for allocating rewards according to different penetration states, and P is the attack state transition probability, which represents the possibility of the attacker transferring between different penetration states; build a prioritized replay experience pool;

[0017] The penetration testing module is used to perform penetration testing based on the Markov model. Based on the current penetration state through the policy network, it calculates and outputs attack actions to the network information environment to obtain the next penetration state. It rewards the attack actions through the value network, saves all attack actions and environmental feedback data in the prioritized replay experience pool, calculates and determines the sampling method based on the priority of the experience to train high-priority samples more frequently, repeats the penetration testing operation of the Markov model until the function converges, outputs the optimal penetration strategy, generates a test case report based on the optimal penetration strategy, and obtains a penetration testing report based on the test case report.

[0018] An embodiment of the present invention also provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the above-mentioned intelligent penetration testing method based on deep reinforcement learning.

[0019] An embodiment of the present invention also provides a computer-readable storage medium. An implementation program for information transmission is stored on the computer-readable storage medium. When the program is executed by the processor, it implements the steps of the above-mentioned intelligent penetration testing method based on deep reinforcement learning.

[0020] Adopting the intelligent penetration testing method based on deep reinforcement learning and improved reward strategy in the embodiments of the present invention can effectively improve the effectiveness of the deep reinforcement learning model in the field of penetration testing. Description of the Drawings

[0021] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a flowchart of the intelligent penetration testing method based on deep reinforcement learning in the embodiments of the present invention;

[0023] Figure 2 It is a specific flowchart of the intelligent penetration testing method based on deep reinforcement learning in the embodiments of the present invention;

[0024] Figure 3 It is a schematic diagram of the improved SAC algorithm structure in the embodiments of the present invention;

[0025] Figure 4 It is a schematic diagram of the implementation program in the embodiments of the present invention;

[0026] Figure 5 It is a schematic diagram of an intelligent penetration testing device based on deep reinforcement learning according to an embodiment of the present invention;

[0027] Figure 6 It is a schematic diagram of an electronic device according to an embodiment of the present invention. Specific embodiments

[0028] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification in conjunction with the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.

[0029] Method embodiments

[0030] According to an embodiment of the present invention, there is provided an intelligent penetration testing method based on deep reinforcement learning, Figure 1 It is a flowchart of an intelligent penetration testing method based on deep reinforcement learning according to an embodiment of the present invention, as Figure 1 shown, the intelligent penetration testing method based on deep reinforcement learning according to an embodiment of the present invention specifically includes:

[0031] Step S101, construct a network information environment and set basic parameters, initialize the network information environment, and construct an actor-critic structure with an independent policy network and value network;

[0032] Step S102, model the Markov decision process of penetration testing as a quadruple <S, A, R, P>, construct a Markov model, where the goal of the Markov model is to select an optimal penetration strategy so that the optimal attack action can be obtained under a given penetration state in order to maximize the reward. Among them, S represents a set of observed penetration states, A represents a set of possible attack actions corresponding to different penetration states, R represents a reward function for allocating rewards according to different penetration states, and P is the attack state transition probability, which represents the possibility of describing the attacker's transition between different penetration states;

[0033] Step S103, construct a priority replay experience pool;

[0034] Step S104: Conduct penetration testing based on the Markov model. Based on the current penetration state through the policy network, calculate and output attack actions to the network information environment to obtain the next penetration state. Reward the attack actions through the value network, save all attack actions and environmental feedback data in the prioritized replay experience pool, calculate and determine the sampling method for preferential selection according to the priority of the experience, so as to train samples with high priority more frequently. Repeat the above penetration testing operation of the Markov model until the function converges, output the optimal penetration strategy, generate a test case report according to the optimal penetration strategy, and obtain a penetration test report according to the test case report.

[0035] It should be noted that the network information environment specifically includes: network topology environment, node information collection environment status, node connection topology matrix, firewall configuration between hosts or subnets, host operating system configuration, process information, and protocol information; the basic parameters specifically include: learning rate, discount factor, target entropy value, soft update coefficient, and random seed; there is one policy network and four value networks; the penetration state specifically includes: initial unpenetrated state, external network information collection state, boundary breakthrough state, internal network penetration state, privilege escalation state, critical target access state, defense detection and blocking state, network environment dynamic change state, penetration failure state, and penetration success state; the attack actions specifically include: network scanning, host scanning, vulnerability exploitation, and privilege escalation.

[0036] In the above processing steps, rewarding the attack actions through the value network specifically includes:

[0037] Let the threat of a certain type of attack be xi, the cost of launching this type of attack be ai, the attack cost of node i be ni, and the attack value be vi. Then calculate the attack cost when attacking a certain node j state. Assuming that on the optimal attack path, the state of node i relative to the attacker is si, then there is formula 1;

[0038] C = min(C) = ∑(ai + ni) * si Formula 1;

[0039] Among them, i is less than j, C is less than or equal to LC, and LC is the established cost, that is, the maximum cost that the resources held by the attacker can consume; C represents the attack cost when attacking a certain node j state;

[0040] Calculate the damage W caused by this type of attack to the network according to Formula 2;

[0041] W = max(W) = ∑( xi + vi ) * pi Formula 2;

[0042] Among them, pi is the node attack success probability, and its value is estimated by statistically analyzing historical attack data, simulating experimental results, or using the exploration feedback of reinforcement learning;

[0043] Calculate the attack reward function under the current penetration state according to Formula 3;

[0044] R = W - C = ∑( xi + vi ) * pi - ( ai + ni ) * si Formula 3.

[0045] Constructing a Markov model specifically includes:

[0046] The goal of constructing a Markov model is to select the optimal policy π such that the optimal action at can be obtained under the given state st to maximize the reward, that is, at = π( st ), so as to maximize the long-term cumulative reward G(s0) in the current state, as shown in Formula 4:

[0047]

[0048] Among them, γ ∈ (0, 1) represents the discount factor, which is used to balance the importance of the current reward and future rewards. E{}· represents the expectation operator, T represents the total number of steps in a round. In the reinforcement learning task with a finite number of steps, the interaction process of the agent ends after T steps. t represents the current step of the agent's interaction with the environment, starting from 0.

[0049] Constructing an actor-critic structure with independent policy networks and value networks specifically includes:

[0050] Determine the update loss expression of the policy network according to Formula 5 and Formula 6:

[0051]

[0052] Among them, represents the prioritized replay experience pool, α is the entropy coefficient, which is used to control the importance of the entropy term lnπ(a′ t |s t ; θ). Its significance increases with the increase of α. st represents the state of the agent at step t of the round, θ represents the parameters of the policy network, at represents the actual executed action, from historical experience, and is used for the update of the Q function. rt represents the reward function value in state st, a′ t represents the action resampled from the current policy π(.| st ; θ) during policy evaluation or improvement, represents representing sampling the action at′ from the current policy π(.| st ; θ) and calculating the expectation, q0(st , a′ t ) represents the Q function, which is used to estimate the expected cumulative reward that can be obtained after taking the action at′ in the state st. represents the set of all possible actions a in the state st, w (0) represents the parameter of the Q function. Entropy represents a measure of the randomness degree of a random variable. If X is a random variable with the probability density function p, then its entropy H(X) is defined as:

[0053]

[0054] Among them, represents taking the expectation of the random variable X under the probability distribution p(x), and p(x) represents the probability density function of the random variable X.

[0055] Constructing the prioritized replay experience pool specifically includes:

[0056] Divide the prioritized replay experience pool into multiple regional replay pools. Among them, each regional replay pool corresponds to a subnet or node;

[0057] Assign a priority to each experience, which is calculated according to the temporal difference TD error. The larger the TD error, the greater the potential for policy improvement and the higher the priority.

[0058] Save all attack actions and environmental feedback data in the prioritized replay experience pool. Calculating and determining the sampling method preferentially according to the priority of the experience specifically includes:

[0059] When updating the penetration strategy each time, sample experiences from the regional replay pool related to the current operation area based on Formula 8, or sample experiences from the regional replay pools in different regions;

[0060]

[0061] Among them, P(i) is the sampling probability, p i represents the priority of experience i, and α is used to control the level of priority. When α = 1, sampling is completely prioritized according to the priority; when α = 0, sampling is uniform;

[0062] Correct the sampling bias, give lower weights to high-priority experiences in the gradient calculation, and calculate the importance sampling weight w(i) according to Formula 9 to balance the sampling bias:

[0063]

[0064] Among them, N is the size of the experience buffer pool, and β is a parameter for adjusting the importance and the degree of sampling correction. When β = 1, the correction is fully applied, and β gradually increases from 0.

[0065] The intelligent network attack strategy discovery method based on Markov process description and Soft Actor-Critic algorithm proposed in the embodiments of the present invention can convert network states into Markov Decision Process (MDP), combine with deep reinforcement learning algorithms to generate optimal strategies, can cope with changing network environments, dynamically adjust strategies, and does not rely on traditional manual design of paths. This method overcomes the defects of static planning and complex graph model construction in traditional methods, can flexibly adapt to environmental changes, and improves the efficiency and accuracy of path discovery.

[0066] Compared with the prior art, the embodiments of the present invention combine Markov process and Soft Actor-Critic algorithm, and avoid the limitations of traditional static path planning methods by dynamically optimizing attack strategies. In addition, the SAC algorithm finds a balance between exploration and exploitation by maximizing entropy, significantly improving the efficiency of policy optimization. Most existing studies use traditional reinforcement learning algorithms, often facing problems such as long training processes and difficult parameter adjustment. However, through the application of the SAC algorithm in the embodiments of the present invention, not only the complexity of training is reduced, but also the parameter adjustment is relatively more stable and predictable, reducing the energy and time invested by researchers in parameter tuning, and providing an efficient, stable and highly adaptable solution for intelligent network attack strategy discovery. The algorithm of the embodiments of the present invention shows significant advantages when dealing with high-dimensional state spaces, can converge in a short training time, and effectively avoids the overfitting problem in traditional deep reinforcement learning methods. This enables the embodiments of the present invention to flexibly respond and generate efficient attack strategies when facing more complex and changing network environments, and has strong practical application prospects. The application of the embodiments of the present invention in network penetration testing also shows significant advantages over traditional methods. Especially in complex network structures, it can quickly and accurately identify potential attack paths, greatly improving the efficiency of path discovery. At the same time, by introducing the SAC algorithm, it can effectively process multi-dimensional state spaces, improving the adaptability and stability of the system, making this method more in line with the attack and defense requirements in real scenarios.

[0067] The above technical solutions of the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0068] The typical penetration testing scenario involved in the embodiments of the present invention consists of an external network, a DMZ, and an internal network subnet. The subnet contains user nodes and data nodes (sensitive nodes). The function of the intelligent penetration agent is to determine the attack target at each step based on the state observation results, reach the final sensitive node, and plan the optimal solution.

[0069] The Markov decision process (MDP) of penetration testing is modeled as a quadruple <S, A, R, P>. The following is the detailed design of each element of this model:

[0070] S represents a set of penetration states observed by the agent, such as network topology, host information, and vulnerability details. These states correspond to different stages in the penetration testing process:

[0071] Initial unpenetrated state: At the starting point of penetration testing, the attacker is outside the network and has no internal network information.

[0072] External network information collection state: Starting from the initial state, the attacker collects external network information through means such as network scanning.

[0073] Boundary breakthrough state: After collecting a certain amount of information, the attacker attempts to break through the network boundary and obtain initial access rights.

[0074] Internal network penetration state: After breaking through the boundary, the attacker enters the internal network and starts penetrating internal nodes.

[0075] Privilege escalation state: In the internal network, the attacker attempts to escalate privileges to gain more control.

[0076] Critical target access state: After privilege escalation, the attacker can access the core network or critical assets.

[0077] Penetration failure state: When the attacker's resources are exhausted or there are no more actions to execute, the attacker cannot continue the penetration and the test ends.

[0078] Penetration success state: When the attacker successfully completes the penetration task and achieves the expected goal, such as stealing important data or completing system destruction, it enters this state.

[0079] A represents the set of attack actions, which represents a set of possible attack actions corresponding to the above different state stages. These operations include network scanning, host scanning, vulnerability exploitation, privilege escalation, and other operations available to the agent.

[0080] R represents the attack reward function, which is the reward function R(s). It assigns rewards according to different penetration states. For example, obtaining the highest level of privilege on a sensitive host gets a reward of 100, destroying a host in the internal network subnet gets a reward of 1, while destroying a DMZ host gets no reward.

[0081] During the evaluation process, the integrity, confidentiality, reliability, and availability of the target under attack can be used as a measure of its security, and the difference in security before and after the attack can be used as an evaluation criterion for the attack effect. Based on this idea, for the MDP process discussed above, penetration testing uses various attack methods to penetrate the network. Suppose the threat of a certain type of attack is Xi, the cost of launching this type of attack is Ai, the attack cost of node i is Ni, and the attack value is Vi. Then the attack cost in the state of attacking a certain node j. Suppose that on the optimal attack path, the state of node i relative to the attacker is S (0 represents no attack action, 1 represents successful attack action, and 2 represents failed attack action).

[0082] Then there is the equation:

[0083] C = min(C) = ∑(ai + ni) * si (i < j) (C ≤ LC)

[0084] LC (Limit Cost) is the established cost. Assume that the resources held by the attacker can consume at most the cost of LC, which can help the intelligent agent to accelerate the convergence speed during the network search attack process.

[0085] The damage W caused by this type of attack to the network = max(W) = ∑( xi + vi ) * pi

[0086] The attack reward function in the current state: R = W - C = ∑( xi + vi ) * pi - ( ai + ni ) * si (i < j) where pi is the probability of node breakthrough, and the attack threat level is related to the CVSS (Common Vulnerability Scoring System) of different nodes in different networks for corresponding service vulnerabilities.

[0087] Attack state transition probability: The attack state transition probability P is used in the modeling of network penetration testing to describe the possibility of an attacker (or intelligent agent) transferring between different penetration states. This probability can be estimated by statistical historical attack data, simulation experiment results, or exploration feedback of reinforcement learning, reflecting the evolution law of attack behavior at different stages and the impact of network defense mechanisms. P represents the state transition function P(s, a, s ) = P( s'|s , a), which defines the probability of transitioning from one state to another after executing a given action. This is usually related to the success rate of the attack operation, and the measurement of this probability was mentioned in the previous subsection.

[0088] The goal of constructing the MDP model is to select the optimal policy π such that the optimal action at can be obtained under the given state st to maximize the reward, i.e., at = π(st), so as to maximize the long-term cumulative reward G(s0) in the current state, as shown in formula (1).

[0089]

[0090] Among them, γ ∈ (0, 1) represents the discount factor, which is used to balance the importance of the current reward and future rewards.

[0091] Since in the enterprise network, the hosts storing sensitive data are often in relatively hidden positions in the network, we need the policy network to explore as many paths as possible, discover more hosts to find the target host, and thus calculate the optimal path and the least cost. As an efficient model-free reinforcement learning algorithm, SAC has obtained state-of-the-art results in many standard environments and is very suitable for intelligent penetration testing in complex network environments. Figure 3 Shows the architecture of the improved ASAC algorithm. The ASAC algorithm consists of a policy network and four value function networks. The policy network (Policy Net), as the actor, calculates and outputs attack actions to the environment to obtain the next state, while the four value networks evaluate the policy network. ASAC uses a prioritized experience replay buffer to store all action and environment feedback data, and calculates and preferentially selects the sampling method according to the priority of the experience, so as to train those samples with high priority more frequently, improving the training efficiency and stability. Update the network. Considering that the behaviors in reinforcement learning are usually highly correlated, this method allows the neural network to achieve more effective training.

[0092] Policy network design: It can be seen that the input of the policy network (Policy Net) as the Actor is the output of the four value networks, i.e., the network state St, and the output is the action policy P(ai|st). The update loss expression of its neural network is as follows:

[0093]

[0094] It should be noted that Q0( st , a‘t ; w(0)) can be replaced by Q1( st , a‘t ; w(1)) because the functions of the two Q - Critic networks are equivalent. Based on the idea of Double DQN, SAC uses two critic networks, but each time it uses a critic network, it selects the network with a smaller value, thus alleviating the overestimation problem.

[0095] Symbol Denotes the experience buffer, which means that when calculating the loss, you need to take the average of the samples extracted from the buffer. This ensures that the expected average meaningfully represents the overall result of the samples. Here, α is the entropy coefficient, which controls the importance of the entropy term lnπ(at+1|st; θ), and its significance increases with the increase of α.

[0096] As we mentioned earlier, entropy represents a measure of the degree of randomness of a random variable. Specifically, if X is a random variable with probability density function p, then its entropy H(X) is defined as:

[0097]

[0098] Based on the optimal Bellman equation, Ut(q)=rt + γV(st+1) is used as the estimate of the true value of state st, and the qi(st,at) values (where i = 0,1) are used as the estimates of the predicted values of state st and actual action at. Finally, the MSE Loss is used as the loss function to train the neural networks Q0 and Q1. Note that using the MSE loss means taking the average of the data sampled from a batch of sample buffers (denoted as ) as follows:

[0099]

[0100] V Critic network design:

[0101] The following entropy formula is used for state value estimation:

[0102]

[0103] The derivation is written as:

[0104]

[0105] Among these terms, π(at’|st; θ), mini = 0,1qi(st; at’; w(i), lnπ(at’|st; θ) exactly matches the loss calculation described in Figure 3 Taking the output of the V Critic network as the predicted value, and finally using the MSE Loss as the loss function to train the V neural network.

[0106] In a large-scale network environment, directly using a single experience replay pool may cause interference among the experiences of different subnets or nodes, thereby reducing the training efficiency. To address this issue, the experience replay pool can be divided into multiple regional replay pools (each region corresponding to a subnet or node), so as to store and utilize experiences more specifically. In a complex network environment, the dynamics of each subnet may not be synchronized. For example, some subnets may change frequently, while others are relatively stable. By storing the experience pool in regions, it is also possible to prevent the negative impact of this non-stationarity on the model and effectively maintain the stability of training.

[0107] Setting up prioritized experience replay requires five major steps: dividing the experience pool, calculating priorities, sampling within regions, cross-regional sampling, and correcting importance sampling weights. The detailed design is as follows:

[0108] Dividing the experience pool: According to the network topology, divide the experience pool into multiple sub-pools. Each sub-pool specifically stores the experiences from a specific subnet or node, avoiding the mixing of experiences from different regions.

[0109] Priority calculation: Assign a priority to each experience, usually calculated based on the temporal difference (TD) error. The larger the TD error, the greater the potential for policy improvement, and thus the higher the priority. The formula is as follows:

[0110]

[0111] Sampling within regions: During each policy update, the SAC agent can sample experiences from the replay pool related to the current operating region. For example, if the agent is currently in subnet A, it will preferentially sample from the experience pool of subnet A. This targeted sampling method can improve the learning efficiency of the model because the sampled experiences are more relevant and can optimize the policy within this region faster. During the sampling process, it is easier to select experiences with higher priorities. This can be achieved through probability sampling, where the sampling probability P(i) is set in proportion to the priority pi. For example, the sampling probability P(i) is defined as:

[0112]

[0113] where pi is the priority of experience pool i, and α controls the level of priority. When α = 1, sampling is completely priority-based; when α = 0, sampling is uniform.

[0114] Cross-regional sampling: In some update steps, sample from the experience pools of different regions to enhance the model's adaptability to the global network. This approach balances the relationship between regional focus and global exploration, enabling the model to not only focus on optimizing the policy in a specific region but also understand the dynamics of other regions, enhancing the generalization ability.

[0115] Importance Sampling Weight Correction: To correct sampling bias, lower weights are given to experiences with high priority during gradient calculation. The importance sampling weight w(i) is used to balance this bias:

[0116]

[0117] where N is the size of the experience buffer, and β is a parameter that adjusts the degree of importance sampling correction. When β = 1, the correction is fully applied. Usually, β starts from 0 and then gradually increases to avoid instability caused by high weights in the early stage of training.

[0118] The process of the improved ASAC algorithm is Algorithm 1, and the specific implementation is as Figure 4 shown.

[0119] Lines 1 - 4: Initialize the critic, actor, and target network parameters, as well as the regional experience replay buffer and temperature parameter.

[0120] Line 5: Start the loop for each episode, sample the initial state, and determine the region. Lines 7 - 10: At each time step, sample an action, execute it, observe the result, and store the experience in the respective regional replay buffer.

[0121] Lines 11 - 19: Perform the training step by sampling a mini - batch with prioritized experience replay, calculating the TD error and priorities, updating the critic and actor networks, and adjusting the temperature parameter.

[0122] Line 20: Update the target network using the soft update mechanism.

[0123] The following further illustrates the detailed steps with examples, as Figure 2 shown. The example of the present invention provides an improved method for deep reinforcement learning algorithm in a network environment scenario, including the following steps:

[0124] Construct a network information environment, and collect the environmental state based on this network topology environment and node information, including the node connection topology matrix, firewall configuration between hosts or subnets, host operating system configuration, process information, protocol information, etc.

[0125] Construct an actor - critic structure with independent policy and value function networks, a heteropolicy formula that can reuse previously collected data to improve efficiency, and entropy maximization that can achieve stability and exploration.

[0126] Construct a prioritized replay experience pool to increase the probability of useful experiences in the case of sparse rewards and accelerate the algorithm convergence performance.

[0127] Specifically, the improved deep reinforcement learning algorithm is as follows:

[0128] Set the environment and basic parameters, including the environment object env providing information on the state space and action space. The basic parameters include the learning rates (actor_lr, critic_lr, alpha_lr), discount factor (gamma), target entropy value (target_entropy), soft update coefficient (tau), and random seed (seed), etc.

[0129] Instantiate the policy network (actor) and value networks (critic_1, critic_2). The policy network is used to output the action selection probability in the current state, generating the probability distribution of each possible action through the Softmax function; the parameters of the two target value networks (target_critic_1, target_critic_2) are initialized the same as the prediction networks and are synchronously updated regularly through soft update. Instantiate the optimizers, setting independent Adam optimizers for the policy network and value networks respectively, and initializing the optimizer for the entropy coefficient (alpha).

[0130] The present invention provides a method for designing the reward function of the deep reinforcement learning algorithm based on the above first aspect, including the following steps:

[0131] Construct the reinforcement learning process as a Markov decision process (MDP) described by the quadruple <S, A, R, P>. Among them, S represents a set of penetration states observed by the agent, such as network topology, host information, and vulnerability details. These states correspond to different stages in the penetration testing process; A represents a set of possible attack actions corresponding to the above different state stages. The operation design includes network scanning, host scanning, vulnerability exploitation, privilege escalation, and other operations available to the agent. R is the reward function R(s), which assigns rewards according to different penetration states. For example, obtaining the highest level of permission on a sensitive host may generate a reward of 100, while compromising an intranet subnet host may obtain a reward of 1, and compromising a DMZ host will not receive any reward. The attack state transition probability P is used in the network penetration testing modeling to describe the possibility of the attacker (or agent) transitioning between different penetration states.

[0132] For the MDP process discussed above, the penetration testing uses multiple attack methods to penetrate the network. Assume that the threat of a certain type of attack is Xi, the cost of launching this type of attack is Ai, the attack cost of node i is Ni, and the attack value is Vi. Then the attack cost in the state of attacking a certain node j. Assume that on the optimal attack path, the state of node i relative to the attacker is S (0 represents no attack action, 1 represents successful attack action, 2 represents failed attack action).

[0133] There is an equation:

[0134] C = min(C) = ∑(ai + ni)*si (i < j) (C = LC)

[0135] LC (Limit Cost) is the established cost. Assuming that the resources held by the attacker can consume a cost of at most LC, it can help the intelligent agent to accelerate the convergence speed during the network search attack.

[0136] The damage W caused by such an attack to the network is W = max(W) = ∑( xi + vi )*pi

[0137] The attack reward function in the current state: R = W - C = ∑( xi + vi )*pi - ( ai + ni )* si (i < j), where pi is the success probability of node attack, and the attack threat degree is related to the corresponding service vulnerability scoring system CVSS of different nodes in different networks.

[0138] To simulate a real observation-based penetration process, it is assumed that the topology and host information must be obtained through the feedback of scanning or attack actions. Therefore, the information collection phase is carried out in 4 steps: (1) subnet scan (subnet_scan) to discover all hosts within the subnet; (2) operating system scan (os_scan) to obtain the operating system type of the target host; (3) service scan (service_scan) to obtain the service type of the target host; (4) process scan (process_scan) to obtain the process information on the target host. By executing the scan operation, the PT Agent can obtain the corresponding host information as the observation state.

[0139] The actions of VE and the actions of PE need to be executed according to specific requirements. In addition, according to the Common Vulnerability Scoring System (CVSS), the corresponding success probability is set to simulate the uncertainty of attacks in reality. By configuring different VE actions and PE actions, the PT agent can be modeled with different attack capabilities.

[0140] The network topology of the simulation experiment in the embodiment of the present invention has a total of 6 subnets. Each subnet is set between the firewall and the access protocol. The yellow host is the sensitive host in the network. The attacker has mastered two available nodes (1,0), (6,0) and 4 scanning means, and finally obtains the permissions and relevant sensitive information of the sensitive nodes through actions such as VE and PE. The host configuration is shown in Table 1.

[0141] Table 1

[0142]

[0143]

[0144] Analysis of Parameter Tuning Results: First, a detailed parameter tuning of the ASAC algorithm was carried out, and a total of four groups of parameters were set. The goal of parameter tuning is to find a set of optimal hyperparameters to ensure the stability and effectiveness of the algorithm in the network environment. A combination of grid search and random search was used to explore different combinations of hyperparameters, including learning rate, discount factor, batch size, etc. Through multiple experiments, a set of optimal hyperparameters was determined and verified in the network environment. The optimized ASAC algorithm has significantly improved in terms of convergence speed and final performance.

[0145] Analysis of Comparison Results of Different Algorithms: To verify the performance advantages of the improved ASAC algorithm in the network environment, the ASAC algorithm was compared with several classic network optimization algorithms, including Q-learning and DQN (Deep Q-Network), which have been studied in intelligent penetration testing.

[0146] According to the experimental data, the variation of the cumulative reward value per round of different algorithms with the number of training rounds was obtained. The learning goal of the agent is to learn to obtain the permissions of all sensitive hosts in the target network with fewer steps, so as to obtain the reward value. Therefore, the reward value can be used to measure the policy level of the agent. In the initial stage of training, the reward value that the agent can obtain per round is very small, but as the training progresses, the reward value continuously increases, indicating that the agent gradually learns the strategy to obtain the maximum reward value. In the early exploration process, the cumulative reward value per round of the ASAC algorithm is significantly lower than that of the DQN and Q-learning algorithms. This is because in the early exploration process of the ASAC algorithm, the agent has no experience to rely on and can only use a random strategy for exploration. It also wastes a large number of attack steps. However, as the training progresses, Q-learning cannot obtain better results due to the explosive growth of the Q-table, and the convergence speed and effect of DQN are not as good as those of the ASAC algorithm. Finally, both the ASAC algorithm and the DQN algorithm converge to a stable cumulative reward value within 600 rounds, and the ASAC algorithm has the best performance and can converge to the optimal value within 300 rounds.

[0147] The experimental results show that the improved ASAC algorithm has obvious advantages in the network environment, including convergence speed, decision-making quality, and resource utilization. The ASAC algorithm is superior to other algorithms in all metrics, especially in dealing with large-scale networks and highly dynamic environments.

[0148] Analysis of experimental results in scenarios of different structures and scales: To verify the scalability of the ASAC algorithm, experiments were conducted in network environments of different scales. The experiments were carried out in a network simulation environment, simulating network topologies of different scales and complexities. Table 3 shows the configurations of five different-scale network environments used in the experiments.

[0149] Table 3

[0150]

[0151] The scale of the experimental network varied from 8 to 64 nodes, covering different scenarios from small local area networks to large wide area networks. The main metrics we focused on were the scores and convergence speeds of the ASAC algorithm under different network conditions. In three heterogeneous networks A, B, and C with the same sensitive hosts and slightly different aggregation points and subnets, the ASAC algorithm model had good robustness. In three similar scenarios, its average cumulative reward value converged to the optimal value within 400 rounds.

[0152] As the number of subnets and hosts increased, the convergence speed of the algorithm slowed down, and the number of hosts and actions that the agent needed to try per round also increased rapidly. From the performance of the ASAC algorithm under different network scales, it can be seen that as the network scale increased, the performance of the ASAC algorithm decreased slightly, but it could still effectively handle complex problems in a large-scale network environment. In a simulated network environment with 40 hosts, the ASAC algorithm could also converge to the optimal solution. The experimental results showed that when the network scale gradually increased, the ASAC algorithm could still maintain good performance. In addition, the ASAC algorithm also performed well in terms of resource utilization and could maintain a high resource utilization rate under different network scales. When dealing with complex network topologies and high loads, the ASAC algorithm could make effective decisions to ensure the stability and efficiency of the network.

[0153] In summary, the embodiment of the present invention is based on the improved ASAC algorithm, which greatly improves the utilization rate of effective experience by designing a prioritized experience replay pool, improves the convergence speed, and experimentally proves that it is superior to model-free deep RL methods. This improved solution shows stronger adaptability and efficiency in the field of network attack and defense. Especially when facing complex dynamic environments and large-scale networks, it can find the optimal penetration path faster, improving the automation and intelligence level of penetration testing.

[0154] The technical key points of the embodiments of the present invention lie in the improved SAC algorithm structure, including the design of an independent policy network and a value function network, the application of a heteropolic formula, the construction of a prioritized replay experience pool, and the introduction of an entropy maximization method. In addition, through the design of a reward function that combines attack benefits and costs, the advantages and disadvantages of different attack paths are accurately reflected. The above technologies are combined with the defense mechanism in the real network environment to provide the agent with efficient learning and decision-making capabilities. The technical solution can effectively improve the application efficiency of deep reinforcement learning in network security penetration testing, and has broad promotion value and practical significance.

[0155] First Embodiment of the Device

[0156] According to an embodiment of the present invention, there is provided an intelligent penetration testing device based on deep reinforcement learning. Figure 5 It is a schematic diagram of the intelligent penetration testing device based on deep reinforcement learning according to an embodiment of the present invention, as Figure 5 shown. The intelligent penetration testing device based on deep reinforcement learning according to an embodiment of the present invention specifically includes:

[0157] A construction module 90, configured to construct a network information environment and set basic parameters, initialize the network information environment, and construct an actor-critic structure with an independent policy network and a value network; model the Markov decision process of penetration testing as a quadruple <S, A, R, P>, and construct a Markov model, where the goal of the Markov model is to select an optimal penetration strategy so as to obtain an optimal attack action under a given penetration state for maximizing the reward. Here, S represents a set of observed penetration states, A represents a set of possible attack actions corresponding to different penetration states, R represents a reward function for allocating rewards according to different penetration states, and P is the attack state transition probability, which represents the possibility of the attacker transitioning between different penetration states; construct a prioritized replay experience pool;

[0158] A penetration testing module 92, configured to perform penetration testing based on the Markov model, calculate and output an attack action to the network information environment based on the current penetration state through the policy network to obtain the next penetration state, reward the attack action through the value network, save all attack actions and environment feedback data in the prioritized replay experience pool, calculate and determine the sampling method for preferential selection according to the priority of the experience, so as to train samples with high priority more frequently, repeat the penetration testing operation of the above Markov model until the function converges, output the optimal penetration strategy, generate a test case report according to the optimal penetration strategy, and obtain a penetration testing report according to the test case report.

[0159] The embodiments of the present invention are device embodiments corresponding to the above method embodiments. The specific operations of each module can be understood with reference to the description of the method embodiments and will not be elaborated here.

[0160] Device Embodiment Two

[0161] An embodiment of the present invention provides an electronic device, such as Figure 6 shown, including: a memory 100, a processor 102, and a computer program stored on the memory 100 and executable on the processor 102. When the computer program is executed by the processor 102, the steps described in the method embodiments are implemented.

[0162] Device Embodiment Three

[0163] An embodiment of the present invention provides a computer-readable storage medium. An implementation program for information transmission is stored on the computer-readable storage medium. When the program is executed by the processor 102, the steps described in the method embodiments are implemented.

[0164] The computer-readable storage medium described in this embodiment includes, but is not limited to, ROM, RAM, magnetic disk, optical disk, etc.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent penetration testing method based on deep reinforcement learning, characterized in that, Including: Construct a network information environment and set basic parameters, initialize the network information environment, and construct an actor-critic structure with an independent policy network and value network; Model the Markov decision process of penetration testing as a quadruple <S, A, R, P>, and construct a Markov model. The goal of the Markov model is to select the optimal penetration strategy to obtain the optimal attack action under a given penetration state in order to maximize the reward. Here, S represents a set of observed penetration states, A represents a set of possible attack actions corresponding to different penetration states, R represents a reward function for allocating rewards according to different penetration states, and P is the attack state transition probability, which represents the possibility of the attacker transferring between different penetration states; Construct a prioritized replay experience pool; Based on the Markov model, conduct penetration testing. Based on the current penetration state through the policy network, calculate and output an attack action to the network information environment to obtain the next penetration state. Reward the attack action through the value network, save all attack actions and environmental feedback data in the prioritized replay experience pool, calculate and determine the sampling method based on the priority of the experience to train samples with higher priority more frequently. Repeat the above penetration testing operation of the Markov model until the function converges, output the optimal penetration strategy, generate a test case report according to the optimal penetration strategy, and obtain a penetration test report according to the test case report.

2. The method according to claim 1, wherein: The network information environment specifically includes: network topology environment, node information collection environment state, node connection topology matrix, firewall configuration between hosts or subnets, host operating system configuration, process information, and protocol information; The basic parameters specifically include: learning rate, discount factor, target entropy value, soft update coefficient, and random seed; There is one policy network and four value networks; The penetration states specifically include: initial unpenetrated state, external network information collection state, boundary breakthrough state, internal network penetration state, privilege escalation state, critical target access state, defense detection and blocking state, network environment dynamic change state, penetration failure state, and penetration success state; The attack actions specifically include: network scanning, host scanning, vulnerability exploitation, and privilege escalation.

3. The method according to claim 1, wherein Rewarding the attack action through the value network specifically includes: Let the threat of a certain type of attack be xi, the cost of launching this type of attack be ai, the attack cost of node i be ni, and the attack value be vi. Then calculate the attack cost when attacking a certain node j state. Assume that on the optimal attack path, the state of node i relative to the attacker is si, then there is formula 1; C = min(C) = ∑(ai + ni) * si Formula 1; Where, i is less than j, C is less than or equal to LC, LC is the established cost, that is, the maximum cost that the attacker's resources can consume; C represents the attack cost when attacking a certain node j state; Calculate the damage W caused by this type of attack to the network according to formula 2; W = max(W) = ∑( xi + vi ) * pi Formula 2; Among them, pi is the node attack success probability, and its value is estimated by statistically analyzing historical attack data, simulating experimental results, or the exploration feedback of reinforcement learning; Calculate the attack reward function under the current penetration state according to Formula 3; R = W - C = ∑( xi + vi ) * pi - ( ai + ni ) * si Equation 3.

4. The method according to claim 1, characterized in that Constructing a Markov model specifically includes: The goal of building a Markov model is to select the optimal policy π such that, given the state st, the optimal action at can be obtained to maximize the reward, i.e., at = π( st ), to maximize the long-term cumulative reward G(s0) in the current state, as shown in Equation 4: Among them, γ ∈ (0, 1) represents the discount factor, which is used to balance the importance of the current reward and the future reward. E{}· represents the expectation operator, and T represents the total number of steps in a round. In the reinforcement learning task with a limited number of steps, the interaction process of the agent ends after T steps. t represents the current step of the agent's interaction with the environment, starting from 0.

5. The method according to claim 1, characterized in that, Constructing an actor-critic structure with independent policy network and value network specifically includes: Determine the update loss expression of the policy network according to Formula 5 and Formula 6: Among them, represents the prioritized replay experience pool, and α is the entropy coefficient, which is used to control the importance of the entropy term lnπ(a′ t |s t ; θ), and its significance increases with the increase of α. st represents the state of the agent at the t-th step of the episode, θ represents the parameters of the policy network, at represents the actually executed action, which comes from historical experience and is used for the update of the Q function, rt represents the reward function value in the state st, and a′ t represents the action resampled from the current policy π(.| st ; θ) during policy evaluation or improvement, represents sampling the action at′ from the current policy π(.| st ; θ) and calculating the expectation. q0(s t , a′ t ) represents the Q function, which is used to estimate the expected cumulative reward that can be obtained after taking the action at′ in the state st, represents the set of all possible actions a in the state st, and w (0) represents the parameters of the Q function. Entropy represents a measure of the randomness degree of a random variable. If X is a random variable with the probability density function p, then its entropy H(X) is defined as: where, denotes taking the expectation of the random variable X under the probability distribution p(x), and p(x) represents the probability density function of the random variable X.

6. The method according to claim 1, characterized in that, Constructing a prioritized replay experience pool specifically includes: Divide the prioritized replay experience pool into multiple regional replay pools. Among them, each regional replay pool corresponds to a subnet or node; Assign a priority to each experience, which is calculated according to the temporal difference TD error. The larger the TD error, the greater the potential for policy improvement and the higher the priority.

7. The method according to claim 1, wherein Save all attack actions and environmental feedback data in the prioritized replay experience pool, and calculate and determine the sampling method of preferential selection according to the priority of the experience specifically includes: During each penetration policy update, based on Formula 8, sample experiences from the regional replay pool related to the current operation area, or sample experiences from the regional replay pools of different regions; where P(i) is the sampling probability, p i represents the priority of experience i, and α is used to control the level of priority. When α = 1, sampling is completely priority-based according to priority; when α equals 0, sampling is uniform; Correct the sampling bias, give a lower weight to high-priority experiences in gradient calculation, and calculate the importance sampling weight w(i) according to Formula 9 to balance the sampling bias: Among them, N is the size of the experience buffer pool, and β is a parameter for adjusting the importance and the degree of sampling correction. When β = 1, the correction is fully applied, and β gradually increases from 0.

8. An intelligent penetration testing device based on deep reinforcement learning, characterized in that, Include: A construction module for constructing a network information environment and setting basic parameters, initializing the network information environment, and constructing an actor-critic structure with independent policy network and value network; modeling the Markov decision process of penetration testing as a quadruple <S, A, R, P>, constructing a Markov model, where the goal of the Markov model is to select the optimal penetration strategy to obtain the optimal attack action under a given penetration state for maximizing the reward. Among them, S represents a set of observed penetration states, A represents a set of possible attack actions corresponding to different penetration states, R represents a reward function for allocating rewards according to different penetration states, and P is the attack state transition probability, which represents the possibility of describing the attacker's transition between different penetration states; constructing a prioritized replay experience pool; The penetration testing module is used to perform penetration testing based on the Markov model. Based on the current penetration state through the policy network, it calculates and outputs attack actions to the network information environment to obtain the next penetration state. It rewards the attack actions through the value network, saves all attack actions and environmental feedback data in the prioritized replay experience pool, calculates and determines the sampling method for priority selection according to the priority of the experience, so as to train the samples with high priority more frequently. Repeat the penetration testing operation of the above Markov model until the function converges, output the optimal penetration strategy, generate a test case report according to the optimal penetration strategy, and obtain a penetration testing report according to the test case report.

9. An electronic device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the intelligent penetration testing method based on deep reinforcement learning according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, An implementation program for information transmission is stored on the computer-readable storage medium. When the program is executed by the processor, it implements the steps of the intelligent penetration testing method based on deep reinforcement learning according to any one of claims 1 to 7.

Citation Information

Cited By

  • Automatic penetration testing method and system based on cognitive decision model

    CN121615150A

  • Power system network intelligent penetration test method based on deep reinforcement learning

    CN121690739A