Penetration test path planning method based on ant colony and reinforcement learning

By combining ant colony algorithm and reinforcement learning network, the penetration test path planning is automated, and the problems of time-consuming and labor-intensive and improper path selection are solved, and the efficiency and accuracy of the test are improved.

CN120030544APending Publication Date: 2025-05-23CHINA ACAD OF SPACE SYST SCI & ENG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411917029.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing penetration testing methods rely on manual operations, are time-consuming and labor-intensive, and are prone to improper path selection or missing important vulnerabilities in complex network environments, resulting in incomplete and inaccurate test results.

Method used

The penetration test path planning method based on ant colony and reinforcement learning is adopted, and the path planning is automated by building an environmental model, using the ant colony algorithm to explore the initial path, and combining the reinforcement learning network to dynamically adjust the action selection.

Benefits of technology

It improves the degree of automation of penetration testing, enhances the comprehensiveness of path discovery, and can find effective penetration testing paths faster in complex network environments, improving the efficiency and accuracy of testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030544A_ABST
    Figure CN120030544A_ABST
Patent Text Reader

Abstract

The invention discloses a penetration test path planning method based on ant colony and reinforcement learning, and belongs to the technical field of information security testing. According to the method, an initial global path is quickly generated through an ant colony algorithm, then path planning is dynamically optimized and adjusted by utilizing a reinforcement learning method, dynamically adjusting action selection exploration factors and adopting a multi-step return priority experience sequence playback method, and efficient and economical automatic penetration test path planning is formed. According to the method, the intelligent level of path planning can be improved, continuously changing network security challenges can be effectively handled, and powerful support is provided for network security protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a penetration test path planning method based on ant colonies and reinforcement learning, and belongs to the technical field of information security testing. Background Art

[0002] With the rapid development of information technology, network security issues are becoming increasingly serious. Penetration testing, as a proactive defense method, helps organizations identify potential vulnerabilities in network systems and has become an important tool for evaluating and strengthening network security. Traditional penetration testing methods mostly rely on manual operations, which are not only time-consuming and labor-intensive, but also easily limited by the professional capabilities of testers. In addition, when facing complex network environments, it is easy to make improper path selections or miss important vulnerabilities, resulting in incomplete and inaccurate test results. Most of the existing path planning methods are based on attack tree models or attack graph models, but these two methods have obvious limitations when dealing with complex attack steps and state transition relationships. Therefore, how to improve the automation of penetration testing, reduce costs and improve efficiency has become an urgent problem to be solved.

[0003] Chinese patent CN118410497A discloses an intelligent penetration test method and system based on deep learning, which collects system business logs and performs deep learning analysis to intelligently select the best penetration method for penetration testing. This method relies on a large amount of high-quality business log data for training and classification. If the data is incomplete or noisy, it may affect the accuracy and effectiveness of the model. Chinese patent CN117614728A discloses an intelligent penetration test method for dynamic network environments, which improves the efficiency of penetration testing by constructing a dynamic network environment and using a Markov model to optimize the penetration path. However, in a dynamic environment, the lack of a real-time feedback mechanism may cause the test results to lag. Chinese patent CN116595536A discloses a penetration test path planning method based on the A3C model, which uses an improved A3C model to generate an effective penetration test path by analyzing vulnerability data and topological attack trees, which can improve the success rate and efficiency of penetration testing. However, the lack of adaptability to different network architectures and configurations limits its versatility. Summary of the invention

[0004] The technical problem solved by the present invention is: to overcome the shortcomings of the prior art, and to propose a penetration test path planning method based on ant colony and reinforcement learning, which can realize efficient automated penetration test path planning for active network defense.

[0005] The technical solution of the present invention is:

[0006] A penetration test path planning method based on ant colony and reinforcement learning, comprising:

[0007] S1: Build a penetration test environment model to simulate real network attack scenarios;

[0008] S2: Based on the environment model of the penetration test, the state space and action space of the path planning problem are established, and the ant colony algorithm is used to explore the initial path to obtain a set of optimal initial paths;

[0009] S3: Initialize the reinforcement learning network, use the best initial path as the input of the reinforcement learning network, select and execute the corresponding action according to the training results, dynamically adjust the action selection exploration factor according to the feedback, and continuously replay the experience sequence of multi-step rewards to update the parameters of the reinforcement learning network;

[0010] S4: Establish an environmental model according to the real network environment, and obtain a set of optimal initial paths through the ant colony algorithm based on the path planning problem defined in step S2; input the optimal initial paths into the reinforcement learning network after updating the parameters to obtain the execution results of each initial path; screen the different effective and feasible paths for comprehensive performance evaluation and select the best attack path.

[0011] The advantages of the present invention compared with the prior art are:

[0012] (1) The present invention utilizes the global search capability of ant colony optimization, which can effectively avoid local optimal solutions, enhance the comprehensiveness of path discovery, and help the system find effective penetration test paths more quickly in complex network environments.

[0013] (2) During the training process, the algorithm adopted by the present invention dynamically adjusts the value of the action selection exploration factor c according to the learning effect and environmental changes, thereby adjusting the balance between exploration and utilization and improving the effectiveness of decision-making.

[0014] (3) The present invention can more comprehensively evaluate the long-term impact of actions and improve the optimization effect of strategies by using multi-step returns in the priority experience sequence playback stage. This method enables the model to better capture time dependencies when dealing with complex sequential decision problems.

[0015] (4) The algorithm adopted by the present invention shows good adaptability in different penetration testing scenarios and can handle dynamically changing network environments, ensuring its effectiveness in various application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0017] Figure 1 This is a flow chart of a penetration test path planning method based on ant colony and reinforcement learning according to an embodiment of the present invention;

[0018] Figure 2 This is a specific structural diagram of the state space of an embodiment of the present invention;

[0019] Figure 3 A flowchart of an ACO algorithm implemented in an embodiment of the present invention;

[0020] Figure 4 is an overall flow chart of the reinforcement learning process of an embodiment of the present invention;

[0021] Figure 5 This is a flow chart of the priority experience sequence playback phase of multi-step reporting in an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0023] The present invention proposes a penetration test path planning method based on ant colony and reinforcement learning, such as Figure 1 As shown, the following steps are included:

[0024] S1: Environment Generation

[0025] Environment generation is the first step in the study of penetration test path planning methods. By constructing a complex and dynamic penetration test environment network model and simulating real network attack scenarios, it helps researchers evaluate the effectiveness and performance of the algorithm under different conditions and provides infrastructure for subsequent testing and verification.

[0026] Environment generation includes network construction and host and vulnerability configuration. Network construction needs to describe the layout of network nodes (such as servers, routers, etc.) and their connection methods, and display the network topology. Host configuration includes address identification and service information. Vulnerability configuration requires setting or identifying and simulating known security vulnerabilities on the host. Through these configurations, a foundation is provided for the construction of the penetration test environment to ensure the effectiveness and comprehensiveness of the test. Testers can simulate real attack scenarios and evaluate the security of the network.

[0027] S2: Data preprocessing

[0028] Data preprocessing is a key step to ensure the validity and availability of input. It can provide a reliable foundation for subsequent model training and testing, ensuring that the algorithm can run on high-quality data, thereby improving the accuracy and efficiency of predictions.

[0029] This step mainly includes data cleaning, feature extraction and data normalization. Data cleaning ensures the accuracy and reliability of data through steps such as deduplication, missing value processing, outlier detection and consistency check; through the extraction of network features and vulnerability features, it can provide the necessary information basis for penetration test path planning. Data normalization ensures that feature data is processed at a unified scale, so that the algorithm can have better performance and faster convergence speed, thereby improving the training efficiency of the model and the accuracy of prediction. The process of data preprocessing can not only improve the reliability of path planning, but also enhance the adaptability of the model in complex network environments.

[0030] S3: Ant Colony Search

[0031] The ant colony algorithm is used to explore the initial path. After initializing the ant colony information, the initial path is constructed; at the end of each round, the path constructed by each ant is evaluated and its fitness is calculated; then the pheromone concentration is updated according to the fitness of each path.

[0032] S4: Reinforcement Learning Process

[0033] Reinforcement learning is the core of the system function, ensuring that the penetration test can be carried out efficiently and accurately. Receive the initial path P generated by the ant colony search phase initial As input, after initializing the network and experience buffer pool, the corresponding action is selected and executed according to the training results, and the action selection exploration factor c is dynamically adjusted according to the feedback. At the same time, the utilization of training data is improved by replaying the priority experience sequence of multi-step returns.

[0034] (1) Action selection with dynamically adjustable parameters: Decisions are made based on the results of ant colony search and the action selection strategy. A set of executable actions is extracted from the path information obtained during the ant colony search phase; the UCB value of each action is calculated, and the UCB algorithm is used to balance exploration and utilization to select actions; the selected actions are executed.

[0035] (2) Priority experience sequence replay for multi-step rewards: After each action is executed, the obtained experience (current state, action, reward, next state) is stored in the experience replay pool; samples are extracted from the experience replay pool, and the priority of each experience is calculated based on the temporal difference error (TD error) of the experience; based on the calculated priority, experience sequences are sampled from the experience replay pool to update the multi-step rewards; at the end of each training cycle, the mean and standard deviation of the reward are calculated, and the exploration factor c is dynamically adjusted.

[0036] S5: Result Generation

[0037] The paths generated in each exploration are recorded, and these paths are evaluated to confirm their effectiveness and feasibility in the real network environment. Multiple valid paths can be selected. Further optimization and integration are carried out by comparing the performance of different paths (such as execution time, success rate, etc.) to ensure that the best path is selected. A comprehensive risk assessment is then performed to ensure that the path will not trigger too many security alerts or cause unnecessary damage when executed, forming the final attack path.

[0038] The present invention will be further described below by specific embodiments:

[0039] S1: Environment generation, including:

[0040] S1.1: Define the overall architecture of the network. Select the main components of the network, including the external network, firewall, DMZ (demilitarized zone), and multiple internal networks.

[0041] S1.2: Draw a network topology diagram to visualize the network structure, and use the adjacency matrix M to record the connection relationship between nodes. Each element of the matrix represents a node in the network.

[0042] S1.3: Set up the segmented structure of the network and determine the number and size of subnets.

[0043] S1.4: Set connectivity between subnets and use firewall rules to manage traffic in and out of each subnet.

[0044] S1.5: Determine the number of hosts in each subnet.

[0045] S1.6: Assign a unique IP address to each host to ensure there are no address conflicts.

[0046] S1.7: Install the appropriate operating system and configure necessary services and applications for each host.

[0047] S1.8: Query the vulnerability database to screen out vulnerabilities related to the host operating system and services and determine the scope and exploitation methods of each vulnerability.

[0048] S1.9: Simulate known vulnerabilities on the host to ensure that the configured vulnerabilities can be effectively exploited.

[0049] S1.10: Configure the vulnerable service or application and then introduce available attack tools.

[0050] S1.11: Create an attribute record for each node containing the specific information of the network environment described above. Combine the network topology and the attribute record of each node to generate the operating environment of the algorithm.

[0051] S1.12: Store the environment information in a configuration file in YAML (YAML Ain't Markup Language) format and upload it to the NASim simulator.

[0052] S2: Data preprocessing, including:

[0053] S2.1: Check the adjacency matrix and remove redundant connections to improve data accuracy; identify isolated nodes and choose to retain or delete them as needed.

[0054] S2.2: Fill or delete missing host attributes. If there are few missing values, you can delete the records containing missing values. Otherwise, you can use the filling method to fill the missing values ​​with the mean, median or mode, or fill them with more complex methods such as interpolation.

[0055] S2.3: Detect outliers, identify unreasonable host configuration data, and delete or replace them with reasonable values ​​to improve the overall quality of the dataset.

[0056] S2.4: Check the consistency of data field types and formats to ensure that units are unified to avoid confusion and improve data usability.

[0057] S2.5: Remove outdated or irrelevant vulnerability information and unify the representation of vulnerability risk levels.

[0058] S2.6: Use network scanning tools to collect node information and extract data features. Integrate the extracted features into a unified feature set to ensure that each sample has a complete feature description and organize them into a structured data set.

[0059] S2.7: Encode the extracted features (such as one-hot encoding or label encoding) to adapt to the input requirements of the model.

[0060] S2.8: Remove redundant or irrelevant features through correlation analysis or other methods to improve the training efficiency and effectiveness of the model.

[0061] S2.9: Select an appropriate normalization method to scale the feature data to a uniform standard range to avoid the impact of feature value differences on model training.

[0062] S2.10: Normalize each feature to ensure that all feature values ​​are on the same scale.

[0063] S2.11: Check the normalized data distribution to ensure that there are no anomalies and it is suitable for model input. The preprocessed network environment is used as the environment for subsequent algorithm operation.

[0064] S3: Ant Colony Search

[0065] First, a series of initialization tasks need to be performed, including initializing the ACO algorithm parameters and defining the state space and action space of the path planning problem. The state space includes all important information in the penetration test environment. The state s can be defined by a two-dimensional array, with a specific structure such as Figure 2 As shown. The first dimension of the array represents different hosts in the network, and the second dimension represents the status of each host, including the subnet address to which the host belongs, the host number, related additional information, and information such as the services and processes being run. The additional information includes whether the host has been compromised, whether it is reachable, whether it has been discovered, the host value, and the access permission level. The action space includes all possible operations, such as vulnerability exploitation, privilege escalation, service scanning, operating system scanning, subnet scanning, process scanning, and no actual action. The action space can be defined as a discrete set. For example: A = {vulnerability exploitation, privilege escalation, service scanning, operating system scanning, subnet scanning, process scanning, no actual action}, stored in an array, when an operation is performed, the element at that position is set to 1, otherwise it is set to 0.

[0066] like Figure 3 As shown, the specific steps in the initial path generation phase of executing the ACO algorithm are as follows:

[0067] S3.1: Initialize each bit of the array representing the action space to 0, and define the state space of the path planning problem. The services being run by the host are represented by an array with a length equal to the number of services. Each bit in the array corresponds to a service. When the service is running, the corresponding position is 1, otherwise it is 0. The representation method of the process is the same as that of the service. Through automatic scanning by the tool, a mapping relationship between the host and the running services and processes is established.

[0068] S3.2: Based on the benefits and costs after successfully executing the action, a mapping between different actions and corresponding rewards is established and stored in a dictionary.

[0069] S3.3: Design a reward system for security testing actions based on the benefits and costs of successfully executing the action.

[0070] S3.4: Initialize the pheromone matrix (a two-dimensional matrix) and set the same initial pheromone concentration τ on all edges ij =1.0(τ ij is the pheromone concentration from node i to node j), ensuring that all paths are equally likely without any prior knowledge. The threshold of pheromone concentration change is set to 0.001 as the termination condition.

[0071] S3.5: Combine the pheromone concentration and heuristic information to calculate the selection probability of each possible action. The specific calculation formula is:

[0072]

[0073] z i =α·log(τ i )+β·log(η i ) (2)

[0074] Where P(i) is the probability of choosing action i.

[0075] z i is the score associated with action i, z j represents all possible actions (including i).

[0076] τ i is the pheromone concentration associated with action i.

[0077] η i is the heuristic information associated with action i (e.g., distance or risk).

[0078] S3.6: Based on the calculated selection probabilities, select the option with a higher P(i).

[0079] S3.7: Record the path of each ant in the form of a four-tuple, including the action and the reward obtained, that is, (node ​​n1, node n2, action a, reward r).

[0080] S3.8: After all ants have completed the path, update the pheromone value based on the total reward for each path. Use the formula:

[0081] τ ij ←(1-ρ)τ ij +Δτ ij (3)

[0082] Among them, τ ij is the pheromone concentration from node i to node j, ρ is the volatility factor, Δτ ij Proportional to the quality of the path.

[0083] S3.9: If the change in pheromone concentration of all paths is less than the threshold value 0.001, the termination condition is met and the algorithm ends; otherwise, return to step five.

[0084] After completing multiple rounds of iterations, a set of optimal initial paths is obtained, defined as P initial . Each path path i is a sequence of four tuples, and the generated initial path P initial as input to the reinforcement learning process.

[0085] S4: Reinforcement Learning Process

[0086] The reinforcement learning process receives the initial path generated by the ant colony search phase as input, and its overall flow chart is as follows: Figure 4 After initializing the network and experience buffer pool, the corresponding action is selected and executed according to the training results. The action selection exploration factor c can be dynamically adjusted according to the feedback, and the multi-step reward experience sequence is continuously replayed, as shown in Figure 5 shown.

[0087] S4.1: First, initialize the Q network and define the structure of the input layer, hidden layer, and output layer. The number of nodes in the input layer matches the dimension of the state space, and the number of nodes in the output layer matches the dimension of the action space. Initialize the target Q network with the same structure as the Q network. Set the threshold of the cumulative reward to 1000.

[0088] S4.2: Initialize the experience replay pool R to be empty, which is used to store the experience tuples (s, a, r, s') of state transitions. Set the maximum capacity of the experience replay pool, for example, 1000 or 2000 experiences. If the capacity limit is reached, a first-in-first-out strategy can be used for replacement.

[0089] S4.3: For each optional action, calculate the UCB value using the formula:

[0090]

[0091] Among them, N a is the number of times action a is selected; c is the exploration factor, which controls the trade-off between exploration and exploitation.

[0092] S4.4: Select the action with the highest UCB value, execute the selected action, receive the reward r and the new state s' and store the experience (s, a, r, s') into the experience buffer pool R.

[0093] S4.5: Maintain a fixed-size window that stores the rewards for the past k time steps.

[0094] S4.6: Calculate the mean μ and standard deviation σ of the current reward as follows:

[0095]

[0096] Among them, r i It’s an instant reward for every step.

[0097] S4.7: Use the following formula to dynamically adjust the action selection exploration factor c.

[0098]

[0099] S4.8: At each time step, calculate the temporal difference error (TD Error) for each experience using the formula:

[0100]

[0101] Among them, r is the immediate reward obtained after performing action a in the current state s, reflecting the direct effect of the action.

[0102] γ is a discount factor, which indicates the importance of future rewards (the value range is (0,1), and the larger the value, the more attention is paid to future rewards).

[0103] Q(s',a') is the maximum Q value of all possible actions a' in the new state s'. It represents the best expected future reward that the agent can get in the new state.

[0104] Q(s,a) is the Q value of executing action a in the current state s, which represents the expected reward of the agent taking this action in the current state.

[0105] S4.9: Calculate the priority of empirical sampling based on TD error:

[0106] P=∑|TD_error i | (10)

[0107] S4.10: Sample the experience sequence from the experience buffer pool R according to the priority.

[0108] S4.11: Calculate multi-step returns:

[0109] R t =r t +γ·r t+1 +γ 2 ·r t+2 +...+γ n-1 ·r t+n-1 (9)

[0110] Among them, R t represents the n-step reward starting from time step t; the discount factor γ is used to reduce the impact of future rewards (the value range is [0,1)). t is the reward obtained by performing the action at time step t.

[0111] S4.12: Calculate the loss and update the network through the back-propagation algorithm. The loss function is calculated as:

[0112] Loss=(Q(s,a)-R t ) 2 (11)

[0113] S4.13: Update the parameters of the target network every F time steps (F is the target network update frequency, usually set to every few time steps, such as 100, 200 or 500 steps).

[0114] S4.14: Determine whether the reward target is achieved. If achieved, terminate the operation and record the obtained optimization path; otherwise, repeat S4.3 to S4.14.

[0115] S5: Result Generation

[0116] Result generation is an important part of the research on penetration test path planning methods. Its main function is to record and evaluate the paths generated during the exploration process and ensure the effectiveness and feasibility of these paths in the real network environment. By analyzing the performance of different paths, it provides a basis for optimizing and integrating paths, and finally forms the best attack path. The specific steps are as follows:

[0117] S5.1: Generate a simulation environment and run the algorithm according to the actual network environment requirements.

[0118] S5.2: Collect the results of path execution and evaluate the effectiveness and feasibility of each path in a real network environment by integrating indicators such as vulnerability coverage, path length, attack failure reasons, and resource consumption.

[0119] S5.3: Compare the comprehensive performance (including cumulative rewards, path length, risk level and key target coverage) of the different effective and feasible paths screened in S5.2, and select the best attack path.

[0120] The above-described embodiments are only preferred specific implementations of the present invention. Common changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.

Claims

1. A penetration test path planning method based on ant colony and reinforcement learning, characterized in that: include: S1: Build a penetration test environment model to simulate real network attack scenarios; S2: Based on the environment model of the penetration test, the state space and action space of the path planning problem are established, and the ant colony algorithm is used to explore the initial path to obtain a set of optimal initial paths; S3: Initialize the reinforcement learning network, use the best initial path as the input of the reinforcement learning network, select and execute the corresponding action according to the training results, dynamically adjust the action selection exploration factor according to the feedback, and continuously replay the experience sequence of multi-step rewards to update the parameters of the reinforcement learning network; S4: Establish an environmental model according to the real network environment, and obtain a set of optimal initial paths through the ant colony algorithm based on the path planning problem defined in step S2; input the optimal initial paths into the reinforcement learning network after updating the parameters to obtain the execution results of each initial path; screen the different effective and feasible paths for comprehensive performance evaluation and select the best attack path.

2. The penetration test path planning method based on ant colony and reinforcement learning according to claim 1 is characterized in that: Establish the state space and action space of the path planning problem, where: The state space includes all important information in the environment model. The state is defined by a two-dimensional array. The first dimension represents different hosts in the network, and the second dimension represents the state of each host, including the subnet address to which the host belongs, the host number, related additional information, and the running services and processes; the related additional information includes whether the host is compromised, reachable, discovered, host value, and access permission level; The action space includes all possible operations, including vulnerability exploitation, privilege escalation, service scanning, operating system scanning, subnet scanning, process scanning and no actual action; the action space is defined as a discrete set and stored in an array.

3. The penetration test path planning method based on ant colony and reinforcement learning according to claim 2 is characterized in that: According to the path planning problem, the ant colony algorithm is used to explore the initial path and obtain a set of optimal initial paths. The specific method is as follows: S3.1: Initialize each bit in the array representing the action space to 0; the services being run by the host are represented by an array with a length equal to the number of services, each bit in the array corresponds to a service, and the corresponding position is 1 when the service is running, otherwise it is 0; the representation method of the process is the same as that of the service; the mapping relationship between the host and the running services and processes is established through automatic scanning by the tool; S3.2: Based on the benefits and costs after successfully executing the action, a mapping between different actions and corresponding rewards is established and stored in a dictionary; S3.3: Initialize the algorithm parameters, pheromone matrix, and threshold value of pheromone concentration change, set the same initial pheromone concentration on all edges, and ensure that all paths are equally likely without any prior knowledge; S3.4: Combine pheromone concentration and heuristic information to calculate the selection probability of each possible action and select the action with the highest probability; record the path of each ant in the form of a four-tuple, including the action and the reward obtained, that is, (node ​​n1, node n2, action a, reward r); S3.5: After all ants have completed the path, update the pheromone value based on the total reward of each path; S3.6: If the change in pheromone concentration of all paths is less than the change threshold, the termination condition is met and a set of optimal initial paths is obtained, each of which is a four-tuple sequence; otherwise, return to S3.4 and continue iterating until the termination condition is met or the set iteration threshold is reached.

4. The penetration test path planning method based on ant colony and reinforcement learning according to claim 2 is characterized in that: Initialize the reinforcement learning network, take the best initial path as input, select and execute the corresponding action according to the training results, and dynamically adjust the action selection exploration factor according to the feedback, including: S4.1: Initialize the Q network and define the structure of the input layer, hidden layer, and output layer. The number of nodes in the input layer matches the dimension of the state space, and the number of nodes in the output layer matches the dimension of the action space. Initialize the target Q network, which has the same structure as the Q network. S4.2: Initialize the experience replay pool R to be empty, which is used to store the experience tuples (s, a, r, s') of state transitions, where s represents the current state, s' represents the new state after executing action a, and r represents the reward for executing action a; set the maximum capacity of the experience replay pool; S4.3: For each optional action, calculate the UCB value: Among them, N a is the number of times action a is selected; c is the exploration factor; S4.4: Select the action with the highest UCB value, execute the selected action, receive the reward r and the new state s' and store the experience (s, a, r, s') into the experience buffer pool R; S4.5: Maintain a fixed-size window, store the rewards of the past k time steps, and calculate the mean μ and standard deviation σ of the reward in the current window; S4.7: Dynamically adjust the action selection exploration factor c: In the formula, r i It’s an instant reward for every step.

5. The penetration test path planning method based on ant colony and reinforcement learning according to claim 4 is characterized in that: Methods for replaying experience sequences with multiple returns include: At each time step, the temporal difference error for each experience is calculated: Where γ is the discount factor; Q(s',a') is the maximum Q value of all possible actions a' in the new state s'; Q(s,a) is the Q value of executing action a in the current state s; r is the reward obtained after executing action a in the current state s; According to the TD obtained at each time step _error , calculate the priority of experience sampling; sample the experience sequence from the experience buffer pool R according to the priority, and calculate the multi-step return: R t =r t +γ·r t+1 +g 2 ·r t+2 +...+c n-1 ·r t+n-1 Among them, R t represents the n-step return starting from time step t; r t is the reward obtained by performing the action at time step t.

6. The penetration test path planning method based on ant colony and reinforcement learning according to claim 5 is characterized in that: Update the parameters of the reinforcement learning network, specifically: Calculate the loss and update the network through the back-propagation algorithm. The loss function is calculated as: Loss=(Q(s,a)-R t ) 2 Update the parameters of the target network every F time steps, where F is the target network update frequency.

7. The penetration test path planning method based on ant colony and reinforcement learning according to claim 1 is characterized in that: Construct an environment model for penetration testing to simulate real network attack scenarios. The construction method of the environment model is as follows: S1.1: Determine the overall architecture of the network and select the components of the network, including external networks, firewalls, DMZs, and multiple internal networks; S1.2: Draw a network topology diagram to visualize the network structure, and use the adjacency matrix M to record the connection relationship between nodes. Each element of the matrix represents a node in the network. S1.3: Set up the segmented structure of the network and determine the number and size of subnets; S1.4: Set up connectivity between subnets and use firewall rules to manage traffic in and out of each subnet; S1.5: Determine the number of hosts in each subnet; S1.6: Assign a unique IP address to each host; S1.7: Install the appropriate operating system and configure necessary services and applications for each host; S1.8: Query the vulnerability database to screen out vulnerabilities related to the host operating system and services and determine the impact scope and exploitation method of each vulnerability; S1.9: Simulate known vulnerabilities on the host; S1.10: Configure vulnerable services or applications and then introduce available attack tools; S1.11: Create an attribute record for each node containing the specific information of the network environment described above; integrate the network topology map and the attribute record of each node to generate the operating environment of the algorithm.

8. The penetration test path planning method based on ant colony and reinforcement learning according to claim 1 is characterized in that: Clean the data in the environment model of the penetration test, including: Check the adjacency matrix and remove redundant connections to improve data accuracy; identify isolated nodes and choose to retain or delete them as needed; Fill in or delete missing host attributes; Detect abnormal values, identify unreasonable host configuration data, and delete or replace them with reasonable values; Check the consistency of data field types and formats; Remove outdated or irrelevant vulnerability information and unify the way vulnerability risk levels are expressed.

9. The penetration test path planning method based on ant colony and reinforcement learning according to claim 8 is characterized in that: After data cleaning, feature extraction is performed: use network scanning tools to collect node information, extract data features, integrate the extracted features into a unified feature set, ensure that each sample has a complete feature description, and organize it into a structured data set.

10. The penetration test path planning method based on ant colony and reinforcement learning according to claim 1 is characterized in that: Screen the different effective and feasible paths for comprehensive performance evaluation, including cumulative rewards, path length, risk level and key target coverage, and select the best attack path.

Citation Information

Patent Citations

  • Penetration test path planning method based on A3C model

    CN116595536A

  • Intelligent penetration testing method for dynamic network environment

    CN117614728A

  • Intelligent penetration testing method and system based on deep learning

    CN118410497A