Automatic penetration testing method for power internet of things
Optimizing the penetration test of the Internet of Things of the power through deep reinforcement learning models and attention mechanisms, the problem of existing methods relying on expert experience and blind exploration is solved, and efficient and adaptive penetration test is achieved.
Patent Information
- Application Number
- CN202510446283.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-25
AI Technical Summary
The existing power IoT penetration testing methods rely highly on expert experience, are time-consuming and labor-intensive, and are difficult to meet the rapidly developing power IoT vulnerability penetration testing needs, and traditional automatic penetration tools have blind exploration behaviors.
The deep reinforcement learning model is used to combine attention mechanisms to build state space and action space, and optimize penetration testing through expert prior knowledge and power IoT knowledge, and automatically explore attack paths to avoid blind exploration.
It realizes adaptive penetration testing without mastering all knowledge of the power IoT system in advance, improves penetration testing efficiency, focuses on target host testing, and reduces dependence on expert skills.
Smart Images

Figure CN120372612A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power information security, and particularly to an automatic penetration testing method for the power Internet of Things. Background Art
[0002] In penetration testing, penetration testers use various attack means and tools to try to crack the security protection mechanism of the system in order to discover and exploit potential security vulnerabilities. Due to the characteristics of the vulnerability such as concealment, vulnerability to attack, and fragility, penetration testers need to master the vulnerability characteristics of the power Internet of Things information system and combine various testing tools for efficient penetration. Therefore, it is often performed by experts with rich development experience and comprehensive technical knowledge. The complexity of the power Internet of Things causes great difficulties in manually discovering hidden vulnerabilities, analyzing the attack surface of the power Internet of Things, and designing penetration testing schemes. For example, in a small program of a certain power system, it is necessary to manually perform multiple packet captures and data modifications to discover a SQL injection method, and this SQL injection attack will cause a large amount of data leakage in the database. This highly expert-experience-dependent and time-consuming and laborious testing method is increasingly difficult to meet the penetration testing requirements of the ever-emerging vulnerabilities in the rapid development process of the power Internet of Things. In order to avoid security risks such as data leakage and system hacking attacks brought by the power Internet of Things, the State Grid Corporation urgently needs an auxiliary tool for penetration testing of the power Internet of Things information system to reduce the skill level requirements for penetration testers.
[0003] The existing technologies are generally divided into two types.
[0004] The penetration testing method based on manual work requires professional technical personnel to select a test tool set according to experience for the characteristics of the target object and keep trying. Eventually, two results are obtained: 1) The target object is penetrated, and the target object needs to be secured according to the test results; 2) The target object is not penetrated; after trying all known penetration methods, the target object still cannot be penetrated, and the conclusion is that the target object does not need to be secured. The penetration testing method based on reinforcement learning replaces manual attempts with an artificial intelligence model, uses the reinforcement learning theory to continuously learn the penetration process, and improves the penetration method until the penetration is completed. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides an automatic penetration testing method for the power Internet of Things, including the following steps:
[0006] Step 1: Establish an initial state space of the learning process according to expert experience;
[0007] Step 2: Introduce an attention mechanism to process the state space, so that the state of each row of the state space is converted into a key-value pair, ensuring that the input parameters required for subsequent deep reinforcement learning networks to learn remain unchanged;
[0008] Step 3: On the basis of constructing the state space, construct a deep reinforcement learning model, including constructing the action space, reward function required for deep reinforcement learning, and the training process to automatically explore the attack penetration path for automatic penetration testing.
[0009] Further, the establishment of the state space in Step 1 specifically includes the following steps:
[0010] Step 101: Establish a matrix M = {M i,j} N×N to represent the connection relationship of the power Internet of Things device nodes, where M i,j represents whether there is a connection relationship between node i and node j, 1 represents a direct connection, -1 represents an indirect connection, and N represents the network scale;
[0011] According to the station control layer - bay layer - process layer topology of the power Internet of Things, for the new nodes discovered during the penetration testing process, after identifying the characteristics of the nodes, divide them into different levels; if the node is an operation station, it belongs to the station control layer; if the node is a protection and measurement control device, a fault filtering device, etc., it belongs to the bay layer; if the node is an intelligent terminal, a merging unit, etc., it belongs to the process layer. The connection relationship between nodes within the same level is 1, and the node relationship between different levels is -1. Generally, the connection between different levels is established by cross-layer gateway devices. Therefore, the relationship between the cross-layer gateway device and the nodes of the corresponding two levels is 1;
[0012] During the automatic penetration testing process, the agent selects the nodes with a connection relationship of 1 as potential next penetration testing targets;
[0013] Step 102: Based on the hardware, operating system, software, ports, and service asset attributes of the nodes, establish an asset attribute vector Ass(M i ) = {d1, d2,..., d m} using the bitmap method, where each element d i represents an attribute, and m represents the number of attributes; this number can be determined by an asset scanning tool;
[0014] According to expert prior knowledge, enumerate the common hardware types in the power Internet of Things system, number them according to the serial numbers 1 to l a , and convert them into binary codes. The coding length is which represents rounding up; similarly, enumerate the common operating system types in the power Internet of Things system, number them according to the serial numbers 1 to l b , and convert them into binary codes. The coding length is Enumerate the common software types in the power Internet of Things system, number them according to the serial numbers 1 to l cThe number is converted into a binary code, and the code length is Enumerate the common port types in the power Internet of Things system, according to the serial numbers 1 to l d The number is converted into a binary code, and the code length is Enumerate the common service types in the power Internet of Things system, according to the serial numbers 1 to l e The number is converted into a binary code, and the code length is Therefore, the total length of each asset attribute vector is
[0015] Step 103: According to the vulnerability information of the node, including attributes such as Vector (attack vector), Complexity (attack complexity), Interactions (interactability), Privileges (required permissions), etc., establish a vulnerability attribute vector Attr(M i ) = {c1, c2,..., c n} by the bitmap method, where each element c i represents an attribute, and n represents the number of attributes; this number can be determined according to the CVE vulnerability library and the CVSS vulnerability scoring system; the binary coding method and the total length of each vulnerability attribute vector refer to Step 103;
[0016] Step 104: For each node i, combine its connection relationship vector with other nodes, that is, M i,j , j = 1, 2,..., N, and the asset attribute vector Ass(M i ) = {d1, d2,..., d m}, and the vulnerability attribute vector Attr(M i ) = {c1, c2,..., c n} into a state vector Finally, obtain the state matrix That is, the state space; where k represents the number of nodes in the topology structure of the power Internet of Things that have been discovered currently:
[0017]
[0018] Furthermore, introduce an attention mechanism to process the state matrix, specifically as follows:
[0019] For the state matrix For each row of Use the Padding mask method to expand it to a fixed length size, and the state matrix The row dimension will increase as the number of discovered device nodes increases; for this purpose, an attention mechanism is introduced to convert the state of each row into a key-value pair Key&Value, where the key vector Key represents the index of the information at each position in the input sequence, that is, the identification feature of the device node; the value vector Value is the carrier of the content or context information associated with each position, containing the information that actually needs to be attended to and combined into the output, that is, the state vector of each device. Therefore, whenever a new host is added, it is equivalent to adding a set of key-value pairs; then the keys are summed to obtain the vector v that provides input to the subsequent neural network:
[0020]
[0021] where x represents the content of each element in the input sequence, that is, the state vector of each device. χ i represents the attention weight; the calculation method of the weight factor can be represented by the softmax function:
[0022]
[0023] where W k and W q are parameters learned automatically, W k is the weight mapping to the key vector Key, while W q is the weight mapping to the query vector query; therefore, through the above processing method, regardless of the number of host nodes discovered by the agent, after expanding the state vector of each device to a fixed length size through the above Padding mask method, the number of input layer neurons in the subsequent deep reinforcement learning network can remain unchanged.
[0024] Furthermore, to build a deep reinforcement learning model, specifically, two multi-layer neural network structures of the same scale are built, one is the target network and the other is the evaluation network; among them, the number of input layers of the neural network is consistent with the dimension of the v vector obtained by using the attention mechanism, and the number of intermediate layers and output layers can be flexibly configured according to actual needs; the current loss is calculated through the calculation formula of the loss function of typical deep reinforcement learning, and then the parameters of the neural network are iteratively optimized; the parameters of the evaluation network are updated in real time during training, while the parameters of the target network are updated according to the parameters of the evaluation network every interval of time τ.
[0025] Furthermore, the construction of the action space is specifically as follows:
[0026] The attack actions are divided into tactical categories such as reconnaissance, initial access, execution, persistence, privilege escalation, defense evasion, discovery, lateral movement, data collection, and command and control. Then, parameterized action encoding is performed on the attack tools, adopting a composite structure of "tactical category + dynamic parameter", including the following core fields: policy, tool, action type, action execution parameter, where the policy includes tactical categories such as reconnaissance, initial access, execution, persistence, privilege escalation, defense evasion, discovery, lateral movement, data collection, and command and control; the tools include Nmap, Metaspolit, Mimikatz, Netcat; the action types include asset scanning, token stealing, brute force cracking, account enumeration, service exploitation, protocol spoofing, etc.; the action execution parameter is the parameter required when the corresponding action is executed.
[0027] Furthermore, for the host node i, a reward function r for action execution is defined (i) :
[0028]
[0029] where Value(i) represents the value of the host node i, which is determined according to its role in the power Internet of Things. Specifically, the SCADA system can be set to 100, the remote terminal unit (RTU) to 50, and the smart meter to 30; or, if there is a specific target test host, it is set to 100.; Score(v i ) is the vulnerability score of the host node i, which is comprehensively calculated through the Common Vulnerability Scoring System CVSS and the power equipment proprietary vulnerabilities. Its calculation formula is:
[0030]
[0031] where Pa represents whether it is a power equipment proprietary vulnerability. If so, Pa = 1.2; otherwise, it is 1;
[0032]
[0033] Scope = Unchanged means that when the vulnerability is exploited, its impact is limited within the initial security trust boundary, that is: the attacker cannot break through to other areas with different security permissions or protection levels through this vulnerability; among them,
[0034] IS = 1 - [(1 - C)·(1 - I)·(1 - A)]
[0035] C, I, A: respectively represent the impact degrees of the vulnerability on confidentiality, integrity, and availability, which can be found from CVSS;
[0036] Ex = 8.22·AV·AC·PR·UI
[0037] AV, AC, PR, and UI represent the values indicating the attack vector, attack complexity, required privileges, and user interaction, respectively, which can be found in CVSS;
[0038] Cost(a) represents the cost of executing attack a, which can be comprehensively calculated based on indicators such as attack execution time and resource consumption. The calculation formula is:
[0039] Cost(a) = ω t ·T(a) + ω r ·R(a)
[0040] ω t and ω r represent the time cost weight and resource cost weight, respectively; T(a) represents the time cost, that is, the expected time required for the attack; R(a) represents the resource cost, that is, the resources consumed during the attack execution, which can be determined by comprehensively considering the CPU occupancy rate and memory occupancy rate.
[0041] Advantages of the present invention:
[0042] Regarding the problem of dynamic change and expansion of the state space, the technical solution of the present invention can be adaptively adjusted according to newly discovered vulnerabilities and newly discovered topologies, without the need to master all the knowledge of the power Internet of Things system in advance;
[0043] In the state space and reward function, the knowledge of the power Internet of Things is incorporated, enabling the system to focus more on the testing of target hosts, avoiding a large number of blind exploration behaviors existing in traditional automatic penetration tools, and improving the efficiency of penetration testing. Description of the Drawings
[0044] Figure 1 Flow diagram of the present invention;
[0045] Figure 2 Schematic diagram of constructing a state matrix based on prior knowledge;
[0046] Figure 3 Schematic diagram of a penetration testing training framework based on deep reinforcement learning. Specific Embodiments
[0047] The principles and features of the present invention will be described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0048] Referring to Figures 1-3 , the present invention provides a method for automatic penetration testing of the power Internet of Things, including the following steps
[0049] Step 1: Establish an initial state space for the learning process based on expert experience;
[0050] Step 2: Introduce an attention mechanism to process the state space, so that the state of each row in the state space is converted into a key-value pair, ensuring that the input parameters to be learned by the subsequent deep reinforcement learning network remain unchanged;
[0051] Step 3: On the basis of constructing the state space, construct a deep reinforcement learning model, including constructing the action space, reward function required for deep reinforcement learning, and the training process to automatically explore the attack penetration path for automatic penetration testing.
[0052] Further, the establishment of the state space in Step 1 specifically includes the following steps:
[0053] Step 101: Establish a matrix \(M = \{M_{ij}\}\) representing the connection relationship of power Internet of Things device nodes, where \(M_{ij}\) represents whether there is a connection relationship between node \(i\) and node \(j\), 1 represents a direct connection, -1 represents an indirect connection, and \(N\) represents the network scale; i,j} N×N where \(M_{ij}\) i,j represents whether there is a connection relationship between node \(i\) and node \(j\), 1 represents a direct connection, -1 represents an indirect connection, and \(N\) represents the network scale;
[0054] According to the station control layer - bay layer - process layer topology of the power Internet of Things, for the new nodes discovered during the penetration testing process, after identifying the characteristics of the node, classify it into different levels; if the node is an operation station, it belongs to the station control layer; if the node is a protection and measurement control device, a fault filtering device, etc., it belongs to the bay layer; if the node is an intelligent terminal, a merging unit, etc., it belongs to the process layer). The connection relationship between nodes within the same level is 1, and the node relationship between different levels is -1. Generally, the connection between different levels is established by cross-level gateway devices. Therefore, the relationship between the cross-level gateway device and the nodes of the corresponding two levels is 1;
[0055] During the automatic penetration testing process, the agent selects the node with a connection relationship of 1 as the potential next penetration testing target;
[0056] Step 102: Based on the hardware, operating system, software, port, and service asset attributes of the node, establish an asset attribute vector \(Ass(M_{ij})=\{d_1,d_2,\cdots,d_m\}\) through the bitmap method, where each element \(d_k\) represents an attribute, and \(m\) represents the number of attributes; this number can be determined by an asset scanning tool; i ) = \{d1,d2,...,d m}, where each element \(d i represents an attribute, and \(m\) represents the number of attributes; this number can be determined by an asset scanning tool;
[0057] According to expert prior knowledge, enumerate the common hardware types in the power Internet of Things system, number them according to the serial numbers 1 to \(l\) a and convert them into binary codes, and the coding length is representing the ceiling function; similarly, enumerate the common operating system types in the power Internet of Things system, number them according to the serial numbers 1 to \(l\) bThe number is converted into a binary code, and the code length is Enumerate the common software types in the power Internet of Things system, according to the serial numbers 1 to l c The number is converted into a binary code, and the code length is Enumerate the common port types in the power Internet of Things system, according to the serial numbers 1 to l d The number is converted into a binary code, and the code length is Enumerate the common service types in the power Internet of Things system, according to the serial numbers 1 to l e The number is converted into a binary code, and the code length is Therefore, the total length of each asset attribute vector is
[0058] Step 103: Based on the vulnerability information of the node, including attributes such as Vector (attack vector), Complexity (attack complexity), Interactions (interactability), Privileges (required permissions), etc., establish a vulnerability attribute vector through the bitmap method, where each element represents an attribute, indicating the number of attributes; this number can be determined according to the CVE vulnerability library and the CVSS vulnerability scoring system; the binary coding method, and the total length of each vulnerability attribute vector refer to Step 102;
[0059] Step 104: For each node i, the connection relationship vector with other nodes, that is, M i,j , j = 1, 2,..., N, and the asset attribute vector Ass(M i ) = {d1, d2,..., d m}, the vulnerability attribute vector Attr(M i ) = {c1, c2,..., c n} are merged into a state vector Finally, the state matrix is obtained That is, the state space; where k represents the number of nodes in the topological structure of the power Internet of Things that have been discovered currently:
[0060]
[0061] As an improvement of the solution, the state matrix will increase as the number of host nodes explored in the penetration testing process increases. Specifically, each row of the state matrix includes three parts: the vector M representing the connection relationship of the host node i , the vector Ass(M i ) representing the asset attributes of the host node, and the vector Attr(M i)。Since the lengths of these vectors change as the penetration testing process progresses, the input dimension of the DQN network will also change accordingly. However, DQN generally uses a fully connected layer neural network representation, and this network topology requires the number of neurons in each layer to be fixed. Therefore, the present invention introduces an attention mechanism to address this problem.
[0062] The attention mechanism is introduced to process the state matrix as follows:
[0063] For the state matrix each row is expanded to a fixed length size using the Padding mask method, and the row dimension of the state matrix will increase as the number of host nodes increases; therefore, an attention mechanism is introduced to convert the state of each row into a key-value pair Key&Value, where the key vector Key represents the index of the information at each position in the input sequence, that is, the identification feature of each position; the value vector Value is the carrier of the content or context information associated with each position, containing the information that is actually to be focused on and combined into the output, that is, the information considered valuable for subsequent processing at each position; therefore, whenever a new host is added, it is equivalent to adding a set of key-value pairs; then the keys are summed to obtain the vector v provided as input to the subsequent neural network:
[0064]
[0065] where x represents the content of each element in the input sequence, that is, the state vector of each device χ i represents the attention weight; the calculation method of the weight factor can be represented by the softmax function:
[0066]
[0067] where W k and W q are automatically learned parameters, W k is the weight mapping to the key vector Key, and W q is the weight mapping to the query vector query; therefore, through the above processing method, regardless of the number of host nodes discovered by the agent, after expanding the state vector of each device to a fixed length size using the above Padding mask method, the number of neurons in the input layer of the subsequent deep reinforcement learning network can remain unchanged.
[0068] As an improvement to the solution, a deep reinforcement learning model is built, specifically, two multi-layer neural network structures of the same scale are built, one is the target network and the other is the evaluation network; among them, the number of input layers of the neural network is the same as that of the vector dimension obtained by using the attention mechanism, and the number of middle layers and output layers can be flexibly configured according to actual needs; the current loss is calculated through the loss function calculation formula of typical deep reinforcement learning, and then the neural network parameters are iteratively optimized; the parameters of the evaluation network are updated in real time through training, while the parameters of the target network are updated according to the parameters of the evaluation network every interval of time τ. The training process continuously optimizes the strategy through interaction with the environment. At the beginning of training, the evaluation network and the target network are randomly initialized. In each round of training, the model will select an action (such as scanning or vulnerability exploitation) according to the current state and execute this action in the simulation environment to observe the results (such as the next state and reward). These interaction data (state, action, reward, next state) will be stored in the experience replay buffer for subsequent model updates. The model updates the evaluation network and the target network through the gradient descent method to maximize the cumulative reward. After training is completed, the model can be used for automated penetration testing to quickly identify vulnerabilities and optimize the attack path.
[0069] As an improvement to the solution, the action space defines all possible operations that an attacker can perform in each state during penetration testing, including scanning, vulnerability exploitation, privilege escalation, data stealing, etc. For example, the scanning action involves using the Nmap tool to detect the open ports and services of the target network; the vulnerability exploitation action involves using the Metasploit module to attack known vulnerabilities; the privilege escalation action involves using the Mimikatz tool to extract the password hash value in the Windows system and passing the hash value to the target system through the Pass-the-Hash technology to obtain administrator privileges; the data stealing action may involve using the FTP tool to transfer sensitive files (such as database files) in the target system to a remote server controlled by the attacker, and using the Netcat tool to package and send the data to the attacker's listening port through the TCP / UDP protocol. The action space includes discrete actions (that is, each action corresponds to a specific operation) and continuous actions (the actions are parameterized, such as "scanning port range 1-1024"). Among them, the construction of the action space is specifically as follows:
[0070] Referring to the MITRE ATT&CK framework, attack actions are divided into tactical categories such as reconnaissance, initial access, execution, persistence, privilege escalation, defense evasion, discovery, lateral movement, data collection, and command and control. Then, parameterized action encoding is performed on the attack tools, adopting a composite structure of "tactical category + dynamic parameter", including the following core fields: strategy, tool, action type, action execution parameter, where the strategy includes tactical categories such as reconnaissance, initial access, execution, persistence, privilege escalation, defense evasion, discovery, lateral movement, data collection, and command and control; the tools include Nmap, Metaspolit, Mimikatz, Netcat; the action types include asset scanning, token stealing, brute force cracking, account enumeration, service exploitation, protocol spoofing, etc.; the action execution parameter is the parameter required when the corresponding action is executed.
[0071] Among them, the reward function is used to evaluate the effect of each action and provide an optimization direction for the reinforcement learning model. Positive rewards are associated with successful operations (such as obtaining permissions, stealing data), while negative rewards are associated with being detected or operation failures. For example, successfully obtaining administrator permissions may receive a reward of +10, while triggering an intrusion detection system (IDS) alert may result in a penalty of -5. Neutral rewards are used for operations with no significant impact (such as scanning and finding no valuable information).
[0072] Among them, for the host node i, the reward function r for action execution is defined (i) :
[0073]
[0074] Among them, Value(i) represents the value of the host node i, which is determined according to its role in the power Internet of Things. Specifically, the SCADA system can be set to 100, the remote terminal unit (RTU) is set to 50, and the smart meter is set to 30; or, if there is a specific target test host, it is set to 100.; Score(v i ) is the vulnerability score of the host node i, which is comprehensively calculated through the Common Vulnerability Scoring System (CVSS) and power equipment-specific vulnerabilities. Its calculation formula is:
[0075]
[0076] Among them, Pa represents whether it is a power equipment-specific vulnerability. If so, Pa = 1.2, otherwise it is 1;
[0077]
[0078] Scope=Unchanged means that when the vulnerability is exploited, its impact is restricted within the initial security trust boundary, that is: the attacker cannot break through to other areas with different security permissions or protection levels through this vulnerability; among them,
[0079] IS=1-[(1-C)·(1-I)·(1-A)]
[0080] C, I, A: respectively represent the impact degrees of the vulnerability on confidentiality, integrity, and availability, which can be found from CVSS;
[0081] Ex=8.22·AV·AC·PR·UI
[0082] AV, AC, PR, UI respectively represent the values of attack vector, attack complexity, required privileges, and user interaction, which can be found from CVSS;
[0083] Cost(a) represents the cost of executing attack a, which can be comprehensively calculated based on indicators such as attack execution time and resource consumption. The calculation formula is:
[0084] Cost(a)=ω t ·T(a)+ω r ·R(a)
[0085] ω t 、ω r respectively represent the time cost weight and the resource cost weight; T(a) represents the time cost, that is, the expected time required for the attack; R(a) represents the resource cost, that is, the resources consumed for the attack execution, which can be determined by comprehensively considering the CPU occupancy rate and the memory occupancy rate.
[0086] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An automatic penetration testing method for the power Internet of Things, characterized in that, It includes the following steps Step 1: Establish the initial state space of the learning process based on expert experience; Step 2: Introduce an attention mechanism to process the state space, so that the state of each row of the state space is converted into a key-value pair, ensuring that the input parameters required for subsequent deep reinforcement learning networks remain unchanged; Step 3: On the basis of constructing the state space, construct a deep reinforcement learning model, including constructing the action space, reward function and training process required for deep reinforcement learning to automatically explore the attack penetration path for automatic penetration testing.
2. An automatic penetration testing method for the power Internet of Things, characterized in that: The establishment of the state space in Step 1 specifically includes the following steps: Step 101: Establish a matrix M = {M i,j} N×N , where M i,j represents whether there is a connection relationship between node i and node j. 1 represents a direct connection, -1 represents an indirect connection, and N represents the network scale; According to the station control layer - bay layer - process layer topology of the power Internet of Things, for the new nodes discovered during the penetration testing process, after identifying the characteristics of the node, divide it into different levels; The connection relationship between nodes within the same level is 1, the node relationship between different levels is -1, and the nodes between different levels are generally connected by cross-level gateway devices. Therefore, the relationship between the cross-level gateway device and the nodes of the corresponding two levels is 1; During the automatic penetration testing process, the agent selects the node with a connection relationship of 1 as the potential next penetration testing target; Step 102: According to the hardware, operating system, software, ports, and service asset attributes of the nodes, establish an asset attribute vector Ass(M i ) = {d1, d2,..., d m} by means of the bitmap method, where each element d i represents an attribute, and m represents the number of attributes; this number can be determined by an asset scanning tool; Enumerate the common types of hardware in the power Internet of Things system according to the prior knowledge of experts, and convert them into binary codes according to the numbers 1 to l a number, and the encoding length is which means rounding up; Similarly, enumerate the types of operating systems commonly used in the power Internet of Things system, numbered from 1 to l b According to the number, convert it into a binary code, and the code length is Enumerate the types of software commonly used in the power Internet of Things system, numbered from 1 to l c According to the number, convert it into a binary code, and the code length is Enumerate the types of ports commonly used in the power Internet of Things system, numbered from 1 to l d According to the number, convert it into a binary code, and the code length is Enumerate the types of services commonly used in the power Internet of Things system, numbered from 1 to l e According to the number, convert it into a binary code, and the code length is Therefore, the total length of each asset attribute vector is Step 103: According to the vulnerability information of the node, including attributes such as Vector (attack vector), Complexity (attack complexity), Interactions (interactiveness), Privileges (required permissions), etc., establish a vulnerability attribute vector Attr(M i ) = {c1, c2, …, c n}, where each element c i represents an attribute, n represents the number of attributes; this number can be determined according to the CVE vulnerability library and the CVSS vulnerability scoring system; the binary encoding method and the total length of each vulnerability attribute vector refer to Step 102; Step 104: For each node i, combine its connection relationship vector with other nodes, i.e., M i,j , j = 1, 2,..., N, and the asset attribute vector Ass(M i ) = {d1, d2,…, d m}, and the vulnerability attribute vector Attr(M i ) = {c1, c2,..., c n} into a state vector Finally, obtain the state matrix That is, the state space; where k represents the number of nodes in the topology of the power Internet of Things that have been discovered currently:
3. The automatic penetration testing method for the power Internet of Things according to claim 1, characterized in that: Introduce an attention mechanism to process the state matrix, specifically as follows: For the state matrix each row of is expanded to a fixed length size using a padding mask, while the state matrix row dimension will increase as the number of explored device nodes increases; Therefore, introduce an attention mechanism to convert the state of each row into a key-value pair Key&Value, where the key vector Key represents the index of the information at each position in the input sequence, that is, the identification feature of the device node; The value vector Value is a carrier of the content or context information associated with each position, containing the information that is actually to be concerned about and combined into the output, that is, the state vector of each device Therefore, whenever a new host is added, it is equivalent to adding a set of key-value pairs; then the key values are summed to obtain the vector v that provides input to the subsequent neural network: where x represents the content of each element in the input sequence, i.e., the state vector of each device χ i represents the attention weight; The calculation method of the weight factor can be represented by the softmax function: Among which W k and W q are parameters for automatic learning. W k is the weight mapping to the key vector Key, while W q is the weight mapping to the query vector query; therefore, through the above processing method, regardless of the number of host nodes discovered by the agent, after expanding the state vector of each device to a fixed length size through the above Padding mask method, the number of input layer neurons in the subsequent deep reinforcement learning network can remain unchanged.
4. The automatic penetration testing method for the power Internet of Things according to claim 1, wherein: Among them, building a deep reinforcement learning model specifically means building two multi-layer neural network structures of the same scale, one is the target network and the other is the evaluation network; among them, the number of input layers of the neural network is consistent with the dimension of the v vector obtained by using the attention mechanism, and the number of intermediate layers and output layers can be flexibly configured according to actual needs; calculate the current loss through the calculation formula of the loss function of typical deep reinforcement learning, and then iteratively optimize the neural network parameters; the parameters of the evaluation network are updated in real time through training, while the parameters of the target network are updated according to the parameters of the evaluation network every interval of time τ.
5. The automatic penetration testing method for the power Internet of Things according to claim 1, wherein: Among them, The construction of the action space is specifically as follows: Divide the attack actions into tactical categories such as reconnaissance, initial access, execution, persistence, privilege escalation, defense evasion, discovery, lateral movement, data collection, and command and control, and then parameterize the action encoding of the attack tools, using a composite structure of "tactical category + dynamic parameter", including the following core fields: policy, tool, action type, action execution parameter, where the policy includes tactical categories such as reconnaissance, initial access, execution, persistence, privilege escalation, defense evasion, discovery, lateral movement, data collection, and command and control; the tools include Nmap, Metaspolit, Mimikatz, Netcat; The action types include: asset scanning, token stealing, brute force cracking, account enumeration, service exploitation, protocol spoofing, etc.; The action execution parameter is the parameter required when the corresponding action is executed.
6. The automatic penetration testing method for the power Internet of Things according to claim 1, characterized in that: For host node i, define the reward function r for action execution (i) : Among them, Value(i) represents the value of host node i, which is determined according to its role in the power Internet of Things. Specifically, the SCADA system can be set to 100, the remote terminal unit (RTU) to 50, and the smart meter to 30; or, if there is a specific target test host, it is set to 100; Score(v i ) is the vulnerability score of host node i, which is comprehensively calculated through the Common Vulnerability Scoring System (CVSS) and power equipment-specific vulnerabilities. Its calculation formula is: Among them, Pa represents whether it is a proprietary vulnerability of power equipment. If so, Pa = 1.2, otherwise it is 1; Scope=Unchanged means that when a vulnerability is exploited, its impact is restricted within the initial security trust boundary, that is: an attacker cannot break through to other areas with different security permissions or protection levels through this vulnerability; among them, IS=1-[(1-C)·(1-I)·(1-A)] C, I, A: respectively represent the degrees of impact of the vulnerability on confidentiality, integrity, and availability, which can be found in CVSS; Ex=8.22·AV·AC·PR·UI AV, AC, PR, UI respectively represent the values of the attack vector, attack complexity, required privileges, and user interaction, which can be found in CVSS; Cost(a) represents the cost of executing attack a, which can be comprehensively calculated based on indicators such as attack execution time and resource consumption. The calculation formula is: Cost(a) = ω t ·T(a) + ω r ·R(a) ω t and ω r respectively represent the time cost weight and the resource cost weight; T(a) represents the time cost, that is, the time expected to be spent on the attack; R(a) represents the resource cost, that is, the resources consumed for the execution of the attack, which can be determined by integrating the CPU occupancy rate and the memory occupancy rate.
Citation Information
Cited By
Power system network intelligent penetration test method based on deep reinforcement learning
CN121690739A