Smart grid honeypot design method and system based on reinforcement learning
By applying reinforcement learning technology in smart grid honeypots, the problem of insufficient interaction in the face of specific industrial protocols and collaborative attacks is solved, deeper attacker interaction and more effective attack data capture are achieved, and targeted protection capabilities are enhanced.
Patent Information
- Application Number
- CN202310243016.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-14
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-03-14
AI Technical Summary
The existing smart grid honeypot design method has problems such as limited interaction, insufficient scalability and lack of physical process simulation when facing industrial protocols, collaborative attacks and deep interactions in specific industries, making it difficult to effectively capture new and collaborative attack information.
Using the smart grid honeypot design method based on reinforcement learning, we generate the correct response expected by the attacker through offline and online data acquisition, design and simulation of honeypot systems, offensive and defensive behavior modeling, and the application of reinforcement learning models, deeply induce attacker interaction, analyze attack behavior and extract features to target the protection of real power equipment.
It improves the depth of interaction between honeypot and attackers, can effectively capture attack data, enhance targeted protection capabilities, and adapt to a wider range of threats and attack modes.
Smart Images

Figure CN116405258B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power grid security, and in particular relates to a smart grid honeypot design method and system based on reinforcement learning. Background Art
[0002] The rapid development of the industrial Internet has increased the risk of cyber attacks on key infrastructure such as power grids year by year. Currently, many power attacks have occurred. With the continuous integration of information systems and physical systems, the originally isolated smart substation power equipment has begun to directly or indirectly access the Internet. Due to weak security measures and important military and economic value, smart grids have quickly become an important target of cyber attacks and are facing an increasing threat of cyber attacks. In recent years, attacks on smart grids have occurred frequently, causing devastating effects.
[0003] Honeypots are an important supplement to existing defense methods (such as intrusion prevention systems and mobile target defense), and are an effective deception defense method. They play a vital role in the security protection of smart grids. By studying smart grid honeypots, we can conduct in-depth analysis of the power information-physical fusion system, explore possible risk points, and provide guidance for the design of future power equipment. At the same time, studying smart grid security defense technology is of great significance to improving the defense capabilities of smart grids and protecting the safe and stable operation of key industrial infrastructure.
[0004] Defects of the prior art:
[0005] 1. Current methods usually have limited interactivity. Since they are designed with more consideration for simulating basic protocols, they are not implemented for industrial protocols in some specific industries. As a result, they are often at a disadvantage in targeted attacks and cannot effectively capture attack information. They cannot effectively capture new attacks and coordinated attacks.
[0006] 2. The existing honeypot design methods in the power industry are usually highly targeted and lack scalability. In the context of the rapid increase in attack and defense in the current network security field, it is necessary to design and implement universal and easily scalable honeypots to deal with more threats.
[0007] 3. In addition, the existing methods often have the problem of not or rarely considering the simulation of the physical process of the power grid industry during the construction of the honeypot. The attacker's request data after establishing a connection cannot be responded to according to the response of the real device, resulting in the inability to interact deeply with the attacker and reducing the probability of discovering vulnerabilities. Summary of the invention
[0008] In view of the problems existing in the prior art, the present invention proposes a smart grid honeypot design method and system based on reinforcement learning, which can generate the correct response expected by the attacker based on the attacker's request, so as to obtain more attack data to specifically protect real power equipment or systems.
[0009] In order to achieve the above object, the present invention adopts the following technical solutions:
[0010] The present invention provides a smart grid honeypot design method based on reinforcement learning, comprising the following steps:
[0011] Obtain a large amount of raw request response data for the power grid from the network and test environment using offline and online methods;
[0012] Design an initial honeypot and deploy it in a validation environment as a trap environment;
[0013] After modeling the attack and defense behaviors, the reinforcement learning method is used to intelligently learn the behavior characteristics of power equipment based on the collected original request response data, and automatically generate the best response to reply to the attacker to lure the attacker into the system;
[0014] Analyze the attack behaviors captured during the interaction process, extract attack features, and protect real power equipment or systems in a targeted manner.
[0015] Furthermore, an offline acquisition method is used to perform traffic analysis based on the existing data set and test environment to extract request and response data. An online acquisition method is used to deploy the honeypot instance on the Internet, and then forward the request data from the attacker's honeypot detection and scanning process to the test environment and real devices available online for response collection.
[0016] Furthermore, the obtained original request and response data are preprocessed to filter out irrelevant and erroneous request and response data.
[0017] Furthermore, the design of the initial honeypot includes simulating the general power protocol and the specific power industry protocol, followed by emulating the communication network, and automatically generating model files for the power equipment in the test environment.
[0018] Furthermore, the optimal response decision problem is modeled as a Markov decision process, and the reinforcement learning method is used to solve the approximate optimal response.
[0019] Furthermore, the optimal response decision problem is modeled as a Markov decision process, and then the reinforcement learning method is used to solve the approximate optimal response, including:
[0020] Step 1: Perform attack surface analysis based on the test environment to implement attack and defense behavior modeling;
[0021] Step 2: By designing a honeypot system, the attacker is induced to turn his target to the honeypot system.
[0022] Step 3: In the reinforcement learning model, the agent interacts with the environment, uses the agent's perception of the environment, inputs the action to be performed, and learns the mapping relationship from state to action through interaction with the dynamic environment;
[0023] Step 4, value iteration update;
[0024] Step 5, automatically generating a candidate response set based on response judgment based on reinforcement learning;
[0025] Step 6, select the best response from the candidate response set by executing the ε-greedy strategy.
[0026] Furthermore, the implementation process of step 3 is as follows:
[0027] State value function V π (s) represents the value of the state after executing strategy π, and the calculation formula is as follows:
[0028] V π (s) = E π [R t |s t =s] (1)
[0029] In the formula, π is the strategy selected in learning, s is any state in the state set S, t represents the state at a certain time, s t is the state at time t, R t is the reward obtained at time t, E π is to adopt strategy π, and the state at time t is s t Time R t expectations;
[0030] The state-action value function, or Q function, represents the value of taking an action when executing the policy π in a certain state. The calculation formula is as follows:
[0031] Q π (s,a)=E π [R t |s t =s,a t =a] (2)
[0032] Where a is any action in the action set A, a t is the action at time t;
[0033] The optimal strategy is the strategy that produces the optimal value function. The optimal value function is the Q function that takes the maximum value. The calculation formula of the optimal value function is as follows:
[0034] Vπ (s)=max a Q π (s,a) (3).
[0035] Furthermore, the value iteration update in step 4 includes:
[0036] Initialize the random value function, that is, the random value of each state; then, use formula (2) to calculate the Q function Q of all state-action pairs π (s,a), and then use the maximum value of Q π (s,a) Update the optimal value function V π (s), repeat the above process until the optimal value function changes very little, then the final optimal value function is obtained.
[0037] Furthermore, step 6 of selecting the best response from the candidate response set by executing the ε-greedy strategy includes:
[0038] Generate a random number ε, ε is small enough, the agent selects the best action according to the Q function of the next state with a probability of 1-ε, that is, argmax a∈A Q π (s,a).
[0039] The present invention also provides a smart grid honeypot system based on reinforcement learning, comprising:
[0040] An information collection module is used to obtain a large amount of original request response data for the power grid from the network and the test environment in an offline and online manner;
[0041] An intelligent decision-making module is used to model the attack and defense behaviors and then use the reinforcement learning method to intelligently learn the behavior characteristics of power equipment based on the collected original request response data, automatically generate the best response to reply to the attacker, so as to lure the attacker into the system;
[0042] The data analysis module is used to analyze the attack behaviors captured during the interaction process, extract attack features, and protect real power equipment or systems in a targeted manner.
[0043] Compared with the prior art, the present invention has the following advantages:
[0044] 1. The reinforcement learning-based smart grid honeypot design method of the present invention simulates the general power protocol and the specific power industry protocol in the honeypot to build a communication network, thereby solving the problem that the current smart grid honeypot design mainly focuses on the simulation of basic protocols and cannot effectively capture targeted attacks, advanced attacks and coordinated attacks.
[0045] 2. The present invention models the attack and defense behaviors, characterizes the image of the attacker, and then models the decision-making choices performed during the interaction process as a Markov decision process. It uses the reinforcement learning method to automatically learn the behavioral characteristics of the power equipment, so that in the actual execution process, it can generate the correct response expected by the attacker based on the attacker's request, greatly improve the interaction depth between the honeypot and the attacker, and induce the attacker to go deep into the system, so as to obtain more attack data to more specifically protect the real power equipment or system. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0047] Figure 1 It is a flowchart of a smart grid honeypot design method based on reinforcement learning according to an embodiment of the present invention;
[0048] Figure 2 It is a structural block diagram of a smart grid honeypot system based on reinforcement learning according to an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] The smart grid honeypot design method based on reinforcement learning in this embodiment includes four core stages: full-factor information collection stage, initial honeypot design stage, intelligent decision-making stage and data analysis stage. Figure 1 As shown, the method includes the following steps.
[0051] Step S1, using offline and online methods to obtain a large amount of original request response data for the power grid from the network and the test environment, and the data can be stored in the attack and defense knowledge base.
[0052] Considering that when the honeypot generates a response, in order to automatically generate the response expected by the attacker as accurately as possible, a large amount of original request response data is required for reinforcement learning.
[0053] An offline acquisition method is used to perform traffic analysis based on existing data sets and test environments to extract request and response data.
[0054] Using online acquisition, we tested on Shodan and FOFA platforms and found that when obtaining information about devices such as Remote Terminal Units (RTUs), almost all of them were marked as honeypots, which means it is difficult to obtain effective responses from real devices. Based on the defects of existing network detection, this example uses a real test environment, deploys a honeypot instance on the Internet, and forwards the request data from the attacker during the honeypot detection scan to the test environment and the real devices available online in the network for response collection.
[0055] In addition, the original request and response data collected above need to be preprocessed, and irrelevant and erroneous request and response data need to be filtered out using clustering, leaving only useful request and response data.
[0056] Step S2, design an initial honeypot and deploy it in the verification environment as a trapping environment.
[0057] 1. Simulation of honeypot protocol and network, simulate the general power protocol and specific power industry protocols (such as IEC61850 and other protocols), and then simulate the communication network to truly reproduce the network traffic.
[0058] 2. The simulation of the honeypot industrial physical process determines the authenticity of the honeypot system. The automatic generation of model files for the power equipment in the test environment can present the most realistic environment to the attacker to pass the attacker's initial reconnaissance.
[0059] Specifically, the designed initial honeypot is deployed in the SCADA host of the monitoring station in the verification environment, which is equivalent to the host having two SCADA systems at the same time, one is the original system and the other is the honeypot instance interface.
[0060] Step S3, after modeling the attack and defense behaviors, the reinforcement learning method is used to intelligently learn the behavior characteristics of the power equipment based on the original request response data collected in step S1, and the best response is automatically generated to reply to the attacker to lure the attacker into the system.
[0061] This example models the optimal response decision problem as a Markov decision process, and then uses the reinforcement learning method to solve the approximate optimal response. The specific steps include:
[0062] Step S301, performing attack surface analysis based on the test environment to implement attack and defense behavior modeling.
[0063] For example, when analyzing the attacker's behavior from the perspective of possible attacks on smart substations, attackers may launch two types of attacks based on a single small smart substation, one of which is to obtain data traffic through the substation host or to conduct penetration tests on the process layer network and then execute attacks.
[0064] Step S302, by designing a honeypot system, the attacker is induced to turn his target to the honeypot system, thereby protecting the real device / system.
[0065] The problem of choosing the best response is modeled as a Markov decision process (MDP for short). In an MDP, there is a set of states that the agent can visit. At each state, the agent can choose an action from a set of actions. The reward function defines the reward assigned after each action. The agent searches for a strategy defined by the relationship between states and actions, and the goal is to find the optimal strategy by maximizing the payoff. Usually, the transition function and reward function in an MDP are unknown. In addition, the computational complexity of the Bellman equation is very high. These two problems lead us to turn to reinforcement learning, which can approximate the optimal strategy. Reinforcement learning contains the following steps S303-S305.
[0066] Step S303: In the reinforcement learning model, the agent interacts with the environment, and inputs the action to be performed in a way that the agent perceives the environment, and learns the mapping relationship from state to action through interaction with the dynamic environment.
[0067] Specifically, the policy function can be expressed as π(s): S->A, which represents the mapping from state to action. The policy function indicates what action to perform in each state. The goal is to find the optimal policy that specifies the correct action in each state, so as to maximize the reward. The state value function determines the best degree of an agent in a specific state under the policy π. It is usually denoted as V π (s), represents the value of the state after executing strategy π, and the calculation formula is as follows:
[0068] V π (s) = E π [R t |s t =s]
[0069] In the formula, π is the strategy selected in learning, s is any state in the state set S, t represents the state at a certain time, s t is the state at time t, R t is the reward obtained at time t, E π is to adopt strategy π, and the state at time t is s t Time R t The formula determines the expected return from starting in state s under policy π.
[0070] The state-action value function, or Q function, determines the optimality of an action performed by an agent in a particular state under strategy π, usually denoted by Q π (s,a), represents the value of the action taken by the execution strategy π in a certain state, and the calculation formula is as follows:
[0071] Q π (s,a)=E π [R t |s t =s,a t =a]
[0072] Where a is any action in the action set A, a t is the action at time t. This formula determines the expected reward obtained by taking action a starting from state s under policy π.
[0073] The problem of selecting the best response is transformed into finding the action corresponding to the optimal strategy. The optimal strategy is the strategy that produces the optimal value function. The optimal value function is the Q function that takes the maximum value. The calculation formula of the optimal value function is as follows:
[0074] V π (s)=max a Q π (s,a).
[0075] Step S303: value iteration update.
[0076] Initialize the random value function, that is, the random value of each state; then calculate the Q function Q of all state-action pairs π (s,a), and then use the maximum value of Q π (s,a) Update the optimal value function V π (s), repeat the above process until the optimal value function changes very little, then the final optimal value function is obtained, and the action a with the largest Q value is selected as the optimal strategy for the state s.
[0077] Step S304: automatically generate a candidate response set based on the response judgment of reinforcement learning.
[0078] The agent selects a response a from the candidate response set, and then measures the effectiveness (i.e., reward) r(s, a) of the response by calculating whether it induces the attacker to further interact based on the current state.
[0079] Step S305, selecting the best response from the candidate response set by executing the ε-greedy strategy.
[0080] Generate a random number ε, ε is small enough, the agent selects the best action according to the Q function of the next state with a probability of 1-ε, that is, argmaxa∈A Q π (s,a), this action is the best response.
[0081] Step S4, analyze the attack behaviors captured during the interaction process, extract attack features, such as identifying the attacker's target, attack intentions and other information, and protect the real power equipment or system in a targeted manner, such as by transferring, isolating and other means to protect the real power equipment or system.
[0082] Corresponding to the above-mentioned smart grid honeypot design method based on reinforcement learning, Figure 2 As shown, this embodiment also proposes a smart grid honeypot system based on reinforcement learning, including:
[0083] The information collection module is used to obtain a large amount of original request response data for the power grid from the network and the test environment in an offline and online manner.
[0084] The intelligent decision-making module is used to model the attack and defense behaviors and then use the reinforcement learning method to intelligently learn the behavior characteristics of power equipment based on the collected original request response data, automatically generate the best response to reply to the attacker, so as to lure the attacker into the system.
[0085] The data analysis module is used to analyze the attack behaviors captured during the interaction process, extract attack features, and protect real power equipment or systems in a targeted manner.
[0086] The honeypot system of this embodiment also includes the simulation of the honeypot protocol and network and the simulation of the honeypot industrial physical process. Specifically, the general power protocol and the specific power industry protocol are simulated, and then the communication network is simulated; the model file is automatically generated for the power equipment in the test environment to present the most realistic environment to the attacker to facilitate the attacker's reconnaissance.
[0087] The present invention not only considers the simulation of general power protocols, but also can simulate specific power industry protocols, so that the honeypot system can capture more types of attacks. The present invention uses the reinforcement learning method to automatically learn the behavioral characteristics of power equipment. In the actual execution process, it can generate the attacker's expected response based on the attacker's request, greatly improve the interaction depth between the honeypot and the attacker, and induce the attacker to go deep into the system, so as to obtain more attack data to more specifically protect the real power equipment or system.
[0088] It should be noted that, in this article, the terms "comprises", "includes" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or apparatus.
[0089] Finally, it should be noted that the above is only a preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention, and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A smart grid honeypot design method based on reinforcement learning, It is characterized in that The following steps are involved: Step 1: Obtain a large amount of original request response data for the power grid from the network and test environment using offline and online methods; Step 2: Design the initial honeypot and deploy it in the verification environment as a trapping environment; Step 3: After modeling the attack and defense behaviors, the reinforcement learning method is used to intelligently learn the behavior characteristics of the power equipment based on the collected original request response data, and the best response is automatically generated to reply to the attacker to induce the attacker to penetrate the system; The optimal response decision problem is modeled as a Markov decision process, and then the reinforcement learning method is used to solve the approximate optimal response, including: Step 301, performing attack surface analysis based on the test environment to implement attack and defense behavior modeling; Step 302, by designing a honeypot system, inducing the attacker to turn his target to the honeypot system; Step 303, in the reinforcement learning model, the agent interacts with the environment, uses the agent to perceive the environment, inputs the action to be performed, and learns the mapping relationship from state to action through interaction with the dynamic environment; Step 304, value iteration update; Step 305, automatically generating a candidate response set based on the response judgment of reinforcement learning; Step 306, selecting the best response from the candidate response set by executing the ε greedy strategy; Step 4: Analyze the attack behaviors captured during the interaction process, extract attack features, and protect real power equipment or systems in a targeted manner.
2. According to the smart grid honeypot design method based on reinforcement learning according to claim 1, It is characterized in that An offline acquisition method is used to perform traffic analysis based on existing data sets and test environments to extract request and response data. An online acquisition method is used to deploy a honeypot instance on the Internet, and then forward the request data from the attacker's honeypot detection and scanning process to the test environment and real devices available online for response collection.
3. According to the reinforcement learning-based smart grid honeypot design method of claim 1, It is characterized in that Preprocess the obtained raw request and response data to filter out irrelevant and erroneous request and response data.
4. The smart grid honeypot design method based on reinforcement learning according to claim 1, It is characterized in that The design of the initial honeypot includes the simulation of the general power protocol and the specific power industry protocol, followed by the simulation of the communication network and the automatic generation of model files for the power equipment in the test environment.
5. The smart grid honeypot design method based on reinforcement learning according to claim 1, It is characterized in that The implementation process of step 3 is as follows: State value function V π (s) represents the value of the state after executing strategy π, and the calculation formula is as follows: V π (s)=Ε π [R t |s t =s] (1) In the formula, π is the strategy selected in learning, s is any state in the state set S, t represents the state at a certain time, s t is the state at time t, R t is the reward obtained at time t, Ε π is to adopt strategy π, and the state at time t is s t Time R t expectations; The state-action value function, or Q function, represents the value of taking an action when executing the policy π in a certain state. The calculation formula is as follows: Q π (s,a)=Ε π [R t |s t =s,a t =a] (2) Where a is any action in the action set A, a t is the action at time t; The optimal strategy is the strategy that produces the optimal value function. The optimal value function is the Q function that takes the maximum value. The calculation formula of the optimal value function is as follows: V π (s)=max a Q π (s,a) (3)。 6. According to the reinforcement learning-based smart grid honeypot design method of claim 5, It is characterized in that The value iteration update in step 4 includes: Initialize the random value function, that is, the random value of each state; then, use formula (2) to calculate the Q function Q of all state-action pairs π (s,a), and then use the maximum value of Q π (s,a) Update the optimal value function V π (s), repeat the above process until the optimal value function changes very little, then the final optimal value function is obtained.
7. The smart grid honeypot design method based on reinforcement learning according to claim 6, It is characterized in that Step 6 selects the best response from the candidate response set by executing the ε greedy strategy, including: Generate a random number ε, ε is small enough, the agent selects the best action according to the Q function of the next state with a probability of 1-ε, that is, argmax a∈A Q π (s,a).
8. A smart grid honeypot system based on reinforcement learning, It is characterized in that The system is used to implement the smart grid honeypot design method based on reinforcement learning as described in any one of claims 1 to 7, comprising: An information collection module is used to obtain a large amount of original request response data for the power grid from the network and the test environment in an offline and online manner; An intelligent decision-making module is used to model the attack and defense behaviors and then use the reinforcement learning method to intelligently learn the behavior characteristics of power equipment based on the collected original request response data, automatically generate the best response to reply to the attacker, so as to lure the attacker into the system; The data analysis module is used to analyze the attack behaviors captured during the interaction process, extract attack features, and protect real power equipment or systems in a targeted manner.