Industrial Internet of Things Honeypot Deployment Method and Device Based on Adversarial Reinforcement Learning

Through the adversarial reinforcement learning method, the problem that smart honeypots are difficult to capture and analyze network attacks in dynamic environments is solved, and effective deployment and attack data analysis of industrial IoT honeypots in complex scenarios is achieved.

CN119341821BActive Publication Date: 2025-07-04ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411490198.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-07-04
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

Existing smart honeypots based on general reinforcement learning cannot effectively deal with the dynamics and uncertainties of cyber attack behavior, making it difficult to capture and analyze attack data in complex scenarios.

Method used

Adversarial reinforcement learning is adopted to collect data from industrial IoT devices, use LDA analysis to cluster data, build a Markov decision model, and use RUQL algorithm to train the model, dynamically adjust the honeypot response strategy to deal with complex and changeable attackers.

Benefits of technology

It improves the ability to capture and analyze cyber attacks under non-stationary conditions, enhances the effectiveness and adaptability of honeypots in dynamic environments, and can better attract and deal with complex attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119341821B_ABST
    Figure CN119341821B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for deploying honeypots in industrial Internet of Things based on adversarial reinforcement learning, belonging to the field of network security technology. It mainly solves the dynamicity and uncertainty of network attack behaviors. The implementation solution is to collect industrial Internet of Things device data, including industrial Internet of Things device IPs, industrial Internet of Things device request data and response data; perform clustering analysis on the request data by using the LDA analysis method based on industrial Internet of Things multi-modal request data to classify the request data into different categories; construct an interaction model between the honeypot and the attacker by using a Markov decision model; and finally train the model by using the adversarial reinforcement learning algorithm Repeated Update Q-learning (RUQL). The present invention can effectively attract and respond to network attacks, and significantly improve the ability to capture and analyze network attacks in an adversarial environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and particularly relates to a method and device for deploying an industrial Internet of Things (IIoT) honeypot based on adversarial reinforcement learning. Background Art

[0002] In recent years, with the popularization of industrial Internet of Things technology, its security problems have become increasingly prominent and have become the focus of attention in the academic and industrial circles. Industrial Internet of Things devices and systems are widely used in fields closely related to personal safety such as intelligent healthcare, security, transportation, or manufacturing. Their security problems not only concern the safety of human life and property, but also are major issues related to social stability and national security.

[0003] Compared with passive protection means such as intrusion detection and firewalls, a honeypot, as an active network security protection means to understand the attack situation, aims to attract and monitor the activities of attackers by deploying a false system disguised as a real system. A successful honeypot can simulate various services of a real system, including operating systems, applications, databases, etc., to provide an interaction experience close to that of real devices. Once an attacker starts to act on the honeypot, the honeypot will record their behavior data. This data includes the attacker's IP address, attack time, attack method, tools used, etc. By analyzing the attack data collected by the honeypot, the security team can understand the types, sources, and trends of current threats. Therefore, in order to better research and analyze new and unknown industrial Internet of Things attacks, the industrial Internet of Things honeypot technology has become an important research field in Internet of Things security.

[0004] In recent years, there have also been innovative studies integrating the concept of reinforcement learning into honeypot technology. An intelligent high-interaction honeypot based on general reinforcement learning uses algorithms to dynamically adjust and optimize the security mechanism of honeypot behavior. It analyzes the attacker's behavior patterns to learn and predict the attacker's next actions. However, in previous intelligent honeypots based on general reinforcement learning, the studied attackers were all in a steady state, that is, it was considered that the state transition of all attackers did not change with time, and the feedback to the same honeypot actions was homogeneous. But in real physical confrontation scenarios, the feedback of attackers to the same honeypot actions may be different. This means that even in the face of the same honeypot actions, the same attacker may give different feedback at different times. Therefore, the intelligent honeypot based on general reinforcement learning does not consider the dynamics and uncertainties of network attack behaviors and cannot cope with complex attack scenarios. Summary of the Invention

[0005] In order to overcome the deficiencies of the above-mentioned prior art, the present invention provides a method and device for deploying an industrial Internet of Things honeypot based on adversarial reinforcement learning, which solves the problem of difficultly effectively capturing and analyzing network attacks under non-stationary conditions.

[0006] The object of the present invention is achieved by the following technical solutions:

[0007] According to the first aspect of this specification, a method for deploying an industrial Internet of Things (IIoT) honeypot based on adversarial reinforcement learning is provided. The method includes the following steps:

[0008] (1) Collect IIoT device data, including IIoT device IPs, IIoT device request data, and response data;

[0009] (2) Based on the IIoT device request data collected in step (1), use the LDA analysis method based on IIoT multimodal request data to perform data clustering and divide the request data into different categories;

[0010] (3) Based on the classified request data in step (2), use a Markov decision model to construct an interaction model between the honeypot and the attacker. The honeypot is the agent, the network attacker who attacks the honeypot is the environment, the state is the classified request of the network attacker, the action is the response data replied by the honeypot, and the rewards induce requests for attacks and extended sessions;

[0011] (4) Based on the Markov decision model constructed in step (3), use the adversarial reinforcement learning algorithm RUQL to train the model, and the honeypot selects the optimal response for attackers with complex and variable behaviors for reply.

[0012] Further, the collection of IIoT device data is specifically: collect real IIoT device request data and IPs; traverse the IPs to open TCP sockets, actively send TCP requests to the IIoT devices, observe and collect their response data, and store them in the response database.

[0013] Further, in step (1), the method for collecting request data is: deploy a low-interaction honeypot instance to listen on the decoy port to achieve request data collection, or use the penetration tool Burpsuit to achieve request data collection.

[0014] Further, step (2) is specifically: regard each IIoT device request as a document, regard the predefined request data categories as the topics of the documents; divide each document into a series of words through the predefined delimiter; combine the words of all documents to form a global corpus; calculate the occurrence frequency of each word in the corpus; use the number of predefined request data categories as the number of topics, and after training the LDA model, the input request data categories can be obtained.

[0015] Further, the categories of the request data include: attack data against industrial Internet of Things devices, attack data containing Linux Bash instructions, attack data of the malware upload type, other attack data, and termination status data.

[0016] Further, in step (2), a data preprocessing step is also included, specifically: removing invalid or incorrect data records, and performing standardization processing on the request data so that the mean of each feature is 0 and the standard deviation is 1.

[0017] Further, in step (3), the goal of state definition is to capture the key features of the interaction between the attacker and the honeypot while keeping the state space relatively compact; the states of the interaction between the attacker and the honeypot are divided into five state sets, namely S = {s1, s2, s3, s4, s5}; s1 is the attack data against industrial Internet of Things devices, including vulnerability exploitation, privilege bypass, brute force cracking, malicious scanning, access attempt; s2 is the attack data containing Linux Bash instructions; s3 is the attack data of the malware upload type, and this type of attack is identified by interpreting the file suffix; s4 is other attack data, referring to requests that cannot be processed, including TLS communication composed of hexadecimal data; s5 is the termination status data, and when the agent does not receive the next state or times out, it will automatically enter the termination state.

[0018] Further, in step (4), the RUQL algorithm updates by considering the difference between the actual return and the predicted Q value and adjusts the update according to the difference between the policy and the behavior. The Q value update formula is:

[0019]

[0020] where the expression measures the difference between the current Q value prediction and the actual Q value, r is the reward obtained after executing action a in state s; γ is the discount factor; is the maximum Q value of all possible actions in the next state s'; α is the learning rate; π(s, a) is the probability that the policy selects action a in state s, is used as a weight to adjust the update to reflect the difference between the policy and the behavior; β is an adjustment parameter used to prevent the policy from approaching zero.

[0021] Further, in step (4), the running steps of the RUQL algorithm are specifically: a. Observe the current state s and arbitrarily initialize the value function Q; b. For each time step t, the agent selects an action a according to the current value function Q and the exploration strategy, outputs the action a and observes a resulting reward r and the next state s' of the environment; update the value function Q using the Q value update formula in RUQL; repeat step b until the value function Q converges.

[0022] According to the first aspect of this specification, there is provided an industrial Internet of Things honeypot deployment device based on adversarial reinforcement learning, including three components that operate independently and share data with each other during the learning process; specifically including:

[0023] A response database that stores the obtained response data related to industrial Internet of Things devices;

[0024] A data receiving module for capturing attack requests and simplifying the status of the requests using the LDA algorithm;

[0025] An adversarial interaction module that trains the model using the RUQL algorithm based on the given response feedback of the attacker; when the honeypot is dealing with attackers with complex and changeable behaviors, it continuously learns and adapts to the changes in network attacks, and dynamically adjusts the behaviors and strategies of the honeypot according to the observed attack behaviors to more effectively attract and respond to attacks.

[0026] In the research on attack traffic classification, it is found that when collecting a large amount of request data of real industrial Internet of Things devices, the original request data packets may lead to a sparse state space. Therefore, after researching a large number of literatures, a machine learning method of generative text clustering is used to perform text clustering on the original request data packets. Through in-depth analysis of the attack traffic classification research, the present invention realizes an industrial Internet of Things honeypot system based on adversarial reinforcement learning, which solves the problem of being difficult to effectively capture and analyze network attacks under non-stationary conditions. At the same time, aiming at improving the quantity of intelligence collection and the depth of attack interaction in the industrial Internet of Things honeypot interaction process, an adversarial reinforcement learning method is selected to build the honeypot system, and the effectiveness of the method is verified. Brief Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0028] Figure 1 It is a flowchart of the industrial Internet of Things honeypot deployment method based on adversarial reinforcement learning provided by the embodiment of the present invention;

[0029] Figure 2 It is a schematic diagram of the clustering analysis principle of LDA request data provided by the embodiment of the present invention;

[0030] Figure 3 It is the construction of the MDP model in a real adversarial scenario example provided by the embodiment of the present invention. Detailed Embodiment

[0031] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0032] Referring to Figure 1 , an embodiment of the present invention provides an industrial Internet of Things honeypot deployment method based on adversarial reinforcement learning. The method includes the following steps:

[0033] (1) Data collection of industrial Internet of Things devices; the ultimate goal of data collection is to collect response resources. Therefore, the first step in building a honeypot is to first collect real industrial Internet of Things device request data and IPs, then traverse the IPs to open TCP sockets, actively send TCP requests to the industrial Internet of Things devices, observe and collect their response data, and store it in the response database.

[0034] The present invention can collect request data in the following ways:

[0035] 1. Deploy a low-interaction honeypot instance on the Tencent Cloud server to collect relevant request data;

[0036] 2. Use the penetration tool Burpsuit to collect relevant request data. This part of the request data has higher quality compared to the request data collected by the low-interaction honeypot.

[0037] In the present invention, for IP collection, the third-party platform Shodan is used to obtain the original IP information. Its working principle is to crawl publicly accessible device information on the Internet, such as the IP addresses and metadata of cameras, printers, and routers. This information can include the device model, manufacturer, open ports, service type, etc.

[0038] In the present invention, for response data collection, determine the IP address of the device to be accessed, and use the Requests library of python to write a program to send requests and receive responses.

[0039] (2) Based on the industrial Internet of Things device request data collected in step (1), use the LDA (Latent Dirichlet Allocation) analysis method based on industrial Internet of Things multi-modal request data for data clustering. The industrial Internet of Things device request data contains multi-dimensional features, including but not limited to attack methods, timestamps, device status, operation types, etc. These features can more comprehensively reflect the operating status and request content of the devices.

[0040] In the data preprocessing stage, first remove invalid or incorrect data records. For example, delete records missing key information and correct obviously incorrect data.

[0041] Secondly, in order to eliminate the influence of dimensionality between the data requests of different devices, the request data is standardized so that the mean of each feature is 0 and the standard deviation is 1.

[0042] Next, each industrial Internet of Things device request is regarded as a document, and the predefined request data categories are regarded as the topics of the document; the categories of request data include: attack data against industrial Internet of Things devices, attack data containing LinuxBash instructions, attack data of malware upload type, other attack data, termination status data, and each of the above categories can further contain several subcategories; the present invention divides each document into a series of words through predefined delimiters; then combines the words of all documents (requests) together to form a global corpus; next, calculates the occurrence frequency (statistical distribution) of each word in the corpus; takes the number of predefined request data categories as the number of topics n, and after training the LDA model, the input request data categories can be obtained.

[0043] The above method applies the LDA model to the clustering analysis of industrial Internet of Things device request data. Through multi-dimensional feature extraction and customized clustering algorithms, it realizes in-depth understanding and dynamic monitoring of device request data, and improves the accuracy and real-time performance of data analysis. Combining the LDA model in the field of natural language processing with the industrial Internet of Things field realizes cross-field technology integration, and this integration provides a new perspective and method for industrial Internet of Things data analysis.

[0044] (3) Based on the classified request data in step (2), the present invention uses a Markov decision model to construct an interaction model between the honeypot and the attacker, where the honeypot is the agent; the network attacker who attacks the honeypot is the environment; the state is the request of the classified network attacker, and more specifically, the state represents the current situation of the attacker interacting with the honeypot.

[0045] The goal of the state definition in the present invention is to capture the key features of the interaction between the attacker and the honeypot while keeping the state space relatively compact. Using the network kill chain model, the present invention divides the interaction process between the attacker and the honeypot into seven stages, namely reconnaissance, weaponization, delivery, exploitation, installation, command and control, and objective achievement. For example, the request classification based on HTTP (such as GET, POST, PUT, etc.) or higher-level protocol types (such as FTP, SSH, etc.) can illustrate the malicious scanning behavior of the attacker in the reconnaissance stage. In addition, the classification based on the request target path (using the target URL path of the request as part of the state) can provide information about the weaponization stage. The goal of the embodiments of the present invention is to capture the key features of the interaction between the attacker and the honeypot while keeping the state space relatively compact. Therefore, the embodiments of the present invention divide the state of the interaction between the attacker and the honeypot into the following 5 state sets, a total of 24 states, that is, S = {s1, s2, s3, s4, s5}.

[0046] s1: Attack data for industrial Internet of Things devices. Such as vulnerability exploitation, privilege bypass, brute force cracking, malicious scanning, access attempts, etc. The specific content is shown in Table 1. s1 contains a total of 15 states.

[0047] Table 1 Attack data for industrial Internet of Things devices

[0048]

[0049]

[0050] s2: Attack data containing Linux Bash commands. According to the threat level, the embodiments of the present invention divide it into 4 states. The specific content is shown in Table 2. From bottom to top, the threat level increases in turn.

[0051] Table 2 Attack data containing Linux Bash commands

[0052] Linux Bash Commands kill, wegt, sudo, mount, umount, passwd cat, rm, tar, chmod, chown cd, cp, mkdir, rmdir, touch ls, pwd, ps, df

[0053] s3: Attack data of the malware upload type. This type of attack is common and harmful in WEB attacks and usually has obvious characteristics in the file suffix. The embodiments of the present invention identify this type of attack by interpreting the file suffix. The specific content is shown in Table 3 and is divided into 3 states.

[0054] Table 3 Attack data of the malware upload type

[0055] Suffix Type Variant asp asp, aspx, asa, asax, ascx, ashx, asmx, aSax, aScx, aShx, aSmx, cEr php php, ppHp3, pHp2, html, htm, phtml, pht, Html, Htm, pHtml jsp jsp, jspa, jspx, jsw, jsv, jspf, jtml, jSp, jSpx, jSpa, jSw, jSv, jSpf, jHtml

[0056] s4: Other attack data. It refers to some requests that cannot be processed, such as some TLS communications composed of hexadecimal data. Therefore, it is taken as a separate request category.

[0057] s5: Termination state data. When the agent does not receive the next state or times out, it will automatically enter the termination state.

[0058] (4) Based on the Markov decision model constructed in step (3), use the adversarial reinforcement learning algorithm RUQL to train the model. The main problem solved by the RUQL algorithm is the policy bias in the action value update. The policy bias problem appears in the Q-learning (QL) algorithm because the value Q of the action is only updated when the action is executed. Therefore, the update rate of the action value directly depends on the probability of selecting that action for execution. Obviously, in a single state, it is unrealistic to update the value Q of all actions at the same time, but the value Q of a certain action can be repeatedly updated under each state, so that the policy bias problem can be alleviated. Among them, the Q-value update formula in RUQL is:

[0059]

[0060] Among them, the expression in parentheses measures the difference between the current Q-value prediction and the actual Q-value, r is the reward obtained after executing action a in state s; γ is the discount factor, which determines the current value of future rewards; is the maximum Q-value of all possible actions in the next state s'. α is the learning rate, which controls the step size of Q-value update. π(s,a) is the probability that the policy selects action a in state s, This weight is used to adjust the update to reflect the difference between the policy and the behavior; if the policy is more likely to select action a, that is, π(s,a) is larger, the weight is smaller; if the policy is less likely to select action a, that is, π(s,a) is smaller, the weight is larger. β is an adjustment parameter to prevent the policy from approaching zero; when approaching zero, the update rate becomes unbounded (approaching infinity), so safeguard conditions must be added in practice. The update formula of the RUQL algorithm updates the Q-value by considering the difference between the actual return and the predicted Q-value and adjusting the update according to the difference between the policy and the behavior.

[0061] The running steps of the RUQL algorithm are as follows:

[0062] 1. Observe the current state s and arbitrarily initialize the value function Q.

[0063] 2. For each time step t, the agent selects action a according to the current value function Q and the exploration policy (such as ε-greedy).

[0064] 3. Output action a and observe a resulting reward r and the next state s of the environment ′ .

[0065] 4. Use the Q-value update formula in RUQL to update the value function Q. It should be noted that this weight is used to adjust the update, and the update is inversely proportional to the probability of selecting this action, so that the actions selected with low probability are updated multiple times.

[0066] 5. Repeat steps 2-4 until the value function Q converges.

[0067] The following will illustrate the advancement of the method of the present invention through the following examples, and test and compare the performance of the adversarial interactive honeypot based on RUQL (the honeypot designed by the present invention, hereinafter referred to as the RUQL honeypot), the honeypot based on the general reinforcement learning QL (hereinafter referred to as the QL honeypot), and the random honeypot after deploying the honeypot system in an actual scenario.

[0068] A. Experimental settings

[0069] The present invention has realized an adversarial interactive honeypot based on RUQL and deployed it on the public Internet. Due to limited budget, an instance with 2 cores and 2G of configuration and a public network broadband of 4M is used to run on the Tencent Cloud server. In order to simulate Hikvision cameras, the present invention has opened ports 80, 81, and 554. In order to verify the performance of the adversarial interactive honeypot based on RUQL, the present invention has also deployed a honeypot based on general reinforcement learning QL (QL honeypot) with the same configuration and a random honeypot to run on the Tencent Cloud server.

[0070] QL honeypot: Deploy the general reinforcement learning algorithm QL.

[0071] Random honeypot: A honeypot that does not deploy an algorithm and each response is selected with equal probability.

[0072] B. Evaluation results

[0073] After the three honeypots have run on the Tencent Cloud server for nearly a month (from April 8, 2024 to May 1, 2024), some basic data are statistically analyzed as shown in Table 4. The average session lengths of the RUQL honeypot, QL honeypot, and random honeypot are 4.9, 3.9, and 2.11 respectively. The numbers of attack sessions are 1908, 1159, and 425 respectively. Long sessions and multi-attack sessions indicate that attackers attempt to deeply probe or exploit vulnerabilities, while otherwise it may be simple scanning or harmless traffic. The random honeypot simulates a bait interface, with an average session length of about 2 times and fewer attack sessions. The two high-interaction honeypots, the RUQL honeypot and the QL honeypot, have long sessions and many attack sessions, provide rich services, and can deeply detect attacks. Since the QL honeypot attempts to learn a fixed optimal strategy, its evaluation metrics may be restricted by the strategy it has learned. Since the RUQL honeypot can adapt to environmental changes, its session length may be longer and the number of attack sessions may be more. When the behavior of the attacker changes, this non-stationary reinforcement learning honeypot of the RUQL honeypot can adjust its strategy to better simulate the real system and attract attackers to conduct deeper interactions.

[0074] Table 4 Basic data of different honeypots

[0075] RUQL Honeypot QL Honeypot Random Honeypot Average Session Length 4.9 3.9 2.11 Number of Attack Sessions 1908 1159 425

[0076] Table 5 shows the quantity and distribution of attack requests collected by different honeypots. The RUQL honeypot far exceeds the QL honeypot in terms of the quantity of attack requests collected. This means that the method of non-stationary reinforcement learning is more effective in honeypot design. The QL honeypot system under general reinforcement learning often attracts attackers based on static or fixed-rule behavior patterns, while the RUQL honeypot system under non-stationary reinforcement learning takes into account the dynamics and uncertainties of cyber-attack behaviors, and thus can attract and respond to these attacks more effectively.

[0077] Table 5 Quantity and distribution of attack requests collected by different honeypots

[0078] RUQL Honeypot QL Honeypot Random Honeypot Exploitation 118 19 0 Privilege Bypass 10 10 0 Brute Force 227 130 13 Malicious Scanning 276 139 21 Access Attempt 369 163 54

[0079] Figure 1 It is the flowchart of the method of the present invention. There are three main components running separately but sharing data with each other during the learning process. The response database mainly stores the obtained response data related to industrial Internet of Things devices. The data receiving module is used to capture attack requests and simplify the status of the requests using the LDA algorithm. On the other hand, this request can be used for later analysis of the attack trajectory, guessing the attacker's intention, and evaluating the analysis algorithm model. The adversarial interaction module uses the RUQL algorithm to train the model according to the given response feedback of the attacker. After several rounds of learning iterations, the honeypot of the present invention can optimize the model to reply to the attacker. From the beginning, the behavior of the honeypot system of the present invention is like that of a low-interaction honeypot because the system starts from zero knowledge of industrial Internet of Things devices and their behaviors. Through experimental evaluation, the industrial Internet of Things honeypot based on adversarial reinforcement learning can continuously learn and adapt to the changes of cyber-attacks when dealing with attackers with complex and changeable behaviors, dynamically adjust the behavior and strategy of the honeypot according to the observed attack behaviors, so as to attract and respond to attacks more effectively.

[0080] Figure 2It is the schematic diagram of the LDA analysis method based on multi-modal request data in the industrial Internet of Things. That is, the LDA model can be used to perform clustering analysis on device request data. This is an innovative attempt because LDA is usually used for the analysis of text data, and applying it to industrial Internet of Things device request data can uncover the potential themes and patterns behind device requests. Specifically, LDA first uses predefined delimiters (such as commas, spaces, etc.) to split each document (request) into a series of words (usually features or tokens). Next, the words of all documents (requests) are combined together to form a global corpus, which is also called a dictionary and contains the vocabulary in all documents. Then, the frequency of occurrence (statistical distribution) of each word in the corpus is calculated. Select a number of topics n (this is usually a hyperparameter that needs to be determined through experiments or domain knowledge). Use the LDA algorithm to assign words to n topics. Finally, after the LDA model is trained, new requests (documents) can be input into the model to see which topic they are most relevant to. The implementation of LDA is mainly based on two core libraries: the jieba word segmentation library and the LDA model library in Gensim. The specific steps are as follows. First, before performing topic modeling, the text needs to be preprocessed, including word segmentation, removing stop words and punctuation marks, etc. Word segmentation can use tools such as jieba, and removing stop words can use the nltk library. After word segmentation and building the dictionary, a document-term frequency matrix is generated through the Gensim library. This matrix counts the frequency of occurrence of each word in each document based on the vocabulary in the dictionary. Then, use the LdaModel class in the Gensim library to run the LDA model for further topic analysis. The number of topics, the dictionary, and the document-term frequency matrix need to be specified as input parameters. The model will automatically learn the distribution of topics and vocabulary. Once the model is trained, the show_topics method can be used to extract the topics. Each topic is represented by a group of high-weight words.

[0081] Figure 3It is the construction of the MDP (Markov Decision Process) model in the real adversarial scenario example provided by the embodiments of the present invention. Since the whole graph is too large, the embodiments of the present invention only emphasize the specific vulnerabilities and attack behaviors utilized. In the figure, the rectangular boxes represent states, the elliptical boxes represent actions, and the triangular boxes represent the actions that can be taken in specific states. In our context, each state is a unique request abstraction, and each operation of the request is a unique response candidate to be replied. The actions responseA, responseB, and responseC that can be selected under the root path ' / ' are shown. ResponseA is an invalid response, which can cause the attacker to end the session prematurely. ResponseB and responseC represent different types of valid responses respectively. As can be seen from the figure, when the RUQL honeypot replies with responseA, the session ends directly; when the RUQL honeypot replies with responseB, the attacker can send ' / image / lgbg.jpg' or ' / favicon.ico' for pre-attack checks. The node ' / image / lgbg.jpg' represents a certain image information of the device, and by analyzing

[0082] the image of / lgbg.jpg to identify that this is a Hikvision camera. The node with ' / favicon.ico' represents that the attacker attempts to access the favicon file, which represents the small icon of the manufacturer; when the RUQL honeypot replies with responseC, in the case where the pre-attack check has been completed, it can guide the attacker to the transition node ' / Security / users' of the vulnerability, and then enter the specific vulnerability node 'CVE-2021-36260'. From Figure 3 it can be seen that the RUQL-based agent in the adversarial environment not only regards responseC as a valid response, but also regards responseB as a kind of valid response, and responseB can effectively avoid the pre-attack check. Thus, it guides the attacker to access the transition node ' / Security / users' of the vulnerability.

[0083] The above description is only a specific example of the present invention, and does not constitute any limitation to the present invention. The step numbers in the specification and claims are only for the convenience of clearly describing the technical solution of the present invention, and their sequence numbers are not limited. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these corrections and changes based on the idea of the present invention are still within the protection scope of the claims of the present invention.

Claims

1. An industrial Internet of Things honeypot deployment method based on adversarial reinforcement learning, characterized in that Including: (1) Collecting industrial Internet of Things (IIoT) device data, including the IP of IIoT devices, request data, and response data of IIoT devices; (2) Based on the IIoT device request data collected in step (1), using the LDA analysis method based on multi-modal request data of the IIoT to perform data clustering, classifying the request data into different categories, treating each IIoT device request as a document, and treating the predefined request data categories as the topics of the documents; Dividing each document into a series of words through predefined delimiters; combining the words of all documents together to form a global corpus; Calculating the occurrence frequency of each word in the corpus; using the number of predefined request data categories as the number of topics, and after training the LDA model, the input request data categories can be obtained; (3) Based on the classified request data in step (2), using the Markov decision model to construct an interaction model between the honeypot and the attacker, where the honeypot is the agent, the network attacker who attacks the honeypot is the environment, the state is the request of the classified network attacker, the action is the response data replied by the honeypot, and the rewards are requests that induce attacks and extend sessions; (4) Based on the Markov decision model constructed in step (3), using the adversarial reinforcement learning algorithm RUQL to train the model, the honeypot selects the optimal response to reply to attackers with complex and changeable behaviors. The RUQL algorithm adjusts the update by considering the difference between the actual reward and the predicted Q value, and according to the difference between the policy and the behavior. The Q value update formula is: Among them, the expression measures the difference between the current Q-value prediction and the actual Q-value, r is the reward obtained after executing action a in state s; γ is the discount factor; is the maximum Q-value of all possible actions in the next state s'; α is the learning rate; π(s,a) is the probability that the policy selects action a in state s, serves as a weight to adjust the update to reflect the difference between the policy and the behavior; β is an adjustment parameter used to prevent the policy from approaching zero.

2. The method for deploying an industrial Internet of Things honeypot based on adversarial reinforcement learning according to claim 1, wherein The specific method of collecting IIoT device data is: collecting real IIoT device request data and IP; traversing the IP to open TCP sockets, actively sending TCP requests to IIoT devices, observing and collecting their response data and storing them in the response database.

3. The method for deploying an industrial Internet of Things honeypot based on adversarial reinforcement learning according to claim 1, wherein In step (1), the method of collecting request data is: deploying a low-interaction honeypot instance to listen on the decoy port to collect request data, or using the penetration tool Burpsuit to collect request data.

4. The method for deploying an industrial Internet of Things honeypot based on adversarial reinforcement learning according to claim 1, wherein The categories of the request data include: attack data against IIoT devices, attack data containing Linux Bash instructions, attack data of malware upload type, other attack data, and termination status data.

5. The method for deploying an industrial Internet of Things honeypot based on adversarial reinforcement learning according to claim 1, wherein Step (2) also includes a data preprocessing step, specifically: removing invalid or incorrect data records, and normalizing the request data so that the mean of each feature is 0 and the standard deviation is 1.

6. The method for deploying an industrial Internet of Things honeypot based on adversarial reinforcement learning according to claim 1, wherein In step (3), the goal of state definition is to capture the key features of the interaction between the attacker and the honeypot while keeping the state space relatively compact. The states of the interaction between the attacker and the honeypot are divided into five state sets, i.e., S = {s1, s2, s3, s4, s5}. s1 is the attack data against industrial Internet of Things devices, including vulnerability exploitation, privilege bypass, brute force cracking, malicious scanning, and access attempts. s2 is the attack data containing Linux Bash instructions. s3 is the attack data of the malware upload type, which is identified by interpreting the file suffix. s4 is other attack data, referring to requests that cannot be processed, including TLS communications composed of hexadecimal data. s5 is the termination state data. When the agent does not receive the next state or times out, it will automatically enter the termination state.

7. The method for deploying an industrial Internet of Things honeypot based on adversarial reinforcement learning according to claim 1, wherein In step (4), the specific running steps of the RUQL algorithm are as follows: a. Observe the current state s and arbitrarily initialize the value function Q. b. For each time step t, the agent selects an action a according to the current value function Q and the exploration strategy, outputs the action a, and observes a resulting reward r and the next state s' of the environment. Update the value function Q using the Q-value update formula in RUQL; repeat step b until the value function Q converges.

8. An industrial Internet of Things honeypot deployment device implemented by the method according to any one of claims 1-7, characterized in that, It includes three components. The three components run independently and share data with each other during the learning process. Specifically, it includes: The response database stores the obtained response data related to industrial Internet of Things devices. The data receiving module is used to capture attack requests and simplify the state of the requests using the LDA algorithm. The adversarial interaction module uses the RUQL algorithm to train the model based on the given response feedback of the attacker. When dealing with attackers with complex and changeable behaviors, the honeypot continuously learns and adapts to the changes in network attacks, and dynamically adjusts the behavior and strategy of the honeypot according to the observed attack behaviors to more effectively attract and respond to attacks.

Citation Information

Patent Citations

  • Internet of Things honeynet system based on reinforcement learning and dynamic scheduling method

    CN116132190A

  • Corpus set classification method based on TF-IDF and LDA topic models

    CN116595178A

  • BERT model-based reinforcement learning honeypot construction method and device

    CN117834228A