Industrial Internet of Things intrusion detection method based on human feedback reinforcement learning
Through the combination of reinforcement learning based on human feedback and large language models, the shortcomings of traditional methods in adaptability and real-time are solved, and efficient and accurate industrial IoT intrusion detection is achieved, adapting to complex attacks and reducing latency.
Patent Information
- Application Number
- CN202510521150.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional industrial IoT intrusion detection methods are difficult to adapt to complex network attack modes, with high false positive rate and weak generalization ability. Deep learning methods lack interpretability and rely on large-scale annotation of data. The design of reward function for reinforcement learning limits the model's adaptability. At the same time, traditional cloud IDS solutions face the problems of limited computing resources and data transmission delay.
Using a method based on human feedback reinforcement learning, we use the reward model to build a reward model and combine it with a large language model (LLM), we automatically extract key features from the industrial environment, dynamically adapt to new attack modes, use human feedback to optimize the reward mechanism, and combine the edge computing architecture to improve real-time and detection accuracy.
It significantly improves the accuracy, response speed and generalization ability of industrial IoT intrusion detection, and can quickly identify multiple attack modes, especially in complex attack scenarios, with high adaptability to variant attacks, and reduces cloud computing latency.
Smart Images

Figure CN120342704A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial Internet of Things intrusion detection, and particularly relates to an industrial Internet of Things intrusion detection method based on human feedback reinforcement learning. Background Art
[0002] With the transformation of global manufacturing towards intelligence and digitization, the Industrial Internet of Things (IIoT) is becoming one of the core technologies driving Industry 4.0 and intelligent manufacturing. The IIoT enhances the level of industrial automation, optimizes production processes, and improves resource utilization efficiency through technologies such as the Internet of Things, cloud computing, big data analysis, and artificial intelligence (AI). Combining the rapid development of artificial intelligence technology in the industrial Internet of Things, large models based on human feedback reinforcement learning (RLHF) have demonstrated powerful capabilities in fields such as natural language processing (NLP) and automated decision-making, and play a role in industrial intelligent scenarios, such as intelligent operation and maintenance, industrial control optimization, etc.
[0003] In the architecture of the industrial Internet of Things, the perception layer deploys devices such as sensors, RFID, and PLC for data collection, the network layer ensures data transmission through protocols such as 5G, LoRa, and MQTT, the platform layer processes data based on cloud computing and edge computing, and finally realizes functions such as intelligent manufacturing, predictive maintenance, and remote monitoring in the application layer. The introduction of artificial intelligence makes the industrial Internet of Things move further towards intelligence. In particular, large models trained by RLHF have demonstrated powerful capabilities in intelligent optimization, fault prediction, industrial robot scheduling, etc.
[0004] However, the widespread deployment of the industrial Internet of Things has also brought new security challenges. Due to the large number of IIoT devices, and most of them being low-power and heterogeneous devices, traditional security protection means are difficult to effectively cope with increasingly complex network attacks. Industrial control systems face threats such as DDoS attacks, data tampering, and malicious code injection, making intelligent intrusion detection (IDS) the key to ensuring the security of the industrial Internet. Traditional signature matching detection methods can no longer cope with the constantly evolving attack patterns. Therefore, AI-based intelligent IDS has become a new development direction, analyzing the network traffic and behavior patterns of IIoT devices through deep learning and reinforcement learning technologies to detect potential threats in real time and improve security protection capabilities.
[0005] The rapid development of the current Industrial Internet of Things (IIoT) has made intelligent manufacturing, remote monitoring, and predictive maintenance possible. However, at the same time, the increasingly complex cyberattack means have posed many challenges to traditional intrusion detection methods. Existing intrusion detection systems (IDSs) based on signature matching or anomaly detection are difficult to adapt to the constantly changing attack patterns, resulting in a high false alarm rate and weak generalization ability. In addition, although deep learning methods can improve the detection accuracy, they often lack interpretability and rely on large-scale labeled data, making it difficult to efficiently adapt to new threats in the industrial environment. Reinforcement learning (RL) has been applied in the field of IDSs, but the fixed reward function design limits the model's adaptability and makes it difficult to efficiently capture the dynamic changes of attack behaviors. At the same time, the limited computing resources of IIoT devices make traditional cloud-based IDS solutions face data transmission delays and privacy risks, making it difficult to meet the high real-time and security requirements. Summary of the Invention
[0006] The purpose of the embodiments of the present invention is to provide an intrusion detection method for the Industrial Internet of Things based on human feedback reinforcement learning, aiming to solve the problems raised in the above background technology.
[0007] The embodiments of the present invention are implemented as follows. The intrusion detection method for the Industrial Internet of Things based on human feedback reinforcement learning includes the following steps:
[0008] Obtain the intrusion detection abnormal traffic data features through the IIoT scenario;
[0009] Construct a reward model based on human feedback reinforcement learning, and dynamically calibrate the reward model according to human preferences by introducing corresponding labels;
[0010] Fine-tune the large language model based on the optimized reward model, and generate intrusion detection prediction results for determination.
[0011] Another purpose of the embodiments of the present invention is to provide an intrusion detection system for the Industrial Internet of Things based on human feedback reinforcement learning, which is used to implement the above-mentioned intrusion detection method for the Industrial Internet of Things based on human feedback reinforcement learning, and includes:
[0012] A data acquisition module, which is used to obtain the intrusion detection abnormal traffic data features through the IIoT scenario;
[0013] A reward model training module, which is used to construct a reward model based on human feedback reinforcement learning, and dynamically calibrate the reward model according to human preferences by introducing corresponding labels;
[0014] An LLM fine-tuning module, which is used to fine-tune the large language model based on the optimized reward model, and generate intrusion detection prediction results for determination.
[0015] The industrial Internet of Things (IIoT) intrusion detection method based on human feedback reinforcement learning provided by the embodiments of the present invention combines an enhanced reward model and a large language model (LLM) through human feedback reinforcement learning (RLHF), significantly improving the accuracy, response speed, and generalization ability of IIoT intrusion detection. First, through stable and reliable structure learning, this method automatically extracts key features from industrial environment samples, avoiding false positives and false negatives caused by data bias in traditional machine learning (ML) and deep learning (DL) methods. At the same time, it uses human feedback to optimize the reward mechanism, enabling the detection system to dynamically adapt to new attack patterns and improve the generalization ability. Second, the combination of reinforcement learning and LLM makes the detection process more efficient. Compared with traditional methods, this method can identify intrusion threats and take response measures faster. At the same time, it combines an edge computing architecture to reduce cloud computing latency and improve real-time performance. In addition, this method successfully identified 14 different types of attacks in experiments, demonstrating detection capabilities superior to traditional IDS, especially in complex attack scenarios such as DDoS, SQL injection, malicious code, and data tampering. At the same time, it uses LLM for semantic analysis and pattern recognition to improve the adaptability to variant attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Schematic diagram of the industrial Internet of Things intrusion detection method based on human feedback reinforcement learning provided by the embodiments of the present invention;
[0017] Figure 2 Categories of the Edge-IIoT dataset provided by the embodiments of the present invention;
[0018] Figure 3 Schematic diagram of fine-tuning the LLM of the human feedback reinforcement learning reward mechanism provided by the embodiments of the present invention;
[0019] Figure 4 Receiver operating characteristic curve provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0021] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.
[0022] As Figure 1 shown, a schematic diagram of the industrial Internet of Things intrusion detection method based on human feedback reinforcement learning provided by an embodiment of the present invention includes the following steps:
[0023] S1. Obtain intrusion detection abnormal traffic data features through IIoT scenarios:
[0024] In the embodiments of the present invention, Edge-IIoTset is utilized. This dataset aims to provide a high-quality and realistic test environment for intrusion detection research in IIoT networks, including comprehensive annotations of network traffic data and attack types. There are 15 industrial IoT devices in total, including industrial sensors, IP cameras, smart meters, industrial robots, and PLCs, etc., serving as attack targets, and are divided into five threats: DoS / DDoS attacks, information collection, man-in-the-middle (MITM), injection attacks, and malware attacks, as Figure 2 shown;
[0025] Given a PCAP file containing network traffic logs, relevant features are extracted from a specific time window. There are approximately 2,219,201 network traffic samples, each sample having 61 features and their corresponding labels. This dataset contains 14 attack categories and one benign traffic category. The extracted features are organized into a CSV file format for analysis. Edge-IIoTset includes functions collected from various sources, including network traffic, logs, system resources, and alerts;
[0026] Ensure the dataset is clean and ready for further use by preprocessing the dataset, including the following steps:
[0027] Data filtering: Clean the data before analysis, including eliminating duplicate rows, empty columns, and incorrect feature columns to ensure the quality and reliability of the dataset;
[0028] Numericalization: Convert non-numerical categories to numerical values using "One-Hot-Encoding";
[0029] Normalization: Use the "MinMaxScaler" function to convert numerical values to a common interval between 0 and 1.
[0030] S2. Build a reward model based on human feedback reinforcement learning, and dynamically calibrate the reward model according to human preferences by introducing corresponding labels:
[0031] Input: Define X network ={x1, x2,..., x n} as the input network feature set, and here corresponding labels L h ={l1, l2,..., l n} are referenced according to human preferences. The input dataset of the reward model is expressed as:
[0032] In the formula, represents the labeled reasonable chosen response; Represented as a marked rejected response, the formula is simplified to
[0033] Output: The evaluation metric is defined as M evaluation ={m1, m2,..., m n} and the generated prediction label and is defined as the reward model's relative quality assessment of the label;
[0034] The reward model embedding function is expressed as:
[0035] In the formula, represents the embedding space; W is the dimension of the embedding space;
[0036] Given the input network feature set X network , each network feature x i is mapped to an n-dimensional vector e i , and the encoded representation in the embedding space e is: e i =f(x i , l i ); In the industrial Internet of Things, embedding can be used to convert network traffic data into a unique vector representation and compare these representations with a database of known attack features to identify threats faster and more accurately;
[0037] The human references the corresponding label and outputs it based on the preferred preference, determining that the reasonable choice is greater than the unreasonable choice The Bradley-Terry model provides a framework for preference modeling based on real-valued rewards. Given a reward function The preference distribution probability is formulated as follows:
[0038]
[0039] In the formula, σ is the sigmoid function;
[0040] Regarding this problem as a binary classification task, the negative log-likelihood loss function of the estimated reward function γ is:
[0041] Then it is used to fine-tune the model while maintaining proximity to the initial model and maximizing the reward objective required for optimizing the model: Implement dynamic model updates;
[0042] Where, η represents the penalty coefficient of Kullback - Leibler; KL represents the Kullback - Leibler divergence, which is used to ensure that as an entropy reward, it maintains generation diversity and prevents the mode from collapsing into a single high - return answer; secondly, it ensures that the output of the reinforcement learning strategy does not deviate significantly from the accurate distribution of the reward model.
[0043] S3. Fine - tune the large - language model based on the optimized reward model, generate intrusion detection prediction results for determination:
[0044] Fine - tuning the LLM involves training the basic LLM model on a dataset in a specific domain to improve the performance of specific tasks. Utilize the human - feedback reinforcement learning reward mechanism to guide the fine - tuning process. By taking the accuracy of industrial Internet of Things intrusion detection threat determination as the reward metric, define the feedback that a simulated human evaluator would provide, use the human - feedback reinforcement learning to generate a reward signal for the metric, and then update the model parameters through iterative interaction. The method of fine - tuning the LLM model with the human - feedback reinforcement learning reward mechanism is as Figure 3 shown. Enhance the performance of the model by providing structured and consistent feedback, and promote a more scalable fine - tuning process. Through this mechanism, the model can gradually improve its output to make it closer to the desired result. The human - feedback reinforcement learning process involves multiple iterations, and the model continuously adjusts its parameters according to the received rewards, thereby iteratively improving its performance.
[0045] By evaluating the similarity between network traffic embeddings, identify the deviations that may indicate abnormal network activities. Define the initialized vector database V, and for the embedding categories of network feature sample pairs Provide network scenarios, including malicious and benign instances, which are determined by the proportion of these two types of instances in adjacent labels;
[0046] Specify Update all vectors in their respective categories and assign weights according to their importance. Evaluate the incoming network traffic by generating their corresponding embeddings and checking whether any similar embeddings are stored in the vector database. This comparison can determine whether the incoming traffic is closely similar to any stored embeddings. Use the Euclidean distance to measure the distance between adjacent vectors A = {a1, a2,..., a n} and B = {b1, b2,..., b n}, and the distance metric calculation is expressed as:
[0047] By randomly selecting network feature samples and initializing different data sets according to the "chosen" and "rejected" sample embeddings selected using the markers, in order to improve the accuracy of the classifier, the sample labels are aligned pairwise to ensure uniformity and fairness in compression and positive and negative example calculations. For example, assuming an application distribution threshold of 10%, it is necessary to ensure that the embedding of the incoming network traffic must be 90% similar to the stored benign network embedding in order to be classified as benign. If the stored network embedding is benign and below this threshold, it indicates that the closest match in the database is not similar enough, indicating that the incoming embedding does not belong to the assumed benign category. The representation of the threshold distance is:
[0048] In the formula, is the clustering variance of joint learning; ρ is the standard deviation of the basic distribution for sampling the clustering mean; α is the hyperparameter that controls the clustering concentration in the Chinese restaurant process.
[0049] It can achieve a balance between fitting simple data distributions with low capacity and complex data distributions with high capacity, and divide the incoming network traffic into malicious traffic for further analysis. The balance of its embedding space is expressed as:
[0050] In the formula, align() represents updating the embedding vector to a new vector with the same maximum length as the longest embedding vector in all , and the elements exceeding the original length are filled with 0. Similarly:
[0051] On the training and test data sets of Edge-IIoTset, the intrusion detection performance of the industrial control network was evaluated. Respectively, the true positive rate (TP): the number of attack samples correctly classified; the false positive rate (FP): the number of benign samples correctly classified; the true negative rate (TN): the number of attack samples misclassified; and the false negative rate (FN): the number of benign samples misclassified, which are used to calculate different metrics:
[0052] Precision:
[0053] Recall:
[0054] False Alarm Rate:
[0055] The measure of test accuracy (F1 score) is defined as the balance between the harmonic mean of precision and recall:
[0056] Accuracy / Success Rate:
[0057] Edge-IIoTset was regularly partitioned, and the cleaned sample subsets were randomly assigned for training and evaluation. The numerical data set composed of network features was transformed into predefined categories with categorical data as groups. The benign category was defined as the "non-malicious" label, and all other 14 attack categories were labeled as the "malicious" label. At the same time, to ensure that each category was equally represented, the proportion sizes of all samples were made similar. The distribution of network attack samples on the training and evaluation data sets is shown in Table 1:
[0058] Table 1
[0059]
[0060] To convert numerical network features into categorical values, a quantile-based method was used to divide the data set into three different groups, and three classification groups of "primary", "intermediate", and "advanced" were set. Then, the paired summaries with manual annotations were calibrated, and each row of data was converted into a descriptive sentence to facilitate the model's understanding of network features and help the model be consistent with human judgment or make effective evaluations during the training process, as shown in Table 2:
[0061] Table 2
[0062]
[0063]
[0064] Experimental Verification:
[0065] Edge-IIoTset was used to test and compare traditional machine learning algorithms (decision tree, random forest, extreme gradient boosting, categorical feature boosting, and adaptive boosting); deep learning (convolutional neural network, deep neural network, time series prediction), natural language models, and the performance of the embodiments of the present invention. The parameter settings of the comparative algorithms in the embodiments of the present invention were used to evaluate the effectiveness of the fine-tuning big data reward model based on human feedback reinforcement learning through the divided training and test ratio subsets. The results are shown in Table 3:
[0066] Table 3
[0067]
[0068] It can be seen that the embodiments of the present invention have good performance in intrusion detection tasks. Compared with machine learning, deep learning, and natural language models, the embodiments of the present invention are higher than other algorithms in all evaluation indicators, and its accuracy rate reaches 98.7%. Similarly, based on the large language model, it can be seen that without any fine-tuning processing, the performance effect of the evaluation index of identifying intrusion detection is poor. This is because the large model cannot correctly understand the restricted devices and self-identification. At the same time, the embodiments of the present invention also add a reward model for comparison and embedding, which relatively improves the accuracy.
[0069] In addition, the area under the receiver operating characteristic curve (ROCAUC) scores for various categories are shown in Figure 4. A value of 1.0 indicates perfection. The AUC score for the "Normal" category is 1.0, the AUC score for the "Backdoor" category is 0.9998, the AUC score for the "DDoS_HTTP" category is 0.9992, the AUC score for the "DDoS_ICMP" category is 0.9994, the AUC score for the "DDoS_TCP" category is 0.9992, the AUC score for the "DDoS_UDP" category is 0.9989, the AUC score for the "MITM" category is 0.9986, the AUC score for the "Password" category is 0.9986, the AUC score for the "Port_Scanning" category is 0.9971, the AUC score for the "Ransomware" category is 0.9997, the AUC score for the "SQL_injection" category is 0.9997, the AUC score for the "Uploading" category is 0.9997, the AUC score for the "Vulnerability_scanner" category is 0.9996, and the AUC score for the "XSS" category is 0.9993. The AUC score for the "Fingerprinting" category is 0.9933, all of which are very close to 1.0. It can be seen that the classification of these categories is almost perfect and the misclassification is very small, meaning that almost perfect or perfect classification has been achieved for all categories using the embodiments of the present invention.
[0070] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An industrial Internet of Things intrusion detection method based on human feedback reinforcement learning, characterized in that, It includes the following steps: Obtain the intrusion detection abnormal traffic data features through the IIoT scenario; Construct a reward model based on human feedback reinforcement learning, and dynamically calibrate the reward model according to human preferences by introducing corresponding labels; Fine-tune the large language model based on the optimized reward model, and generate intrusion detection prediction results for determination.
2. The industrial Internet of Things intrusion detection method based on human feedback reinforcement learning according to claim 1, wherein The step of obtaining the intrusion detection abnormal traffic data features through the IIoT scenario specifically includes: Obtain the network traffic logs from the PCAP files of Edge-IIoTset, and extract the intrusion detection abnormal traffic data features; Preprocess the data features, including data filtering, data values, and normalization; Generate the input network feature set X network = {x1, x2,..., x n}.
3. The industrial Internet of Things intrusion detection method based on human feedback reinforcement learning according to claim 2, characterized in that, The step of constructing a reward model based on human feedback reinforcement learning and dynamically calibrating the reward model according to human preferences by introducing corresponding labels specifically includes: Generate sample labels based on the annotation information of Edge-IIoTset L h = {l1, l2,..., l n}, The input data set of the reward model is represented as: In the formula, represents the reasonably selected response to be marked; represents the unreasonably selected response to be marked; The reward model embedding function is expressed as: In the formula, is represented as the embedding space; W is the dimension of the embedding space; Given the input network feature set X network , map each network feature x i to an n-dimensional vector e i . The encoded representation in the embedding space e is: e i = f(x i , l i ); Use the Bradley-Terry model to calculate the preference probability distribution, expressed as: In the formula, the reward function σ is the sigmoid function; Fine-tuning model while maintaining proximity to the initial model to maximize the reward objective required for optimizing the model: achieve dynamic model updates; In the formula, η represents the penalty coefficient of Kullback-Leibler; KL represents the Kullback-Leibler divergence.
4. The industrial Internet of Things intrusion detection method based on human feedback reinforcement learning according to claim 3, wherein The Bradley-Terry model optimizes the reward function γ through a binary classification task. The negative log-likelihood loss function of the reward function γ is as follows:
5. The industrial Internet of Things intrusion detection method based on human feedback reinforcement learning according to claim 3, wherein, In the step of fine-tuning the large language model based on the optimized reward model to generate intrusion detection prediction results for determination, the Euclidean distance is used to measure the embedding similarity to determine whether the network traffic is a malicious attack. The similarity determination uses a threshold, which is expressed as follows: where is the clustering variance of the collaborative learning; ρ is the standard deviation of the basic distribution for sampling the clustering mean; α is the hyperparameter that controls the cluster concentration in the Chinese restaurant process.
6. An industrial Internet of Things intrusion detection system based on human feedback reinforcement learning, which is used to implement the industrial Internet of Things intrusion detection method based on human feedback reinforcement learning as described in any one of claims 1-5, and is characterized in that, It includes: A data acquisition module for obtaining the intrusion detection abnormal traffic data features through the IIoT scenario; A reward model training module for constructing a reward model based on human feedback reinforcement learning and dynamically calibrating the reward model according to human preferences by introducing corresponding labels; An LLM fine-tuning module for fine-tuning the large language model based on the optimized reward model and generating intrusion detection prediction results for determination.