Transform-based Internet of Things intrusion detection method

By applying the Transformer model in IoT intrusion detection and combining data preprocessing and optimization strategies, the difficulties in the existing technology are solved when processing high-dimensional nonlinear data, and efficient, accurate and interpretable intrusion detection are achieved, adapting to the characteristics of the IoT environment.

CN120017345APending Publication Date: 2025-05-16AIR FORCE UNIV PLA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510142452.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-16

Smart Images

  • Figure CN120017345A_ABST
    Figure CN120017345A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of environmental science, in particular to an Internet of Things intrusion detection method based on Transform, which comprises the following steps: an acquisition step: acquiring network traffic data of Internet of Things equipment; a processing step of executing a data preprocessing operation based on the network flow data; according to the preprocessed data, feature extraction and intrusion detection are carried out through a Transform model; and an output step of generating an intrusion detection result and outputting a detection report. The method not only can effectively detect known and unknown network attacks, but also can provide valuable insight for security analysis, and provides a powerful tool for improving the security of the Internet of Things. The comprehensive and innovative method is expected to bring breakthrough progress in the field of Internet of Things security, and makes an important contribution to construction of a safer and more reliable Internet of Things ecosystem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to a Transformer-based Internet of Things intrusion detection method. Background Art

[0002] With the rapid development of IoT technology, more and more devices are connected to the network, which not only brings convenience but also brings unprecedented security challenges. Due to the limitations of their computing power and storage resources, IoT devices often lack a complete security protection mechanism and become vulnerable targets of network attacks. Therefore, an efficient and accurate IoT intrusion detection system becomes particularly important.

[0003] Traditional intrusion detection methods are mainly divided into signature-based detection and anomaly-based detection. Although signature-based detection methods are good at identifying known attacks, they are powerless against new or variant attacks. Although anomaly-based detection methods have the potential to discover unknown attacks, they often face the problem of high false alarm rates.

[0004] In recent years, with the advancement of machine learning technology, intrusion detection methods based on machine learning have gradually become a research hotspot. Among them, deep learning methods have attracted widespread attention due to their powerful feature extraction capabilities. However, existing deep learning methods still face many challenges when applied to IoT intrusion detection:

[0005] First, network traffic data in the IoT environment is characterized by high dimensionality, nonlinearity, and temporal dependency, which makes it difficult for traditional deep learning models to effectively capture the complex patterns. For example, although convolutional neural networks (CNNs) perform well in spatial feature extraction, they have difficulty in processing long-distance dependencies; and although recurrent neural networks (RNNs) can process sequence data, they are prone to gradient vanishing or gradient exploding problems when processing long sequences.

[0006] Secondly, the types of attacks in the IoT environment are diverse and constantly evolving, which requires the intrusion detection system to have strong generalization and adaptability. However, existing deep learning methods often perform poorly when facing unknown attacks or attack variants, making it difficult to meet the needs of practical applications.

[0007] Furthermore, the resource limitations of IoT devices require intrusion detection systems to have low computational complexity and memory usage. However, many advanced deep learning models often have large number of parameters and are difficult to deploy and run on resource-constrained IoT devices.

[0008] In addition, existing intrusion detection methods generally suffer from poor interpretability. In safety-critical IoT applications, it is crucial to understand the basis on which the model makes decisions, but most deep learning models are "black box" in nature and it is difficult to provide clear decision explanations.

[0009] Finally, data in IoT environments often have serious class imbalance problems, that is, normal traffic far outnumbers abnormal traffic. This imbalance causes the model to tend to misclassify minority classes, thus affecting detection performance. Summary of the invention

[0010] In view of the above problems, there is an urgent need for a new IoT intrusion detection method that can effectively process high-dimensional nonlinear data, has strong feature extraction and generalization capabilities, and at the same time has high computational efficiency, strong interpretability, and can effectively deal with data imbalance problems.

[0011] The present invention proposes a Transformer-based Internet of Things intrusion detection method, comprising:

[0012] The acquisition steps include:

[0013] Obtain network traffic data of IoT devices;

[0014] Processing steps include:

[0015] Based on the network traffic data, perform data preprocessing operations;

[0016] Based on the preprocessed data, feature extraction and intrusion detection are performed through the Transformer model;

[0017] Output steps include:

[0018] Generate intrusion detection results and output detection reports.

[0019] Preferably, the data preprocessing operation specifically includes:

[0020] Remove outliers;

[0021] Divide the dataset into training set, test set and validation set;

[0022] Normalize the data.

[0023] Preferably, the normalization process adopts a maximum and minimum value normalization method, and its calculation formula is:

[0024]

[0025] Among them, x j is the original eigenvalue, x′ j is the normalized value, xj,min represents the minimum value of the jth feature, x j,max Represents the maximum value of the j-th feature.

[0026] Preferably, the Transformer model comprises:

[0027] an encoder and a decoder;

[0028] Multi-head attention mechanism;

[0029] Positional encoding;

[0030] Residual connections and layer normalization.

[0031] Preferably, the calculation process of the multi-head attention mechanism includes:

[0032] Convert the input sequence into query (Q), key (K) and value (V) respectively;

[0033] Calculate attention weights;

[0034] Concatenate and linearly transform the results of multiple attention heads.

[0035] As an advantage, the model training step is also included:

[0036] Use cross entropy as loss function;

[0037] Adam optimizer is used to update parameters;

[0038] Apply labelsmoothing technology to alleviate the problem of data imbalance.

[0039] Preferably, the model training step further comprises:

[0040] Use early stopping strategy to prevent overfitting;

[0041] Adopt L2 regularization technique;

[0042] Use residual connections to accelerate model convergence.

[0043] Preferably, the model evaluation step is also included:

[0044] Calculate accuracy, false positive rate, and missed negative rate;

[0045] The Grad-CAM method is used to generate saliency maps to improve model interpretability.

[0046] Preferably, the input of the Transformer model includes:

[0047] Take the sample label as the query (Q);

[0048] The multi-dimensional features of the sample are taken as key-value pairs (KV).

[0049] Preferably, the step of sequence length processing is also included:

[0050] For sequences shorter than a predetermined length, a padding operation is performed;

[0051] For sequences longer than a predetermined length, a truncation operation is performed;

[0052] The predetermined length is determined based on the sequence length distribution in the training data set.

[0053] The Transformer-based IoT intrusion detection method proposed in this paper is designed to solve these technical problems. This method innovatively applies the Transformer model to the IoT intrusion detection task and combines a series of optimization strategies to effectively overcome the limitations of existing technologies.

[0054] Specifically, the method of the present invention first improves the quality and representativeness of the input data through a carefully designed data preprocessing step. Then, the powerful self-attention mechanism of the Transformer model is used to effectively capture long-distance dependencies and complex patterns in network traffic data. This not only improves the accuracy of detection, but also enhances the model's generalization ability to unknown attacks.

[0055] In addition, the present invention adopts a series of innovative training strategies, such as label smoothing and early stopping strategies, which effectively solve problems such as data imbalance and overfitting. At the same time, by introducing Grad-CAM visualization technology, the interpretability of the model is greatly improved, enabling security analysts to understand the decision basis of the model.

[0056] The method of the present invention also takes into special consideration the characteristics of the Internet of Things environment. Through a clever sequence length processing strategy, the model can adapt to network traffic data of different lengths, thereby enhancing its applicability in practical applications.

[0057] In general, the Transformer-based IoT intrusion detection method proposed in this paper has achieved significant improvements in accuracy, generalization ability, computational efficiency, interpretability, and practicality. It can not only effectively detect known and unknown network attacks, but also provide valuable insights for security analysis, providing a powerful tool for improving IoT security. This comprehensive and innovative approach is expected to bring breakthrough progress in the field of IoT security and make important contributions to building a safer and more reliable IoT ecosystem. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 This is the main process of the present invention;

[0059] Figure 2 The detailed process of data preprocessing of the present invention;

[0060] Figure 3 It is a logic block diagram of the Transformer model processing of the present invention;

[0061] Figure 4 A logic block diagram for model evaluation of the present invention; DETAILED DESCRIPTION

[0062] Please refer to Figure 1-4 The present invention provides an Internet of Things intrusion detection method based on Transformer. The method makes full use of the advantages of the Transformer model in processing sequence data and makes innovative improvements based on the characteristics of the Internet of Things environment.

[0063] First, the method of the present invention includes an acquisition step, a processing step, and an output step. In the acquisition step, the method acquires network traffic data of IoT devices. These data may come from various IoT devices, such as smart home devices, industrial IoT sensors, etc. Preferably, the present invention adopts real-time data acquisition technology to ensure that potential intrusion behaviors can be captured in a timely manner.

[0064] In the processing step, the method first performs data preprocessing operations based on the acquired network traffic data. Data preprocessing is a key link in the entire intrusion detection process, which directly affects the performance of the subsequent model. The present invention adopts a series of innovative preprocessing techniques to improve data quality and enhance the robustness of the model.

[0065] Specifically, data preprocessing operations include outlier removal, data set partitioning, and normalization. In the process of removing outliers, this method uses an outlier detection algorithm based on statistical principles. For example, the Z-score method can be used to treat data points that deviate from the mean by more than 3 standard deviations as outliers and delete them. This approach can effectively remove noise data and improve the accuracy of the model.

[0066] Next, the method divides the data set into a training set, a test set, and a validation set. Preferably, the division is performed by stratified sampling to ensure the consistency of the category distribution in each subset. For example, the division can be performed in a ratio of 60%, 20%, and 20%. This division method can provide sufficient data for model evaluation and tuning while ensuring sufficient model training.

[0067] In terms of normalization, the present invention adopts the maximum and minimum value normalization method. The calculation formula of this method is as follows:

[0068]

[0069] Among them, x j is the original eigenvalue, x′ j is the normalized value, x j,min represents the minimum value of the jth feature, x j,max Represents the maximum value of the jth feature. The advantage of this normalization method is that it can map all feature values ​​to the [0,1] interval, effectively eliminating the dimensional differences between different features, allowing the model to better learn the relationship between features.

[0070] After completing data preprocessing, this method uses the Transformer model for feature extraction and intrusion detection. The core of the Transformer model lies in its self-attention mechanism, which can effectively capture long-distance dependencies in sequence data. This feature is particularly important in the IoT intrusion detection scenario because attack behaviors are often reflected in the temporal patterns of network traffic.

[0071] The Transformer model of the present invention includes an encoder and a decoder. The encoder is responsible for encoding the input sequence and extracting deep features; the decoder decodes based on the output of the encoder to generate the final detection result. The model also includes key components such as multi-head attention mechanism, position encoding, residual connection and layer normalization.

[0072] The multi-head attention mechanism allows the model to focus on different parts of the input sequence at the same time and extract features from multiple perspectives. Position encoding provides the model with the position information of each element in the sequence, which is crucial for capturing temporal features. The introduction of residual connections and layer normalization helps alleviate the gradient vanishing problem in deep networks and improves the training efficiency and performance of the model.

[0073] Finally, in the output step, this method generates intrusion detection results and outputs a detection report. The detection results include not only the binary classification results of whether there is an intrusion behavior, but also the specific attack type and confidence. This fine-grained output can provide network administrators with more valuable information and help them take targeted defense measures in a timely manner.

[0074] Through the above steps, the method of the present invention realizes efficient and accurate IoT intrusion detection. Compared with traditional methods, this method has the following advantages: first, by comprehensively applying multiple preprocessing techniques, the quality of input data is significantly improved; second, by utilizing the powerful feature extraction capability of the Transformer model, complex attack patterns can be better captured; finally, through fine-grained output, more support is provided for subsequent security decisions.

[0075] In practical applications, this method can be flexibly adjusted according to the specific IoT environment. For example, for resource-constrained edge devices, a lightweight Transformer variant can be considered; for scenarios with particularly high security requirements, additional security measures can be added, such as introducing federated learning technology, to improve detection effects while protecting privacy.

[0076] In summary, the Transformer-based IoT intrusion detection method provided by the present invention innovatively applies deep learning technology to the field of network security, providing an effective solution to the increasingly severe IoT security problem.

[0077] A core innovation of the present invention lies in the architecture design of its Transformer model. The model includes an encoder and a decoder, a multi-head attention mechanism, position encoding, residual connection and layer normalization. This design fully considers the characteristics of IoT intrusion detection tasks and can effectively handle the complexity and timing of network traffic data.

[0078] In a preferred embodiment of the present invention, the encoder is responsible for mapping the input sequence to a high-dimensional representation space. Specifically, the encoder is composed of multiple identical layers stacked together, each of which contains two sublayers: a multi-head self-attention layer and a feedforward neural network layer. This structure enables the model to extract more abstract and meaningful features layer by layer. The structure of the decoder is similar to that of the encoder, but an additional attention layer is added after the self-attention layer to focus on the output of the encoder. This design enables the decoder to make full use of the feature information extracted by the encoder.

[0079] The multi-head attention mechanism is one of the core components of the Transformer model. Its calculation process includes converting the input sequence into query (Q), key (K) and value (V), calculating the attention weight, and concatenating and linearly transforming the results of multiple attention heads. Preferably, the present invention uses 8 attention heads, each with a dimension of 64. Such a setting enables the model to learn the relationship between features from different representation subspaces, greatly enhancing the ability of feature extraction.

[0080] Specifically, the calculation of multi-head attention can be expressed as:

[0081]

[0082] MultiHead(Q,K,V)=Concat(head1,...,head h )W O ,

[0083] Where Q, K, and V represent query, key, and value matrices, respectively. kis the dimension of the key, h is the number of attention heads, and W O is the output linear transformation matrix. Position encoding is to solve the problem that the Transformer model itself does not have the ability to perceive the sequence order. The present invention uses a combination of sine and cosine functions to generate position encoding. This method can not only provide position information for the model, but also has good extrapolation. The calculation formula of position encoding is as follows:

[0084]

[0085] Among them, pos represents the position, i represents the dimension, and d model is the dimension of the model. The introduction of residual connections and layer normalization effectively alleviates the gradient vanishing problem in deep network training. Residual connections allow information to be directly transferred from shallow layers to deep layers, while layer normalization helps stabilize the distribution of activation values ​​in deep networks. The combination of these two technologies greatly improves the training efficiency and final performance of the model.

[0086] In terms of model training, the present invention adopts a series of advanced training strategies. First, cross entropy is used as the loss function. For multi-class classification problems, the calculation formula of the cross entropy loss function is:

[0087]

[0088] Where C is the number of categories, y i is the true label, is the predicted probability.

[0089] This method selects the Adam optimizer for parameter update. The Adam optimizer combines the advantages of the momentum method and the adaptive learning rate method, and can converge quickly and adapt to the learning rates of different parameters. In one embodiment of the present invention, the initial learning rate is set to 0.001, β1=0.9, β2=0.999. In order to solve the data imbalance problem that is prevalent in the Internet of Things environment, the present invention innovatively introduces label smoothing technology. Labelsmoothing increases the robustness of the model and reduces the risk of overfitting by converting hard labels into soft labels. Specifically, for the original one-hot encoded label y, the new label y′ after label smoothing is calculated as follows:

[0090] y′ i =(1-α)y i +α / K,

[0091] Among them, α is a smoothing parameter, usually set to 0.1, and K is the number of categories. The present invention also uses techniques such as early stopping strategy, L2 regularization and residual connection to further optimize the model training process. The early stopping strategy decides when to stop training by monitoring the performance on the validation set, which effectively prevents overfitting. In the implementation of this method, if the validation set loss does not decrease for 5 consecutive epochs, the training is stopped. L2 regularization limits the complexity of the model by adding the L2 norm of the parameter to the loss function. The modified loss function can be expressed as:

[0092] L′=L+λ∑ w w 2 ,

[0093] Where L is the original loss, λ is the regularization coefficient, which is usually set to 0.0001, and w represents the model parameters. The introduction of residual connections not only accelerates the convergence of the model, but also alleviates the gradient vanishing problem to a certain extent. In the implementation of the present invention, the output of each sub-layer is added to its input to form a residual connection:

[0094] output=LayerNorm(x+Sublayer(x)),

[0095] Among them, x is the input of the sublayer, Sublayer(x) is the output of the sublayer, and LayerNorm represents the layer normalization operation. Through the above innovative design and optimization strategies, the Transformer-based IoT intrusion detection method of the present invention can effectively process the complex and changeable network traffic data in the IoT environment and achieve high-precision and high-efficiency intrusion detection.

[0096] The method of the present invention not only focuses on the performance of the model, but also pays special attention to its interpretability and practicality. The method includes a comprehensive model evaluation step, which is crucial to ensure the reliability of the model in a real IoT environment.

[0097] In the model evaluation stage, this method first calculates the three key indicators of accuracy, false positive rate and false negative rate. The accuracy reflects the overall ability of the model to correctly classify, and the calculation formula is:

[0098]

[0099] Among them, TP (True Positive) represents correctly identified attacks, TN (True Negative) represents correctly identified normal traffic, FP (False Positive) represents normal traffic misjudged as attacks, and FN (False Negative) represents unidentified attacks. The false positive rate and false negative rate reflect the sensitivity and omission of the model respectively, and their calculation formulas are as follows:

[0100]

[0101] In a preferred embodiment of the present invention, the goal of the model is to keep the false alarm rate and the missed alarm rate at a low level (such as below 5%) while maintaining a high accuracy rate (such as above 95%). This balance is crucial for the practical application of IoT intrusion detection systems, because too high a false alarm rate will lead to unnecessary alarms and waste of resources, while too high a missed alarm rate may expose the system to security risks. In addition to these conventional indicators, the present invention also innovatively introduces the Grad-CAM (Gradient-weighted Class Activation Mapping) method to improve the interpretability of the model. Grad-CAM is a visualization technology that can generate heat maps to show the input areas that the model focuses on when making decisions. In the scenario of IoT intrusion detection, this technology can help security analysts understand which features the model is based on to make judgments, thereby increasing the trust in the model's decisions. The core idea of ​​Grad-CAM is to use gradient information to determine the importance of each neuron to the final decision. Specifically, for the target category c, the class activation heat map generated by Grad-CAM The calculation is as follows:

[0102]

[0103] Among them, A k is the feature map, It is the feature map A k The importance weight of category c is calculated by k Global average pooling is performed with respect to the gradient of category c:

[0104]

[0105] Here, y c is the score of category c, and Z is the normalization factor. This approach not only improves the transparency of the model, but also enhances the user’s trust in the model’s prediction results, especially in critical safety application scenarios.

[0106] By analyzing these heat maps, we can identify which network traffic characteristics play a key role in the model's decision-making, which not only helps us understand and verify how the model works, but also provides guidance for further optimizing the model and designing more targeted defense strategies.

[0107] The Transformer model of the present invention is also unique in its input design. Specifically, the method uses the sample label as the query (Q) and the multi-dimensional features of the sample as the key-value pair (KV). The advantage of this design is that it enables the model to pay more attention to the characteristic patterns related to specific attack types.

[0108] In actual operation, suppose there are n samples, each sample has d features, and there are m categories in total. Then, the dimension of the query matrix Q is [n,m], and the dimension of the key-value matrix K and V are both [n,d]. With this setting, the model can learn the association between different categories (including normal traffic and various attack types) and features, thereby improving the accuracy and specificity of detection.

[0109] Finally, the present invention also includes an innovative sequence length processing step. This step is introduced to solve the problem of inconsistent network traffic data length in the Internet of Things environment.

[0110] In one embodiment of the present invention, a predetermined length is first determined based on the sequence length distribution in the training data set. The predetermined length can be selected as the 95% quantile of the sequence length distribution to cover most cases while avoiding excessive padding.

[0111] For sequences shorter than a predetermined length, this method uses a padding operation. Padding usually uses a special padding token, such as a zero vector or a special embedding learned. The padding operation can be expressed as:

[0112] X padded =[x1,x2,...,x n ,pad,pad,…,pad],

[0113] where x1,x2,...,x n is the original sequence, and pad is the padding token. For sequences longer than the predetermined length, this method uses truncation. In the scenario of IoT intrusion detection, it is generally believed that the most recent network activity is more representative, so the second half of the sequence is retained:

[0114] X truncated =[x m-L+1 ,x m-L+2 ,...,x m ],

[0115] Where m is the original sequence length and L is the predetermined length. This processing method ensures that all input sequences meet the expected input size requirements of the model while retaining the most informative parts as much as possible. The padding operation adds sufficient pad tokens to make the shorter sequence reach the preset length, avoiding data processing errors caused by inconsistent sequence lengths. The truncation strategy for overly long screen columns gives priority to the data at the end of the sequence, because these data often better reflect the current state or mode, which is especially important for IoT intrusion detection systems with high real-time requirements. This can not only ensure the efficiency of the model training and reasoning process, but also improve the model's sensitivity and responsiveness to the latest changes in network activities.

[0116] In general, the Transformer-based IoT intrusion detection method proposed in this paper effectively solves many challenges faced by intrusion detection in the IoT environment through a series of innovative designs and optimizations. From data preprocessing, model architecture design, training strategy optimization to result interpretation and practical considerations, this method has made comprehensive and in-depth innovations, providing a powerful tool for improving the security of the IoT.

[0117] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A Transformer-based IoT intrusion detection method, characterized in that: include: The acquisition steps include: Obtain network traffic data of IoT devices; Processing steps include: Based on the network traffic data, perform data preprocessing operations; Based on the preprocessed data, feature extraction and intrusion detection are performed through the Transformer model; Output steps include: Generate intrusion detection results and output detection reports.

2. The method according to claim 1, characterized in that The data preprocessing operation specifically includes: Remove outliers; Divide the dataset into training set, test set and validation set; Normalize the data.

3. The method according to claim 2, characterized in that The normalization process adopts the maximum and minimum value normalization method, and its calculation formula is: Among them, x j is the original eigenvalue, x′ j is the normalized value, x j,min represents the minimum value of the jth feature, x j,max Represents the maximum value of the j-th feature.

4. The method according to claim 1, characterized in that The Transformer model includes: an encoder and a decoder; Multi-head attention mechanism; Positional encoding; Residual connections and layer normalization.

5. The method according to claim 4, characterized in that The calculation process of the multi-head attention mechanism includes: Convert the input sequence into query (Q), key (K) and value (V) respectively; Calculate attention weights; Concatenate and linearly transform the results of multiple attention heads.

6. The method according to claim 1, characterized in that Also includes the model training steps: Use cross entropy as loss function; Adam optimizer is used to update parameters; Apply label smoothing technology to alleviate the problem of data imbalance.

7. The method according to claim 6, characterized in that The model training step also includes: Use early stopping strategy to prevent overfitting; Adopt L2 regularization technique; Use residual connections to accelerate model convergence.

8. The method according to claim 1, characterized in that A model evaluation step is also included: Calculate accuracy, false positive rate, and missed negative rate; The Grad-CAM method is used to generate saliency maps to improve model interpretability.

9. The method according to claim 1, characterized in that: The input of the Transformer model includes: Take the sample label as the query (Q); The multi-dimensional features of the sample are taken as key-value pairs (KV).

10. The method according to claim 1, characterized in that It also includes a sequence length processing step: For sequences shorter than a predetermined length, a padding operation is performed; For sequences longer than a predetermined length, a truncation operation is performed; The predetermined length is determined based on the sequence length distribution in the training data set.

Citation Information

Cited By

  • Mosquito killer lamp fault early warning and operation and maintenance system based on big data

    CN122090565A