Internet of Things Intrusion Detection Method Based on Self-Supervised Learning and Self-Knowledge Distillation
By using self-supervised learning and self-knowledge distillation technology in the Internet of Things intrusion detection system, the lightweight intrusion detection model is trained, and the existing system's insufficient efficiency and generalization capabilities on resource-constrained devices are solved, and efficient and real-time intrusion detection is achieved.
Patent Information
- Application Number
- CN202210446932.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-04-26
AI Technical Summary
Existing IoT intrusion detection systems are difficult to achieve lightweight, real-time and unsupervised intrusion detection on resource-constrained devices, and they rely too much on label data and lack generalization capabilities.
Using self-supervised learning and self-knowledge distillation methods, through self-supervised comparative learning and knowledge distillation technology, lightweight intrusion detection models are trained to reduce dependence on label data, and the generalization ability and characterization learning ability of the model are improved.
It realizes that without reducing model efficiency, compressing the model size, improving generalization capabilities, avoiding excessive dependence on label data, improving the detection speed of intrusion detection, and reducing model complexity.
Smart Images

Figure CN114861875B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Internet of Things, and in particular relates to an Internet of Things intrusion detection method based on self-supervised learning and self-knowledge distillation. Background Art
[0002] The Internet of Things is an extension of the Internet, an Internet where everything is connected. Its core and foundation is still the Internet, and it is an extension and expansion of the Internet. The Internet of Things is a network that connects any object to the Internet through sensor devices and agreed protocols to exchange information and communicate, so as to realize the intelligent identification, positioning, tracking, monitoring and management of objects. The rise of the Internet of Things technology has changed the new direction of the information world and is considered to be the third wave of information development after computers and the Internet. Today, the Internet of Things technology is quietly changing our lifestyle and providing various conveniences for our lives, such as smart cities, health care, smart homes, smart wearable devices, etc. However, due to the lack of relevant network security knowledge among the owners of Internet of Things devices, hackers attack network physical devices, such as wearable devices, medical pacemakers, car autonomous driving, or expensive industrial processes controlled by connected devices, etc., and the privacy data of individuals or companies is stolen, resulting in huge property losses, and sometimes even serious life safety accidents.
[0003] Although network security experts have taken many efforts to improve the security of the Internet of Things, including encrypting data transmitted over the network, regularly updating firmware, using strong passwords and security keys, etc. However, even with the above countermeasures, IoT devices are still vulnerable to various network attacks due to their diversity. How to reduce the harm of IoT device intrusion has become a focus of close attention in the industry. As an important part of network security, intrusion detection systems have become an important means of network attack detection. According to different detection technologies, intrusion detection systems can be divided into misuse-based intrusion detection and anomaly-based intrusion detection. However, misuse-based intrusion detection systems are highly dependent on existing signature knowledge bases, making it difficult to detect zero-day attacks and cannot be applied to the detection of unknown attacks. Anomaly-based intrusion detection detects abnormal behavior of the system. When the detected behavior deviates greatly from the normal behavior, an alarm message is issued. At the same time, it can rely on existing intrusion detection data sets to train machine learning and deep learning algorithms to identify specific network attack categories.
[0004] In recent years, a large number of intrusion detection systems based on machine learning and deep learning have been widely used in attack detection of IoT devices. However, there are still many challenges in detecting abnormal traffic in IoT. First, network nodes in IoT are usually deployed in devices with limited resources (such as limited power, limited computing, communication and storage capabilities, etc.); second, the cost of obtaining attack tag data is expensive and time-consuming, and it requires the assistance of network security experts to determine whether network traffic is a new attack method; in addition, IoT networks use different protocol stacks and standards. These requirements require the intrusion detection system to design corresponding security mechanisms. Therefore, a good IoT intrusion detection system needs to meet the characteristics of lightweight, real-time and unsupervised. However, most of the existing intrusion detection systems only meet one of the three characteristics, and there are few IoT intrusion detection systems that meet all three characteristics.
[0005] Therefore, it is worth studying how to compress the size of the model without reducing the efficiency of the intrusion detection model, improve the generalization ability of the model, and avoid over-reliance on labeled data intrusion detection technology. Summary of the invention
[0006] To this end, the present invention provides an Internet of Things intrusion detection method based on self-supervised learning and self-knowledge distillation, which realizes lightweight, real-time and unsupervised Internet of Things intrusion detection, reduces excessive dependence on labels, and improves generalization ability.
[0007] In order to achieve the above object, the present invention provides the following technical solution: an Internet of Things intrusion detection method based on self-supervised learning and self-knowledge distillation, comprising:
[0008] (1) performing data preprocessing on the intrusion detection data set, wherein the data preprocessing includes character data hot-single encoding and data normalization processing;
[0009] (2) First phase training of lightweight intrusion detection model:
[0010] (21) Determine the network structure of the online network and the target network, and use the weights of the online network to initialize the parameters of the target network;
[0011] (22) Input the enhanced data into the online network and the target network for training respectively;
[0012] (23) adjusting the error of the training process according to the loss value obtained by the loss function of self-supervised contrastive learning until the online network reaches convergence;
[0013] (24) Save the weights of the online network locally for use in the second stage of training;
[0014] (3) Second stage training of lightweight intrusion detection model:
[0015] (31) Determine the network structure of the student network and load the online network weights obtained from the first stage of training into the teacher network;
[0016] (32) Inputting the enhanced data into the student network and the teacher network for training;
[0017] (33) Adjust the error of the training process according to the loss value obtained from the loss function of self-knowledge distillation until the student network reaches convergence;
[0018] (34) Save the student network weights locally for lightweight intrusion detection model testing.
[0019] As a preferred solution for the Internet of Things intrusion detection method based on self-supervised learning and self-knowledge distillation, the online network and the target network are both asymmetric neural networks, and both the online network and the target network include a feature encoder and a feature projector; the feature projector of the online network is also added with a feature predictor.
[0020] As a preferred solution of the IoT intrusion detection method based on self-supervised learning and self-knowledge distillation, in step (22), a first-in-first-out memory queue is maintained, and the memory queue is composed of feature embeddings encoded by feature encoders of the online network.
[0021] As a preferred solution of the IoT intrusion detection method based on self-supervised learning and self-knowledge distillation, the feature encoder, feature projector and feature predictor of the online network update parameters by back-propagating loss;
[0022] The feature encoder feature projector in the target network updates parameters by momentum updating.
[0023] As a preferred solution for the IoT intrusion detection method based on self-supervised learning and self-knowledge distillation, the feature encoder is composed of a convolutional neural network; the feature projector and the feature predictor are both composed of a multi-layer perceptron, which includes a hidden layer, a BN layer, a ReLU activation function and a hidden layer.
[0024] As a preferred solution of the IoT intrusion detection method based on self-supervised learning and self-knowledge distillation, in step (32), a first-in-first-out memory queue is constructed, and during training, the latest batch is put into the memory queue, and the oldest batch in the memory queue is taken out of the memory queue, and a group of unlabeled network traffic is sent to the teacher network pre-trained by self-supervised comparative learning, and the obtained feature embedding is added to the memory queue, and at the same time, the unlabeled network traffic is sent to the student network to obtain another group of feature embedding;
[0025] The supervisory signal for the knowledge distillation process is obtained by constraining the distance between the two sets of feature embeddings and the feature embeddings in the memory queue.
[0026] As a preferred solution for the IoT intrusion detection method based on self-supervised learning and self-knowledge distillation, in step (33), the weights of the student network are updated by the back-propagation algorithm, and the characteristic representation of network traffic learned by the student network is transferred to the anomaly detection of the intrusion detection dataset.
[0027] As a preferred solution for IoT intrusion detection based on self-supervised learning and self-knowledge distillation, deep separable convolution is integrated into anomaly detection of intrusion detection;
[0028] Depth-wise separable convolution decouples a complete convolution operation into two steps. For the multi-channel feature maps from the previous layer, they are first split into feature maps of a single channel and perform single-channel convolution separately, and then they are stacked together again for channel-by-channel convolution.
[0029] As a preferred solution of the IoT intrusion detection method based on self-supervised learning and self-knowledge distillation, in step (1), character data is converted into numerical data by one-hot encoding of the character data;
[0030] In step (1), in the data normalization process, a hybrid data normalization method is used to normalize the data.
[0031] As a preferred solution for the IoT intrusion detection method based on self-supervised learning and self-knowledge distillation, it also includes:
[0032] (4) Lightweight intrusion detection model testing process: load the student network weights and input the preprocessed test data set into the student network to obtain the classification results of each data.
[0033] The invention has the following advantages: data preprocessing is performed on an intrusion detection data set, and the data preprocessing includes character data hot encoding and data normalization processing; the first stage training of a lightweight intrusion detection model: determining the network structure of an online network and a target network, and using the weight of the online network to initialize the target network parameters; inputting enhanced data into the online network and the target network for training respectively; adjusting the error of the training process according to the loss value obtained by the loss function of self-supervised contrastive learning until the online network reaches convergence; saving the weight of the online network locally for the second stage training; the second stage training of a lightweight intrusion detection model: determining the network structure of a student network, loading the online network weight obtained by the first stage training into the teacher network; inputting enhanced data into the student network and the teacher network for training; adjusting the error of the training process according to the loss value obtained by the loss function of self-knowledge distillation until the student network reaches convergence; saving the student network weight locally for lightweight intrusion detection model testing. The invention avoids over-reliance on label data without reducing the abnormal detection capability of the model; improves the generalization capability and the ability of characterizing learning of the model; can improve the detection speed of model intrusion detection, and reduces the complexity of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the drawings required for the implementation methods or the description of the prior art are briefly introduced below. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.
[0035] Figure 1 A schematic diagram of an Internet of Things intrusion detection method based on self-supervised learning and self-knowledge distillation provided by an embodiment of the present invention;
[0036] Figure 2 A lightweight intrusion detection model framework in an Internet of Things intrusion detection method based on self-supervised learning and self-knowledge distillation provided in an embodiment of the present invention;
[0037] Figure 3 The self-supervised contrastive learning training process in the Internet of Things intrusion detection method based on self-supervised learning and self-knowledge distillation provided in an embodiment of the present invention;
[0038] Figure 4 A feature encoder framework in an Internet of Things intrusion detection method based on self-supervised learning and self-knowledge distillation provided in an embodiment of the present invention;
[0039] Figure 5 The present invention provides a self-knowledge distillation training process in an Internet of Things intrusion detection method based on self-supervised learning and self-knowledge distillation. DETAILED DESCRIPTION
[0040] The following is a description of the implementation of the present invention by specific embodiments. People familiar with the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] Assistance Figure 1 The specific steps of the present invention using the lightweight intrusion detection model for intrusion detection are as follows:
[0042] (1) performing data preprocessing on the intrusion detection data set, wherein the data preprocessing includes character data hot-single encoding and data normalization processing;
[0043] (2) First phase training of lightweight intrusion detection model:
[0044] (21) Determine the network structure of the online network and the target network, and use the weights of the online network to initialize the parameters of the target network;
[0045] (22) Input the enhanced data into the online network and the target network for training respectively;
[0046] (23) adjusting the error of the training process according to the loss value obtained by the loss function of self-supervised contrastive learning until the online network reaches convergence;
[0047] (24) Save the weights of the online network locally for use in the second stage of training;
[0048] (3) Second stage training of lightweight intrusion detection model:
[0049] (31) Determine the network structure of the student network and load the online network weights obtained from the first stage of training into the teacher network;
[0050] (32) Inputting the enhanced data into the student network and the teacher network for training;
[0051] (33) Adjust the error of the training process according to the loss value obtained from the loss function of self-knowledge distillation until the student network reaches convergence;
[0052] (34) Save the student network weights locally for lightweight intrusion detection model testing.
[0053] See also Figure 2The lightweight intrusion detection model used in the present invention is named CL-SKD model, which can detect intrusion behavior of IoT devices. The lightweight intrusion detection model is divided into two stages. The first stage uses self-supervised contrastive learning to learn the feature representation of the essence of network traffic. The second stage uses self-knowledge distillation to transfer the feature representation of network traffic learned by the large convolutional neural network model to a small deep separable convolutional network model.
[0054] During the implementation of the present invention, not only the views of the same network traffic image under different enhancements are brought closer, but also the nearest neighbors under the enhanced views are brought closer, because the nearest neighbors under the enhanced view of the network traffic may be of the same category as it, so the nearest neighbors cannot be judged as negative examples, thereby increasing the distance between the network traffic image and them. The present invention maintains two asymmetric neural networks, namely the online network (Online Network) and the target network (TargetNetwork). Both the Online Network and the Target Network are composed of a feature encoder (Encoder) and a feature projector (Projector). However, the Online Network also adds a feature predictor (Predictor) after the feature projector. Assume that the Online Encoder and the Target Encoder are respectively composed of f θ and f ζ Indicates that OnlineProjector and Target Projector are respectively represented by g θ and g ζ Online Predictor is composed of q θ Indicates, where θ and ζ represent the weights of the Online Network and Target Network respectively. In order to save the nearest neighbors under different enhanced views, the present invention maintains a first-in-first-out memory queue, which is composed of feature embeddings encoded by the Online Encoder. During training, the latest batch is queued and the oldest batch in the queue is dequeued.
[0055] See also Figure 3 Given an unlabeled network traffic image x, the present invention performs two different data augmentation operations T on x. 1 and T 2 Get two different views of the network traffic T 1 (x) and T 2 (x), then T 1 (x) and T 2 (x) are respectively sent to the feature encoder f θ and f ζ We get two different sets of feature embeddings y 1and 2 , that is, y 1 =f θ (T 1 (x)), y 2 =f ζ (T 2 (x)), and then embed the two different sets of features into y 1 and 2 Feed into feature projector g θ and g ζ Get two different sets of feature projections z 1 and z 2 , that is, z 1 =g θ (y 1 ), z 2 =f ζ (y 2 ), then z 1 Send it to the feature predictor to get the feature query query, that is, q = q θ (z 1 ), then q and z 2 Do it once each 2 Regularization, that is:
[0056]
[0057]
[0058] The present invention first Add to the memory queue, and then find in the memory queue The K nearest neighbors of get a set of feature embeddings as Due to need and The feature embedding average distance in is the smallest, and the present invention can minimize the following loss function:
[0059]
[0060] Wherein, dist(p, q) represents the distance metric between two feature embeddings. The present invention can use MSE loss, that is, as the distance between two feature embeddings.
[0061] Different from the prior art, the present invention takes into account that the feature encoder cannot obtain good feature embedding of the network traffic image in the early stage of training, so it finds The feature embedding obtained by the K nearest neighbors It does not represent this set of feature embeddings and For the same class of network traffic, on the contrary, the present invention finds A set of feature embeddings obtained by the K farthest neighbors This set of features is embedded and The average distance L farthest As part of the model loss function, and as the model is continuously trained, the feature encoder can get a good feature embedding of the network traffic image, L farthest The loss weight gradually decreases to 0, so the loss function of the improved self-supervised contrastive learning of the present invention is:
[0062]
[0063] Where α is the weight coefficient. The present invention uses a linear reduction method to control the size of α, that is, the value of α at the tth epoch can be calculated using the following formula:
[0064]
[0065] Where T and t are the total number of rounds in training and the number of rounds in current training, respectively. After a large number of experiments, the present invention finds that when t=T / 2, the feature encoder can obtain better feature embedding of the network traffic image, so in the subsequent model training, the present invention sets α to 0.
[0066] In the Online Network, the feature encoder fθ and the feature projector g θ and feature predictor q θ The parameters are updated by back-propagating the loss, and the feature encoder f in the Target Network ζ and feature projector g ζ The parameters are updated by momentum update, that is:
[0067] ζ←η*ζ+(1-η)*θ (6)
[0068] Among them, η∈[0, 1] controls the degree to which the parameters in the Target Network depend on the current parameters and is a manually set hyperparameter. The specific steps of self-supervised contrastive learning are shown in Algorithm 1:
[0069]
[0070]
[0071] bmm:batch matrix multiplication
[0072] The present invention takes into account that the convolutional neural network has excellent feature extraction capabilities, so the feature encoder is composed of a convolutional neural network. The specific feature encoder architecture is shown in Figure 4 ,The feature projector and feature predictor are both composed of multi-layer perceptrons, i.e. hidden layer + BN layer + ReLU activation function + hidden layer.
[0073] In this embodiment, knowledge distillation is a common model compression method, which was first proposed by Hinton in the image classification task. Different from pruning and quantization in model compression, knowledge distillation is to build a lightweight small model and use the supervision information of a large model with better performance and pre-trained on a large data set to train the small model so that the small model can achieve better performance and accuracy. The large model is usually called the teacher network and the small model is called the student network. Due to the limitations of power, computing, communication and storage capabilities, it is impossible to deploy a large model on an IoT device. Therefore, it is necessary to transfer the knowledge learned by the large model to the small model through "distillation" so that it can be deployed on a resource-constrained IoT device to detect intrusion behavior.
[0074] However, traditional knowledge distillation algorithms require label information of network traffic to guide the process of knowledge distillation. As we all know, abnormal data in IoT devices is difficult to obtain. Therefore, how to construct the supervisory signal in the process of knowledge distillation is critical. The present invention constructs a first-in-first-out memory queue, puts the latest batch into the queue during training, and takes the oldest batch out of the queue. A group of unlabeled network traffic is sent to the pre-trained teacher network of self-supervised comparative learning to obtain feature embedding and added to the memory queue. At the same time, this group of network traffic is sent to the student model to obtain another group of feature embedding. The supervisory signal of the knowledge distillation process is obtained by constraining the distance between the two groups of feature embeddings and the feature embeddings in the memory queue.
[0075] See also Figure 5 , the present invention represents the teacher network and the student network as and Where θ and ζ are the weights of the teacher network and the student network respectively.
[0076] Given an unlabeled network traffic image x, a data augmentation operation is first performed on x to obtain x′, and then x′ is sent to the teacher network and student networks Get two sets of feature embeddings and Right now Then, respectively and Do it once 2 Regularization, that is:
[0077]
[0078]
[0079] Then Add to the memory queue, assuming that the memory queue is represented by Q = {q 1 ,q 2 ,q 3 , ..., q K}, q j is the feature embedding obtained by the teacher network. Then calculate the feature embedding of the teacher network The distance P to all feature embeddings in the memory queue Q T (x i ,θ,Q) is:
[0080]
[0081] where τ T is the temperature parameter of the teacher network, (·) is the inner product between two feature embeddings, and K is the length of the memory queue.
[0082] Similarly, the feature embedding of the student network The distance P to all feature embeddings in the memory queue Q S (x i ,ζ,Q) is:
[0083]
[0084] where τ S is the temperature parameter of the student network.
[0085] The present invention is to make P T (x i ,θ,Q) and P S (x i ,ζ,Q) are similar in distribution, and the cross entropy loss of the two is used as the loss function of self-knowledge distillation, that is:
[0086]
[0087] Usually, in order to maintain the consistency of the memory queue, the K value needs to be set very large, so that the model can observe more negative samples and thus improve the performance of the model. The present invention fully considers the limited computing power of IoT devices. The larger the K value is set, the more computational effort the model will have. In addition, most of the feature embeddings maintained in the memory queue Q are random and have nothing to do with the target feature embeddings. Therefore, when calculating P T (x i ,θ,Q) and P S (x i ,ζ,Q), most of the elements are very small, which leads to T (x i ,θ,Q) and P S (x i,ζ,Q), most elements contribute very little to the overall loss and can be basically ignored. Calculating the loss of these elements is a waste of computational power.
[0088] Therefore, the present invention can be used in the feature embedding of the teacher network The distance P to all feature embeddings in the memory queue Q T (x i ,θ,Q), find the feature embedding The k nearest neighbors of get the feature embedding distance P′ T (x i ,θ,Q T ),in Then in the feature embedding of the student network The distance P to all feature embeddings in the memory queue Q S (x i ,ζ,Q) to find the feature embedding The k nearest neighbors are used to obtain the feature embedding distance P′ S (x i ,θ,Q S ),in Where k<<K, then calculate P′ T (x i ,θ,Q T ) and P′ S (x i ,θ,Q S )The cross entropy loss of the two is used as the loss function of self-knowledge distillation, that is:
[0089]
[0090] After obtaining the loss function of self-knowledge distillation, the present invention updates the weights of the student network through the back propagation algorithm. The student network finally obtained can learn the excellent feature representation of network traffic and can be transferred to the anomaly detection of other intrusion detection data sets. The specific steps of self-knowledge distillation are shown in Algorithm 2:
[0091]
[0092]
[0093] bmm: batch matrix multiplication
[0094] In the embodiment of the present invention, in order to achieve real-time intrusion detection, the model complexity and the amount of model calculation are reduced from two aspects: first, considering that knowledge distillation can "distill" the feature representation learned by the complex teacher network with strong learning ability, and pass it to the student network with small parameters and weak learning ability, the present invention uses knowledge distillation to transfer the representation of network traffic learned by the large model to the small model; second, considering that the deep separable convolution can both extract data features and reduce the computational burden at the same time, the present invention replaces the traditional convolution with the deep separable convolution, which greatly reduces the parameter amount and computational cost of the model.
[0095] Assistance Figure 4 , the standard convolution operation extracts features from all three dimensions of each image, including the width and height of the image and the channel dimension, while the depthwise separable convolution decouples a complete convolution operation into two steps. For the multi-channel feature maps from the previous layer, they are first split into feature maps of a single channel, and single-channel convolution is performed on them separately, and then stacked together again, which is the so-called channel-by-channel convolution. In the channel-by-channel convolution, only the size of the feature map from the previous layer is adjusted, and the number of channels does not change. So the feature map obtained by the channel-by-channel convolution is convolved for the second time. The convolution kernel size of this convolution process is 1×1, and the filter contains the same number of convolution kernels as the number of output channels of the previous layer. One filter outputs a feature map, so multiple channels require multiple filters, which is pointwise convolution. By decoupling a complete convolution, the depthwise separable convolution can avoid extracting some redundant features and greatly reduce the number of parameters required, thereby reducing the risk of model overfitting.
[0096] For channel-by-channel convolution, the calculation method of the convolution parameter is determined by formula (13):
[0097] num_params = W k *H k *in_channels (13)
[0098] The amount of calculation is determined by formula (14):
[0099] Flops = W k *H k *W img *H img *in_channels (14)
[0100] Among them, W k and W img are the width of the convolution kernel and the input feature map, respectively,k and H img They are the height of the convolution kernel and the input feature map respectively, and in_channels is the number of input channels.
[0101] For point-by-point convolution, the calculation method of the convolution parameter is determined by formula (15):
[0102] num_params=1*1*in_channels*out_channels (15)
[0103] The amount of calculation is determined by formula (16):
[0104] Flops=1*1*W feature *H feature *in_channels*out_channels (16)
[0105] Among them, W feature and H feature The width and height of the input feature map are respectively, in_channels and out_channels are the number of input channels and output channels respectively.
[0106] For traditional convolution, the calculation method of the convolution parameter is determined by formula (17):
[0107] num-params=W k *H k *in_channels*out_channels (17)
[0108] The amount of calculation is determined by formula (18):
[0109] Flops = W k *H k *W feature *H feature *in_channels*out_channels (18)
[0110] After preprocessing, the network traffic is converted into a grayscale format of 14×14×1. If the network traffic is not enough to be converted into 14×14×1, the insufficient part is filled with 0. In the self-knowledge distillation process, the preprocessed network traffic is input into the teacher network and the student network respectively after data enhancement. Therefore, the parameters and calculation amount required for the teacher network and the student network can be calculated respectively by formulas (13)-(18). The parameters required for the teacher network and the student network (excluding the parameters in the classification header) are 346912 and 13401 respectively, and the calculation amount required for the teacher network and the student network (excluding the parameters in the classification header) are 15893632 and 673828 respectively. Therefore, the present invention can conclude that the parameters required for the student network only account for 3.9% of the teacher network, and the calculation amount required for the student network only account for 4.2% of the teacher network. It can be seen that by replacing traditional convolution with deep separable convolution, 96.1% of the parameters and 95.8% of the computational effort can be saved. Moreover, through knowledge distillation, the present invention can transfer the representation of network traffic learned by the teacher model to the student model to improve the intrusion detection performance of the student model. Therefore, deep separable convolution can be applied to the intrusion detection model, which can be deployed in nodes with limited computing power and storage capacity in the Internet of Things network. This can greatly reduce the time required for intrusion detection and meet the real-time requirements.
[0111] In order to verify the excellent anomaly detection ability and excellent generalization ability of the CL-SKD model in IoT intrusion detection, a large number of binary and multi-classification experiments will be conducted on the IoT intrusion detection datasets KDD CUP99, NSL-KDD, CIC IDS2017, UNSW-NB15 and CIDDS-001. Since the UNSW-NB15 dataset contains a relatively comprehensive range of attack types, rich feature information and sufficient data, the present invention performs self-supervised comparative learning on UNSW-NB15 to obtain the feature representation of network traffic, and uses self-knowledge distillation to migrate it to the student network.
[0112] Since the present invention adopts a convolutional neural network as the backbone network of the intrusion detection model, the input network traffic must conform to the input format of the convolutional neural network, so the intrusion detection data set needs to be preprocessed, which mainly includes two steps: one-hot encoding processing of character data and data normalization processing.
[0113] One-hot encoding of character data. Taking the UNSW-NB15 dataset as an example, the three features of proto, state, and service are of character type, while the input data of the convolutional neural network must be numeric. Therefore, it is necessary to convert the character data into numeric data. There are usually two ways to achieve this goal, namely one-hot encoding and ordinal encoding. However, experiments have found that one-hot encoding has better results, so one-hot encoding is used to convert character data into numeric data.
[0114] In the data normalization process, since different features have different dimensions, in order to eliminate the influence of different dimensions, data normalization is required. Min_Max normalization, as the most commonly used normalization method in machine learning, has certain defects. That is, if the maximum and minimum values in the data are very far apart, it is easy to make the normalization result unstable. Therefore, a hybrid data normalization method is used to normalize the data. The specific method is shown in formula (19):
[0115]
[0116] in, is the normalized result, x i is the value of the i-th feature, is the minimum value of the i-th feature, is the maximum value of the i-th feature.
[0117] All experiments of the present invention are simulated on Windows 10 operating system, using Python 3.7 as programming language, Pytorch 1.7 as deep learning framework, Scikit-learning 0.23.2 as machine learning framework, and RTX 2070 graphics card to accelerate training. A large number of binary and multi-classification experiments on the CL-SKD model for IoT intrusion detection datasets KDD CUP99, NSL-KDD, UNSW-NB15, CIC IDS2017 and CIDDS-001 show that the self-supervised learning and self-knowledge distillation schemes of the present invention have strong feasibility; they can outperform the SOTA models in recent years in terms of accuracy, precision, recall and F1-measure; the CL-SKD model learns the representation of network traffic with strong generalization ability, and can achieve relatively excellent intrusion detection performance without changing the weight of the feature extraction layer and only training the classification head.
[0118] Although the present invention has been described in detail above by general description and specific embodiments, it is obvious to those skilled in the art that some modifications or improvements can be made to the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all belong to the scope of protection claimed by the present invention.
Claims
1. IoT intrusion detection method based on self-supervised learning and self-knowledge distillation, It is characterized in that include: (1) performing data preprocessing on the intrusion detection data set, wherein the data preprocessing includes character data hot-single encoding and data normalization processing; (2) First phase training of lightweight intrusion detection model: (21) Determine the network structure of the online network and the target network, and use the weights of the online network to initialize the parameters of the target network; (22) Input the enhanced data into the online network and the target network for training respectively; (23) adjusting the error of the training process according to the loss value obtained by the loss function of self-supervised contrastive learning until the online network reaches convergence; (24) Save the weights of the online network locally for use in the second stage of training; (3) Second stage training of lightweight intrusion detection model: (31) Determine the network structure of the student network and load the online network weights obtained from the first stage of training into the teacher network; (32) Inputting the enhanced data into the student network and the teacher network for training; (33) Adjust the error of the training process according to the loss value obtained from the loss function of self-knowledge distillation until the student network reaches convergence; (34) Save the student network weights locally for lightweight intrusion detection model testing.
2. The method for intrusion detection in the Internet of Things based on self-supervised learning and self-knowledge distillation according to claim 1, It is characterized in that Both the online network and the target network are asymmetric neural networks, and both the online network and the target network include a feature encoder and a feature projector; the feature projector of the online network is also added with a feature predictor.
3. The method for intrusion detection in the Internet of Things based on self-supervised learning and self-knowledge distillation according to claim 2, It is characterized in that In step (22), a first-in-first-out memory queue is maintained, and the memory queue is composed of feature embeddings encoded by the feature encoder of the online network.
4. The method for intrusion detection in the Internet of Things based on self-supervised learning and self-knowledge distillation according to claim 3, It is characterized in that The feature encoder, feature projector and feature predictor of the online network update parameters by back-propagating losses; The feature encoder feature projector in the target network updates parameters by momentum updating.
5. The method for intrusion detection in the Internet of Things based on self-supervised learning and self-knowledge distillation according to claim 4, It is characterized in that The feature encoder is composed of a convolutional neural network; the feature projector and feature predictor are both composed of a multi-layer perceptron, which includes a hidden layer, a BN layer, a ReLU activation function and a hidden layer.
6. The method for intrusion detection in the Internet of Things based on self-supervised learning and self-knowledge distillation according to claim 5, It is characterized in that In step (32), a first-in-first-out memory queue is constructed. During training, the latest batch is put into the memory queue, and the oldest batch in the memory queue is taken out of the memory queue. A set of unlabeled network traffic is sent to the teacher network pre-trained by self-supervised contrastive learning, and the obtained feature embedding is added to the memory queue. At the same time, the unlabeled network traffic is sent to the student network to obtain another set of feature embeddings. The supervisory signal for the knowledge distillation process is obtained by constraining the distance between the two sets of feature embeddings and the feature embeddings in the memory queue.
7. The method for intrusion detection in the Internet of Things based on self-supervised learning and self-knowledge distillation according to claim 6, It is characterized in that In step (33), the weights of the student network are updated through the back propagation algorithm, and the characteristic representation of network traffic learned by the student network is transferred to the anomaly detection of the intrusion detection dataset.
8. The method for intrusion detection in the Internet of Things based on self-supervised learning and self-knowledge distillation according to claim 7, It is characterized in that Incorporating depthwise separable convolution into anomaly detection for intrusion detection; Depth-wise separable convolution decouples a complete convolution operation into two steps. For the multi-channel feature maps from the previous layer, they are first split into feature maps of a single channel and perform single-channel convolution separately, and then they are stacked together again for channel-by-channel convolution.
9. The method for intrusion detection in the Internet of Things based on self-supervised learning and self-knowledge distillation according to claim 1, It is characterized in that In step (1), character data is converted into numerical data by one-hot encoding; In step (1), in the data normalization process, a hybrid data normalization method is used to normalize the data.
10. The method for intrusion detection in the Internet of Things based on self-supervised learning and self-knowledge distillation according to claim 1, It is characterized in that Also includes: (4) Lightweight intrusion detection model testing process: load the student network weights and input the preprocessed test data set into the student network to obtain the classification results of each data.