Internet of Things intrusion detection method integrating label propagation and fuzzy label distribution
By integrating the method of label propagation and fuzzy label allocation, accurate pseudo-labels are allocated to unlabeled traffic data in IoT intrusion detection, solving the problems of insufficient labeled data and low pseudo-label accuracy, and improving detection performance.
Patent Information
- Application Number
- CN202510067733.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
In IoT intrusion detection, the pseudo-label accuracy of labeled traffic data and the unlabeled traffic data is low, resulting in a degradation of model overfitting and detection performance.
Using the method of integrated label propagation and fuzzy label allocation, pseudo-labels are assigned to unlabeled traffic data through label propagation and ambiguity calculation, and a variety of pseudo-label allocation methods are integrated through integrated learning to improve the accuracy of pseudo-labels.
It improves the pseudo-label accuracy of unlabeled traffic data, enhances the detection performance of the IoT intrusion detection model, and makes full use of unlabeled data.
Smart Images

Figure CN119995953A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intrusion detection of Internet of Things, and specifically relates to an intrusion detection method of Internet of Things integrating label propagation and fuzzy label allocation. The method proposes two pseudo-label allocation mechanisms for unlabeled traffic data, one based on the label propagation idea, and the other based on fuzziness calculation. Afterwards, the pseudo-label allocation results are integrated based on ensemble learning technology to complete the training of a semi-supervised intrusion detection model, so that the unlabeled data can be fully utilized and the accuracy of the intrusion detection model of the Internet of Things can be improved. Background Art
[0002] With the rapid development of Internet of Things (IoT) technology and its wide application in various fields, the number and types of IoT devices have shown a significant growth trend in recent years. However, since IoT devices are usually small in size, low in manufacturing cost, rely on battery power and are difficult to maintain, their design process often does not fully consider security, resulting in many security vulnerabilities in the devices. These vulnerabilities make IoT devices easy targets for cyber attacks, and attackers often use zero-day vulnerabilities or distributed denial of service (DDoS) attacks to carry out damage. Most of the IoT devices on the Internet are caused by DDoS and botnet attacks launched by malware such as Mirai and Gafgyt. Therefore, efficiently detecting potential intrusions in the IoT environment and responding to possible security threats in a timely manner is one of the core challenges to ensure the security of IoT systems and promote their healthy development.
[0003] With the development of machine learning and deep learning technologies, abnormal behavior detection methods based on deep learning have achieved good results in learning complex attack patterns and zero-day attack detection. However, the model training of the above two traditional detection methods mainly uses a single-source data set, that is, training based on traffic data collection in a single period, but most of the collected data is unlabeled data. Many organizations may only have unlabeled data because labeled data may leak user privacy and the workload is large. If the detection model lacks sufficient information, it often causes problems such as model overfitting and poor generalization, which ultimately significantly affects the detection performance of the intrusion detection model. In the current IoT intrusion detection method based on semi-supervised learning, pseudo-labels are assigned to unlabeled traffic data, and some inaccurate pseudo-labels still exist. The accuracy of pseudo-labels has an important impact on the detection results. Inaccurate pseudo-labels may introduce noise, resulting in a decrease in model detection performance. Summary of the invention
[0004] Aiming at the problem of insufficient labeled traffic data and low accuracy of pseudo-labels for unlabeled traffic data in IoT intrusion detection, the present invention proposes an IoT intrusion detection method that integrates label propagation and fuzzy label allocation. The client uses the similarity between labeled traffic data and unlabeled traffic data to assign pseudo-labels to unlabeled traffic data. In addition, an ensemble learning method is used to jointly determine labels through multiple pseudo-label allocation methods, further improving the accuracy of pseudo-labels and thus improving the model detection performance.
[0005] The present invention proposes an IoT intrusion detection method that integrates label propagation and fuzzy label allocation. Figure 1 The method consists of three stages:
[0006] 1. Pseudo-label generation stage: In order to make full use of the information of unlabeled traffic data and use both labeled and unlabeled traffic data to train the intrusion detection model, two methods, label propagation and fuzzy label assignment, are used here to assign pseudo labels to the unlabeled traffic data.
[0007] (1) Label propagation method
[0008] The client has labeled and unlabeled traffic data Tr and Ts, initializes the feature extraction model E, and extracts the feature representation R of the labeled and unlabeled traffic data UO and R UN Splicing to get R U =[R UO ; R UN ], for the feature representation of each flow data sample, the Gaussian kernel function is used to measure the distance between flow data:
[0009] W ij =exp(-||R U (i,:)-R U (j,:)|| 2 / σ 2 )(i≠j)#(1),
[0010] Where R U (i,:) is the matrix R U The element of the i-th row, R U (j,:) is the matrix R UO The element in the jth row. Select the k traffic data samples closest to it as its nearest neighbor according to the distance sorting, and fill their distances into the K nearest neighbor matrix, and the rest of the matrix elements are zero. The symmetric normalized form of the matrix W is
[0011]
[0012] Where D is the degree matrix, and its diagonal elements D ii is the sum of the elements in the i-th row of matrix W.
[0013] The client defines a label matrix Y, where the rows corresponding to the labeled traffic data are the corresponding one-hot encoded vectors, and the rows of the remaining unlabeled traffic data are zero. The assignment of pseudo labels to unlabeled traffic data is achieved by calculating the probability matrix
[0014] Z=(I-αS) -1 Y#(3),
[0015] The conjugate gradient method is used to solve the linear equation system (I-αS)Z=Y, where I is the identity matrix and α (α∈(0,1)) is used to adjust the relative importance of the traffic data sample and its adjacent traffic data samples. Finally, the unlabeled traffic data pseudo label can be obtained from the probability matrix Z
[0016]
[0017] in The pseudo labels obtained through this process cannot be guaranteed to be completely accurate. Using incorrect pseudo labels for training will affect the performance of the client intrusion detection model. To address this problem, the entropy value of the traffic data sample is used to select the pseudo labels. to ensure its accuracy.
[0018] (2) Fuzzy label assignment method
[0019] The client has labeled and unlabeled traffic data Tr and Ts, initializes the classifier C, trains the classifier using the labeled traffic data, and uses the classifier C to calculate the membership vector V = {u 1 ,u 2 ,…,u n}, using the calculated ambiguity
[0020]
[0021] Where F(V) is the ambiguity value of the membership vector V. According to the ambiguity value, the samples are divided into low ambiguity groups FG low , medium fuzziness group FG mid and high fuzziness group FG high , and will belong to FG mid and FG high The samples with the highest ambiguity are merged into the dataset Tr for retraining the classifier C, and the steps are repeated until the ambiguity group FG is mid Finally, the classifier is used as an intrusion detection model to give pseudo labels to the unlabeled traffic data Ts.
[0022] 2. Pseudo-label integration stage:
[0023] Select the traffic data that has the same pseudo-label prediction results for the unlabeled traffic data Ts by the above two methods, and use the prediction results as label information to add these unlabeled traffic data to Tr. In this way, iteratively update the feature extraction model E in the label propagation phase based on the K-neighbor matrix and the classifier C in the label propagation phase based on fuzziness until all the unlabeled traffic data are labeled.
[0024] 3. Intrusion detection model training phase:
[0025] For the labeled flow data T r , this part of data is used for supervised learning, for the unlabeled traffic data T S ,This part of data, obtains pseudo labels through the above pseudo label integration, and the client integrates the labeled traffic data and the unlabeled traffic data that successfully assigns pseudo labels T = [T r ; T S ]. The client initializes the intrusion detection model and uses the traffic data T to train the intrusion detection model. Through semi-supervised learning, the model integrates the labeled traffic data and unlabeled traffic data information for training, making full use of large-scale unlabeled traffic data to improve the detection performance. Finally, the trained intrusion detection model is used to accurately identify normal traffic and abnormal traffic.
[0026] Beneficial effects of the present invention:
[0027] 1. Among the existing traditional classical IoT intrusion detection methods based on semi-supervised learning, the label propagation method based on the K nearest neighbor matrix is easily affected by the distribution of traffic data. IoT traffic data usually has a large number of normal traffic and unbalanced sample categories. Therefore, the pseudo-labels assigned by the label propagation method based on the K nearest neighbor matrix have low accuracy. The pseudo-label assignment method based on fuzziness is more dependent on hyperparameters. Inaccurate measurement and grouping of fuzziness will lead to the assignment of incorrect pseudo-labels and introduce noise.
[0028] 2. The accuracy of pseudo labels has an important impact on the detection results. Inaccurate pseudo labels may introduce noise, resulting in a decrease in the model detection performance. Based on the idea of ensemble learning, we use two pseudo label assignment methods to jointly determine the labels and iteratively update the classifier and feature extraction model, which can further improve the accuracy of pseudo label assignment and make full use of unlabeled data information to improve the performance of intrusion detection models. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a flow chart of the label propagation method based on K nearest neighbor matrix;
[0030] Figure 2 Interactive flowchart for the fuzziness-based pseudo-label assignment method;
[0031] Figure 3 Schematic diagram of label propagation process for IoT intrusion detection. DETAILED DESCRIPTION
[0032] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0033] The specific implementation process of the IoT intrusion detection method integrating label propagation and fuzzy label allocation of the present invention is as follows: Figure 1 As shown, the following steps are included:
[0034] (1) Step 1: Pseudo label generation;
[0035] 1) Label propagation method
[0036] a) The client initializes the feature extraction model E, which has L hidden nodes and the activation function is relu; the feature extraction model E is used to extract the feature representation R of the labeled and unlabeled traffic data UO and R UN , concatenate the two to get R U =[R UO ; R UN ], used to construct the client's K nearest neighbor matrix W;
[0037] b) For each traffic data sample, use the Gaussian kernel function W ij =exp(-||R U (i,:)-R U (j,:)|| 2 / σ 2 )(i≠j) measures the distance between flow data, where R U (i,:) is the matrix R U The element of the i-th row, R U (j,:) is the matrix R UO The elements of row j. The similarity between samples is measured by distance;
[0038] c) Construct a K nearest neighbor matrix, select the k traffic data samples closest to it as its nearest neighbors according to the distance sorting, and fill their distances into the K nearest neighbor matrix, and the rest of the matrix elements are zero. Symmetrically normalize the matrix W Where D is the degree matrix, and its diagonal elements D ii is the sum of the elements in the i-th row of matrix W.
[0039] d) The client defines a label matrix Y, where the rows corresponding to the labeled traffic data are the corresponding one-hot encoding vectors, and the rows of the remaining unlabeled traffic data are zero. The assignment of pseudo labels to unlabeled traffic data is done by calculating the probability matrix Z = (I-αS) -1Y is implemented, where the conjugate gradient method is used to solve the linear equation system (I-αS)Z=Y, where I is the identity matrix and α (α∈(0,1)) is used to adjust the relative importance of the flow data sample and its adjacent flow data samples.
[0040] e) Finally, the probability matrix Z can be used to obtain the pseudo label of the unlabeled traffic data:
[0041] The pseudo labels obtained through this process cannot be guaranteed to be completely accurate. Using incorrect pseudo labels for training will affect the performance of the client intrusion detection model. To address this problem, the entropy value of the traffic data sample is used to measure uncertainty and pseudo labels with high certainty are selected. to ensure its accuracy.
[0042] 2) Fuzzy label assignment method;
[0043] a) The client initializes the classifier C. The hidden layer of the classifier has L hidden nodes, the activation function is sigmoid, and the hyperparameter low fuzziness threshold F is set l and high blur threshold F h , train the classifier using labeled traffic data;
[0044] b) Obtain the membership vector V of the unlabeled traffic data through the classifier C = {u 1 ,u 2 ,…,u n}, using the membership vector and The equation calculates the ambiguity F(V);
[0045] c) According to the low blur threshold F l and high blur threshold F h , the unlabeled traffic data is divided into low ambiguity groups FG low , medium fuzziness group FG mid and high fuzziness group FG high ;
[0046] d) Extract the low ambiguity group FG low and high fuzziness group FG high The unlabeled traffic data is merged into the labeled dataset and updated. The classifier C is trained using the updated dataset until the ambiguity group FG mid is empty;
[0047] e) Use the trained classifier C as the intrusion detection model to give pseudo labels to the unlabeled traffic data Ts
[0048] (2) Step 2: Pseudo-label integration stage;
[0049] 1) Assume that the pseudo-label prediction results of the unlabeled traffic data Ts through the above two stages are and If they are consistent, the pseudo-label of the traffic data is used, and the corresponding unlabeled traffic data is added to the labeled traffic data set;
[0050] 2) Iteratively update the feature extraction model E in the label propagation phase based on the K-neighbor matrix and the classifier C in the label propagation phase based on fuzziness until all unlabeled traffic data are labeled;
[0051] (3) Step 3: Intrusion detection model training phase;
[0052] 1) The client initializes the intrusion detection model.
[0053] 2) Each client integrates unlabeled traffic data T r and labeled traffic data T S , using the flow data T = [T r ; T S ]Training intrusion detection models.
[0054] 3) Until the intrusion detection model converges. After the training is completed, the client uses the final trained intrusion detection model to accurately identify normal traffic and abnormal traffic.
Claims
1. An IoT intrusion detection method integrating label propagation and fuzzy label assignment, characterized in that: The method comprises three stages: pseudo label generation stage, pseudo label integration stage and intrusion detection model training stage; the client uses the similarity between the labeled traffic data and the unlabeled traffic data and adopts two pseudo label assignment methods to assign pseudo labels to the unlabeled traffic data; The results of two unlabeled traffic data pseudo-label assignment methods are integrated using ensemble learning to improve the accuracy of unlabeled traffic data pseudo-labels.
2. The IoT intrusion detection method integrating label propagation and fuzzy label allocation according to claim 1 is characterized in that: The following steps are involved: Step 1: Pseudo label generation; 1) Label propagation method; a) The client initializes the feature extraction model E, which has L hidden nodes and the activation function is relu; the feature extraction model E is used to extract the feature representation R of the labeled and unlabeled traffic data UO and R UN , concatenate the two to get R U =[R UO ; R UN ], used to construct the client's K nearest neighbor matrix W; b) For each traffic data sample, use the Gaussian kernel function W ij =exp(-||R U (i,:)-R U (j,:)|| 2 / σ 2 )(i≠j) measures the distance between flow data, where R U (i,:) is the matrix R U The element of the i-th row, R U (j,:) is the matrix R UO The elements of the jth row; the similarity between samples is measured by distance; c) Construct a K nearest neighbor matrix, select the k traffic data samples closest to it as its nearest neighbors according to the distance sorting, and fill their distances into the K nearest neighbor matrix, and the rest of the matrix elements are zero; perform symmetric normalization on the matrix W Where D is the degree matrix, and its diagonal elements D ii is the sum of the elements in the i-th row of matrix W; d) The client defines a label matrix Y, where the rows corresponding to the labeled traffic data are the corresponding one-hot encoded vectors, and the rows of the remaining unlabeled traffic data are zero; The assignment of pseudo labels to unlabeled traffic data is done by calculating the probability matrix Z = (I-αS) -1 Y is implemented, where the conjugate gradient method is used to solve the linear equation system (I-αS)Z=Y, where is the identity matrix, and α (α∈(0,1)) is used to adjust the relative importance of the flow data sample and its adjacent flow data samples; e) Finally, the unlabeled traffic data pseudo-label is obtained from the probability matrix Z: Use the entropy value of the traffic data sample to measure uncertainty and select pseudo labels with high certainty To ensure its accuracy; 2) Fuzzy label assignment method; a) The client initializes the classifier C. The hidden layer of the classifier has L hidden nodes, the activation function is sigmoid, and the hyperparameter low fuzziness threshold F is set l and high blur threshold F h , train the classifier using labeled traffic data; b) Obtain the membership vector V = {u1, u2, ..., u n }, using the membership vector and The equation calculates the ambiguity F(V); c) According to the low blur threshold F l and high blur threshold F h , the unlabeled traffic data is divided into low ambiguity groups FG low , medium fuzziness group FG mid and high fuzziness group FG high ; d) Extract the low ambiguity group FG low and high fuzziness group FG high The unlabeled traffic data is merged into the labeled dataset and updated; the classifier C is trained using the updated dataset until the medium ambiguity group FG mid is empty; e) Use the trained classifier C as the intrusion detection model to give pseudo labels to the unlabeled traffic data Ts Step 2: Pseudo-label integration stage; 1) Pseudo-label prediction results for unlabeled traffic data Ts and If they are consistent, the pseudo-label of the traffic data is used, and the corresponding unlabeled traffic data is added to the labeled traffic data dataset; 2) Iteratively update the feature extraction model E in the label propagation phase based on the K-neighbor matrix and the classifier C in the label propagation phase based on fuzziness until all unlabeled traffic data are labeled; Step 3: Intrusion detection model training phase; 1) The client initializes the intrusion detection model; 2) Each client integrates unlabeled traffic data T r and labeled traffic data T S , using the flow data f = [T r ; T S ]Training intrusion detection models; 3) Until the intrusion detection model converges; after the training is completed, the client uses the final trained intrusion detection model to accurately identify normal traffic and abnormal traffic.