Abnormal sample detection method for Internet of Things traffic
Through class mean clustering and loss function separation between groups within the group, combined with KDE adaptive threshold method, the abnormal sample detection problem of IoT traffic under the cloud edge architecture is solved, and efficient and robust OOD detection is achieved.
Patent Information
- Application Number
- CN202510553575.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-25
AI Technical Summary
Under the cloud-edge and end collaboration architecture, IoT systems face difficulties in detecting unknown or new traffic samples (OOD). Traditional methods perform poorly in IoT scenarios with unbalanced categories, complex data and fast changes, making it difficult to accurately identify abnormal samples, threatening the security and stability of the system.
The class-mean-value clustering algorithm is used to form feature groups, and the classifier model is optimized based on the compact and separated distance loss function within the group, and the sample deviation degree is quantified using Mahalanobis distance, and the abnormal samples are dynamically determined in combination with the KDE adaptive threshold method.
It significantly improves the accuracy and robustness of off-distribution sample detection, with AUROC increased by at least 33%, and FPR95 reduced by at least 10%, which is suitable for real deployment scenarios of complex IoT traffic.
Smart Images

Figure CN120378336A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network security technology, and particularly relates to a method for detecting abnormal samples in Internet of Things (IoT) traffic, which is applicable to an abnormal traffic detection system in an edge-cloud collaborative environment. Background Art
[0002] With the rapid development of Internet of Things (IoT) technology, intelligent terminal devices are widely deployed in multiple scenarios such as industrial control, smart city, smart home, and vehicle networking. These devices generate a large amount of network traffic data during operation. Coupled with the diversification of communication protocols and the complexity of data types, it significantly increases the difficulty of network communication and security management. To ensure the normal operation of the system, the industry generally adopts network security models based on supervised learning or deep learning for tasks such as intrusion detection, anomaly recognition, and device classification.
[0003] In practical applications, IoT systems usually adopt a cloud-edge-terminal collaborative computing architecture to balance computing resources, communication costs, and response efficiency. The terminal (end) is responsible for data collection and preliminary processing, the edge device (edge) is responsible for local analysis and rapid response, while the cloud (cloud) undertakes tasks such as unified training, model update, and policy distribution. This cloud-edge-terminal collaborative architecture can improve the real-time performance and security of data processing without increasing latency.
[0004] However, this cloud-edge-terminal collaborative architecture faces a key challenge: the model cannot predict all device types or network behaviors, resulting in frequently encountering unknown or new traffic samples, namely "Out-of-Distribution (OOD)" samples after deployment. These samples may come from new attack behaviors, unregistered device types, or faulty nodes. If not accurately detected, they will seriously threaten the security and stability of the system.
[0005] Traditional out-of-distribution detection methods such as MSP, ODIN, Energy, etc. mainly rely on the output confidence of the classification model or its transformed form to determine whether a sample is OOD. These methods generally assume that the training sample distribution is balanced, the number of categories is small, and the data features are stable, so they perform well on standard data sets such as images. However, in actual IoT scenarios, the distribution characteristics are highly complex and there are the following prominent problems: the class imbalance is serious, the number of categories continues to grow, the data changes rapidly and uncontrollably; and there are time delays in cloud-edge model synchronization, data privacy, and bandwidth limitations, etc.
[0006] Therefore, in the cloud-edge-terminal architecture, designing an efficient OOD detection method with "strong discriminative power, flexible deployment, and the ability to adaptively adjust thresholds" has become an important technical requirement for IoT security protection. Summary of the Invention
[0007] Aiming at the problems of numerous categories of IoT traffic, unbalanced samples, and difficulty in identifying abnormal samples by traditional methods in the cloud-edge-end architecture, the present invention proposes a method for detecting abnormal samples in IoT traffic, realizing the detection and identification of abnormal samples in the multi-category unbalanced scenario.
[0008] The method for detecting abnormal samples in IoT traffic is as follows:
[0009] Step 1: Collect the original IoT traffic data and upload it to the edge computing node or the central server. After feature extraction and the class mean clustering algorithm, form a structured feature grouping.
[0010] Step 101: From the collected IoT traffic packets, use the CICFlowMeter tool to extract general traffic features as training data.
[0011] Step 102: Use the class mean clustering method to group the training data, divide the similar classes into one group, and obtain the group labels corresponding to each category.
[0012] First, initially manually label the class labels for the training data.
[0013] Then, calculate the feature means of the classes corresponding to different labels.
[0014] The feature mean of category i is calculated as follows:
[0015]
[0016] where X train,i represents the set of data samples belonging to the i-th category in the training set.
[0017] Subsequently, perform hierarchical clustering on the mean set {μ i} of all categories, and divide the similar categories into one group.
[0018] The calculation formula is as follows:
[0019] group c ←HierarchicalClustering({μ i})
[0020] HierarchicalClustering represents the hierarchical clustering function, and group c represents the set of group labels corresponding to all categories, and c is the total number of original categories.
[0021] Step 103: Assign the respective within-group category labels to each class within each group.
[0022] Specifically, first initialize the class counter of each group to 0; traverse each class within the group and update the class counter of the group: And set the within-group label of this class to the current count value:
[0023] group i Indicates the group label to which the i-th class belongs; Indicates the group label group i The corresponding class counter, and each group corresponds to a counter class_group; it increases continuously by traversing the number of classes in the group. Indicates the within-group class label of the i-th class.
[0024] Finally, the classes within each group are represented using independent local labels, thus forming an aggregation structure in the feature space with groups as units.
[0025] Step 2: Optimize the classifier model using a distance loss function that is compact within groups and separated between groups for feature grouping.
[0026] The output of the classifier model includes: intermediate features (features of the layer before the output layer), and the final classification result logits. The intermediate features can be used for anomaly detection, and logits are used for classification tasks.
[0027] The distance loss function is:
[0028]
[0029] Where Is the cross-entropy loss:
[0030] G is the number of groups for feature grouping, C g Is the number of classes in the g-th group, Is the probability after softmax of the final output value logits of the classifier model, Is the within-group class label of the i-th class in the g-th group, and N is the total number of training data.
[0031] L in-group Is the within-group loss:
[0032] Where |G g | represents the number of all training data of all classes within the g-th group, and G g Is the set of all training data of all classes within the g-th group; the mean of the g-th group is ||·|| represents the Euclidean norm. Is the j-th intermediate feature vector within the g-th group output by the classifier model;
[0033] is the inter-group constraint loss:
[0034] m g is the training data set that does not belong to the group G g ; is the feature mean of the non-group training data m g ;
[0035]
[0036] Step 3: Deploy the optimized classifier model to the cloud or edge device, input the sample to be tested to obtain the intermediate feature vector, calculate the Mahalanobis distance between the feature means of each feature group output by the classifier model respectively, and select the minimum distance as the anomaly score;
[0037] The calculation formula is:
[0038]
[0039] D M (x, G g ) is the Mahalanobis distance from the sample x to be tested to the group G g ;
[0040] Step 4: Use the KDE adaptive threshold method to perform probability density estimation analysis on the anomaly scores to dynamically determine whether the sample to be tested is an anomaly.
[0041] First, use the kernel density estimation KDE to model the anomaly scores;
[0042] The formula is:
[0043] represents the probability density value corresponding to the anomaly score s, n is the total number of anomaly scores, h is the bandwidth parameter that controls the smoothness of the estimation, K(·): is the Gaussian kernel function, and s k is the k-th anomaly score.
[0044] Then, calculate the corresponding probability density for the anomaly scores of all samples to be tested to obtain the probability density set
[0045] Determine the KDE threshold according to the probability density set; for any sample to be tested, if its anomaly score is greater than the KDE threshold, it is determined to be an abnormal sample; if its anomaly score is less than or equal to the threshold, it is determined to be a normal sample.
[0046] Specifically: First, draw the KDE curve according to the probability density set. The difference in the abnormal score distribution between normal and abnormal samples causes an obvious local minimum in the "transition zone" of the KDE curve for the two types of samples. Take this first local minimum as the demarcation point to determine the adaptive threshold for anomaly detection.
[0047] The threshold calculation formula is:
[0048] γ represents the probability density value corresponding to the first local minimum, that is, the KDE threshold.
[0049] Among them, Minimas is the local minimum index; the formula is:
[0050] Among them, the argrelextrema function returns the positions of local minima in the array; the np.less parameter specifies finding local minima.
[0051] The advantages of the present invention are as follows:
[0052] 1), A method for detecting abnormal samples in Internet of Things traffic, introducing a supervised feature clustering and grouping mechanism simplifies the discrimination boundary; designing a grouping distance loss function improves the feature separability.
[0053] 2), A method for detecting abnormal samples in Internet of Things traffic, improves the discriminability of in-distribution samples through feature clustering and grouping and grouping distance loss optimization, combines kernel density estimation to achieve adaptive threshold selection, significantly improves the detection accuracy and robustness of out-of-distribution samples in complex Internet of Things traffic, the AUROC is increased by at least 33%, the FPR95 is reduced by at least 10%, and it is applicable to real cloud-edge-end deployment scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a flowchart of a method for detecting abnormal samples in Internet of Things traffic according to the present invention;
[0055] Figure 2 It is a flowchart of using the class mean clustering method to group the training data to obtain group labels and within-group class labels according to the present invention;
[0056] Figure 3 It is a schematic diagram of the constraint of the classifier model according to the present invention;
[0057] Figure 4 It is a flowchart of using the KDE adaptive threshold method to determine whether a sample to be tested is abnormal according to the present invention;
[0058] Figure 5 It is a comparison diagram of the present invention and the traditional out-of-distribution detection method based on a classifier. DETAILED DESCRIPTION OF THE INVENTION
[0059] In order to facilitate those skilled in the art to understand and implement the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. Obviously, the described embodiments are only partial embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of the present invention.
[0060] The present invention proposes a method for detecting abnormal samples in IoT traffic, and provides accurate and efficient out-of-distribution sample identification capabilities for complex, heterogeneous, and unbalanced network traffic data generated by IoT devices under a cloud-edge-end architecture. The method uploads the collected original IoT traffic data to an edge computing node or a central server, and forms a structured feature grouping representation through feature extraction and class mean clustering algorithm modeling training, and optimizes the classifier model with the help of a specially designed distance loss function with compactness within the group and separation between groups. After the classifier model is deployed, the sample to be tested is measured by the Mahalanobis distance to measure its degree of deviation from the group center, so as to determine whether it is an out-of-distribution sample. Finally, the KDE adaptive threshold method is used to perform probability density estimation analysis on the detection score, and the optimal discrimination threshold is dynamically determined without manual intervention, which is suitable for adaptive deployment requirements in large-scale heterogeneous devices and dynamic scenarios.
[0061] like Figure 1 As shown, the specific steps are as follows:
[0062] Step 1: Collect the original IoT traffic data and upload it to the edge computing node or central server. After feature extraction and class mean clustering algorithm, a structured feature grouping is formed.
[0063] Step 101: Use the CICFlowMeter tool to extract common traffic features from the collected IOT traffic packets as training data;
[0064] Common traffic features include total number of forward packets, total number of backward packets, total length of forward packets, total length of backward packets, flow duration, protocol type, etc.
[0065] Step 102: group the training data using the class mean clustering method, divide similar classes into one group, and obtain group labels corresponding to each class;
[0066] like Figure 2 As shown, first, the training data is manually labeled;
[0067] The original classification of the training data was artificially collected by constructing corresponding experimental scenarios. Through manual routine operations, IOT traffic packets under the normal behavior of the collection device were collected, and normal traffic packets of different devices were obtained for tagging and classification.
[0068] Then, calculate the feature means of the classes corresponding to different tags;
[0069] The feature mean of class i is calculated as follows:
[0070]
[0071] where X train,i represents the set of data samples belonging to the i-th class in the training set.
[0072] Subsequently, perform hierarchical clustering on the set of means {μ i} for all classes, and divide similar classes into one group;
[0073] The calculation formula is as follows:
[0074] group c ←HierarchicalClustering({μ i})
[0075] HierarchicalClustering represents the hierarchical clustering function, which is a clustering method that gradually aggregates to form a tree structure based on the similarity between samples. group c represents the set of group labels corresponding to all classes, and c is the total number of original classes.
[0076] Step 103: Assign respective within-group class labels to each class within each group.
[0077] Specifically, first initialize the class counter for each group to 0; traverse each class within the group and update the class counter for that group: And set the within-group label of this class to the current count value:
[0078] group i represents the group label to which the i-th class belongs; represents the class counter corresponding to the group label group i Each group corresponds to a counter class_group; it increases continuously by traversing the number of classes in the group. represents the within-group class label of the i-th class.
[0079] Ultimately, the entire category set is divided into several groups, each of which consists of multiple categories; categories within each group are represented using independent local labels, thereby forming a clustering structure based on small groups in the feature space, providing simplified support for subsequent model discrimination and decision-making.
[0080] Step 2: Optimize the classifier model using the distance loss function that optimizes the intra-group compactness and inter-group separation of feature grouping.
[0081] Unlike conventional classifiers that only output classification results (logits), this model outputs two results at the same time: one is the intermediate feature (the feature of the layer before the output layer), and the other is the final classification logits. The intermediate feature can be used for anomaly detection, and logits are used for classification tasks.
[0082] The classifier model constraints described in the present invention are divided into two categories: intra-group constraints and inter-group constraints. The specific process is as follows: Figure 3 As shown:
[0083] First, based on the clustering algorithm of class mean, group labels and intra-group class labels are constructed for all classes. The feature vector f of the penultimate layer of the classifier model, i.e., the layer before the output layer, is grouped according to the predefined grouping interval [group slice [i][0],group slice [i][1]] is intercepted to obtain the sub-features corresponding to each group, which are called group features. By imposing constraints on these group feature spaces, the ability to distinguish between groups can be enhanced, making the model decision more accurate and robust.
[0084] At the same time, in order to cluster samples in the same group as much as possible in the feature space, thereby improving the consistency within the group, the following intra-group loss L is introduced in-group :
[0085]
[0086] Where |G g | represents the number of all training data of all classes in the g-th group, G g is the set of all training data of all classes in the g-th group; the mean of the g-th group is ||·|| represents the Euclidean norm. is the jth intermediate feature vector in the gth group output by the classifier model;
[0087] This constraint aims to reduce the distance between each sample and the group mean of the group to which it belongs, so that the samples within the group are tightly aggregated and the characteristics of the same group are similar and compact.
[0088] The goal of the inter-group constraint is to enhance the distinction between different groups. For each group Gg , construct a set m of training data that does not belong to this group in the current batch g , and calculate the feature means of these non-group samples
[0089]
[0090] Then, use the Euclidean distance between the center of this group and as the inter-group separation degree, and construct an inter-group constraint loss:
[0091]
[0092] This loss term continuously increases the minimum value in the inter-group distance, effectively preventing samples from different groups from being too close, thus widening the boundaries between groups and enhancing the model's discriminative ability for abnormal samples.
[0093] Finally, obtain the distance loss function as:
[0094]
[0095] where is the cross-entropy loss to ensure the normal progress of the classification task; its calculation formula is:
[0096]
[0097] G is the number of groups of feature grouping, C g is the number of categories in the g-th group, is the probability after softmax of the final output value logits of the classifier model, is the intra-group category label of the i-th category in the g-th group, and N is the total number of training data.
[0098] Actually, although the original features are divided based on the class mean clustering method, the boundaries are still blurred. Through the grouping distance loss constraint, the intra-group distance is reduced and the inter-group distance is enlarged, making the groups show a cluster structure in the model-optimized feature space, that is, the decision boundaries between groups are clearer. When abnormal samples appear in the feature space, due to feature differences, they will fall outside the decision boundary, so the abnormal detection ability is improved, and the model generates a feature for each sample.
[0099] Step 3: Deploy the optimized classifier model to the cloud or edge device, input the sample to be tested to obtain the intermediate feature vector, calculate the Mahalanobis distance with the feature means of each feature grouping output by the classifier model respectively, and select the minimum distance as the anomaly score;
[0100] The classifier model can handle scenarios with a large number of class labels and a high degree of imbalance between classes. In the training process of this model, a distance loss function based on group division is used for optimization. Therefore, when detecting a sample to be tested, the Mahalanobis distance is used to measure the deviation degree of the sample from each group as the anomaly score. The calculation formula is as follows:
[0101]
[0102] D M (x, G g ) is the Mahalanobis distance from the sample x to be tested to the group G g ;
[0103] The calculation formula is represents the mean of the feature vectors extracted after the training data of the g-th group is trained by the classifier model, represents the inverse covariance matrix of the features of the g-th group, and x is the feature vector extracted from the sample to be tested in the classifier model, is the intermediate feature vector obtained by the sample to be tested through the classifier model.
[0104] When the sample to be tested is a normal sample and is close to the center of a certain group, its Mahalanobis distance is small and the anomaly score is low; on the contrary, if the sample deviates from the centers of all groups, the distance is large, the anomaly score is high, and it is more likely to be judged as an abnormal sample. Therefore, the minimum value of the distances to each group is selected as the anomaly score. Since the anomaly score is only a value measuring its anomaly degree, but it cannot directly determine whether the sample to be tested is abnormal or normal, a threshold needs to be selected.
[0105] Step 4: Use the KDE adaptive threshold method to perform probability density estimation analysis on the anomaly score to dynamically determine whether the sample to be tested is abnormal.
[0106] As Figure 4 shown, first, use Kernel Density Estimation (KDE) to model the anomaly score;
[0107] The formula is:
[0108] represents the probability density value corresponding to the anomaly score s, n is the total number of anomaly scores, h is the bandwidth parameter controlling the smoothness of the estimation, K(·): is the Gaussian kernel function, and s k is the k-th anomaly score.
[0109] Then, calculate the corresponding probability density for the anomaly scores of all samples to be tested to obtain the probability density set
[0110] Determine the KDE threshold according to the probability density set; for any sample to be tested, if its anomaly score is greater than the KDE threshold, it is determined as an abnormal sample; if its anomaly score is less than or equal to the threshold, it is determined as a normal sample.
[0111] Specifically: First, draw the KDE curve according to the probability density set; since normal samples are closer to the center of their respective groups, their Mahalanobis anomaly scores are usually smaller and concentrated, so they show a high-density region similar to a normal distribution in the KDE curve. Abnormal samples, on the other hand, are far from the centers of all groups, resulting in larger anomaly scores, sparse distributions, low density values, and being concentrated at the far end of the KDE curve.
[0112] The difference in the distribution of anomaly scores between normal and abnormal samples causes an obvious local minimum in the "transition region" of the KDE curve; taking this first local minimum as the demarcation point, an adaptive threshold for anomaly detection is determined;
[0113] The formula for determining the local minimum index is
[0114] where Minimas is the local minimum index; the argrelextrema function returns the positions of local extrema in the array; the np.less parameter specifies to find local minima.
[0115] The probability density value corresponding to the first local minimum, that is, the formula for calculating the KDE threshold is:
[0116] γ represents the probability density value corresponding to the first local minimum, that is, the KDE threshold.
[0117] The anomaly detection process is that for any sample to be tested, if its anomaly score is greater than the KDE threshold, it is determined as an abnormal sample; if its anomaly score is less than or equal to the threshold, it is determined as a normal sample. This method is essentially an unsupervised adaptive threshold selection mechanism, avoiding manual setting or strong dependence on the validation set, and is more suitable for anomaly detection tasks in an open environment.
[0118] Based on the feature clustering grouping distance loss function and the KDE adaptive threshold, the present invention proposes a multi-level out-of-distribution sample detection system framework, which mainly includes the following three modules: a category clustering grouping module, a grouping distance loss optimization module, and an adaptive threshold selection module.
[0119] Specifically as follows:
[0120] Category Clustering and Grouping Module: Based on the category labels in the training set, calculate the class mean vectors of each category in the feature space, and perform clustering analysis on all categories. By dividing similar categories in the feature space into the same group, local aggregation and overall differentiation of multiple categories in the high-dimensional space are achieved. This clustering and grouping strategy effectively reduces the complexity of the classification boundary, improves the interpretability and stability of feature representation, and provides a structural basis for subsequent distance optimization. This module can be implemented by methods such as hierarchical clustering, K-means, or graph partitioning. The number of clusters can be flexibly set according to the task scenario and has scalability.
[0121] The principle of the category clustering and grouping module is as follows: Since the number of categories is large, the decision boundary of the model becomes complex. Therefore, feature clustering and grouping are used to alleviate the risk of overfitting and increase the generalization ability. Different from traditional unsupervised clustering, the present invention calculates the category feature means, clusters the category means, and realizes supervised clustering of the entire training data, and re-divides the feature space of the training data. According to the clustering results, set the group label of the sample according to the group to which each sample belongs, and replace the original category label with the within-group category label.
[0122] Group Distance Loss Optimization Module: Based on the clustering results, introduce a distance loss function that is compact within groups and separated between groups, which is used to constrain the spatial layout of the neural network during the feature extraction stage. Specifically, the system divides the group features according to the grouping at the penultimate layer of the network (i.e., the feature layer before the classification layer) and applies this loss, so that the feature of the same-group samples converges to the corresponding group center, and at the same time significantly increases the minimum distance between different group centers. This process does not depend on the classification confidence of the model output, so it is more robust than traditional confidence-based out-of-distribution detection methods. During the training process, the loss function can be jointly optimized with classification, so that the features focus on each group, making the feature differentiation between groups large and the boundaries clear, and improving the discrimination ability of the model under complex, multi-class, and unbalanced traffic.
[0123] The principle of the group distance loss optimization module is as follows: After the network outputs features at the penultimate layer, based on the pre-generated category clustering and grouping labels (obtained by hierarchical clustering), the feature space is divided into multiple groups, and feature focusing is achieved through dual loss constraints:
[0124] Among them, the within-group compactness constraint calculates the distance loss between the feature of the within-group sample and the group center, forcing the within-group features to aggregate. By continuously reducing the distance between the within-group sample points and the within-group center, the samples of one group are made as close as possible.
[0125] Inter-group separability constraint, maximizing the minimum distance between different group centers, enhancing the inter-group boundary. By continuously expanding the minimum distance from out-of-group features to the group center, the distance between groups is widened, and finally there is an obvious boundary between groups. Abnormal samples, unable to match any intra-group distribution, will fall into the inter-group boundary area, thus achieving highly robust anomaly recognition.
[0126] Adaptive threshold selection module: Aiming at the problem of threshold sensitivity in out-of-distribution sample detection in actual deployment, a distribution adaptive threshold selection strategy based on kernel density estimation (KDE) is proposed. During the inference phase of the system, based on the Mahalanobis distance, the deviation degree of each test sample from the nearest group center is calculated, and the distance scores of all samples are collected. The KDE is used to smoothly model its distribution, and the inflection point in the probability density curve is identified as the discrimination threshold, thereby automatically separating in-distribution and out-of-distribution samples. This method does not require manual parameter adjustment, can dynamically adapt to different data distributions, is suitable for application scenarios with diverse device types and continuously changing data in the IoT environment, and significantly improves the practicality and deployment efficiency of the detection system.
[0127] The principle of the adaptive threshold selection module is as follows: The adaptive threshold selection module based on dynamic probability density analysis is used to solve the threshold sensitivity problem in out-of-distribution sample detection. This module first calculates the minimum Mahalanobis distance from the test sample to the centers of each group as the anomaly score to quantify its deviation degree from the known distribution. The greater the distance, the more significant the deviation from the known distribution. Then, the KDE is used to perform non-parametric probability density modeling on the anomaly scores of all samples, and the optimal bandwidth is automatically determined through the Scott criterion; the first inflection point is used as the automatically determined threshold boundary. This method achieves adaptability and can dynamically respond to changes in data distribution, especially suitable for application scenarios with diverse IoT device types.
[0128] Figure 5Shows the comparison between this framework and traditional classifier-based out-of-distribution detection methods. The abscissa represents different classifier-based out-of-distribution detection methods, and the ordinate represents FPR95 (i.e., the proportion of misclassifying normal samples (ID) as OOD on the premise of correctly detecting 95% of OOD samples). As can be seen from the figure, the FPR95 of this framework is the lowest, showing the best performance. It shows that in the scenario with a large number of labels and class imbalance, the performance of traditional methods is poor and it is difficult to adapt to the cloud-edge-end environment. While this method optimizes the classifier model through the class mean clustering algorithm and the classification distance loss function, effectively alleviating the pressure of a large number of labels. At the same time, it makes the intra-group features more compact and the inter-group features more distinguishable, thus improving the model's ability to distinguish out-of-distribution samples from normal samples. At the same time, the Mahalanobis distance quantifies the feature distance as the anomaly score, and the KDE adaptive threshold method is used to dynamically determine the optimal discrimination threshold, finally achieving high-performance and high-robustness anomaly detection.
Claims
1. An abnormal sample detection method for Internet of Things traffic, characterized in that, The specific steps are as follows: Step 1: Collect the original IoT traffic data and upload it to the edge computing node or the central server. After feature extraction and the class mean clustering algorithm, form a structured feature grouping; Step 2: Optimize the classifier model using the distance loss function with in-group compactness and between-group separation of the feature grouping; The output of the classifier model includes: intermediate features and the final classification result logits. The intermediate features are used for anomaly detection, and the logits are used for classification tasks; The distance loss function is: wherein is the cross-entropy loss; L in-group is the within-group loss; is the between-group constraint loss; Step 3: Deploy the optimized classifier model to the cloud or edge device. Input the sample to be tested to obtain the intermediate feature vector, calculate the Mahalanobis distance between the intermediate feature vector and the feature means of each feature grouping output by the classifier model respectively, and select the minimum distance as the anomaly score; Step 4: Use the KDE adaptive threshold method to perform probability density estimation analysis on the anomaly score, and dynamically determine whether the sample to be tested is an anomaly; First, use the kernel density estimation KDE to model the anomaly score; The formula is: represents the probability density value corresponding to the anomaly score s, n is the total number of anomaly scores, h is the bandwidth parameter that controls the smoothness of the estimation, K(·) is the Gaussian kernel function, and s k is the k-th anomaly score; Then, calculate the corresponding probability density for the anomaly scores of all samples to be tested, and obtain a set of probability densities Determine the KDE threshold according to the probability density set; for any sample to be tested, if its anomaly score is greater than the KDE threshold, it is determined as an abnormal sample; if its anomaly score is less than or equal to the threshold, it is determined as a normal sample.
2. The anomaly sample detection method for Internet of Things traffic according to claim 1, wherein The specific content of Step 1 is as follows: Step 101: From the collected IoT traffic packets, use the CICFlowMeter tool to extract general traffic features as training data; Step 102: Use the class mean clustering method to group the training data, divide the similar classes into one group, and obtain the group labels corresponding to each class; First, manually label the class labels for the training data initially; Then, calculate the feature means of the classes corresponding to different labels; The calculation formula for the feature mean of class i is as follows: Among them, X train,i represents the set of data samples belonging to the i-th category in the training set; Subsequently, hierarchical clustering is performed on the set of means {μ i} for all categories, and similar categories are grouped together; The calculation formula is as follows: group c ←HierarchicalClustering({μ i}) HierarchicalClustering represents the hierarchical clustering function, and group c represents the set of group labels corresponding to all classes, and c is the total number of original classes; Step 103: Assign the respective in-group class labels to each class within each group; Specifically, first initialize the category counter of each group to 0; traverse each category within the group and update the category counter of this group: And set the within-group label of this category to the current count value: group i Indicates the group label to which the i-th category belongs; Indicates the group label group i The corresponding category counter, each group corresponds to a counter class_group; it is continuously incremented by traversing the number of categories in the group; Indicates the in-group category label of the i-th category; Finally, each in-group class is represented by an independent local label, thus forming an aggregation structure in the feature space with groups as units.
3. The anomaly sample detection method for Internet of Things traffic according to claim 1, wherein The cross-entropy loss calculation in the second step is as follows: G is the number of groups for feature grouping, C g is the number of classes in the g-th group, is the probability after softmax of the final output value logits of the classifier model, is the within-group class label of the i-th class in the g-th group, and N is the total number of training data; The within-group loss is calculated as: where |G g | represents the number of all training data of all classes within the g-th group, and G g is the set of all training data of all classes within the g-th group; the mean of the g-th group is ||·|| represents the Euclidean norm; is the j-th intermediate feature vector within the g-th group output by the classifier model; The inter-group constraint loss is calculated as follows: m g is a training data set that does not belong to the group G g ; is the feature mean of the non-group training data m g .
4. The anomaly sample detection method for Internet of Things traffic according to claim 1, characterized in that The calculation formula for the anomaly score in Step 3 is: D M (x, G g ) is the Mahalanobis distance from the sample x to be measured to the group G g ; the calculation formula is: denotes the mean of the feature vectors extracted after the training data of the g-th group is trained by the classifier model, denotes the inverse covariance matrix of the features of the g-th group, and x is the feature vector extracted from the sample to be tested in the classifier model, is the intermediate feature vector obtained from the sample to be tested through the classifier model.
5. The anomaly sample detection method for Internet of Things traffic according to claim 1, wherein, In Step 4, determining the KDE threshold according to the probability density set is specifically: First, draw the KDE curve according to the probability density set. The difference in the anomaly score distribution between normal and abnormal samples causes an obvious local minimum in the "transition zone" of the KDE curve for the two types of samples; use this first local minimum as the demarcation point to determine the adaptive threshold for anomaly detection; The threshold calculation formula is as follows: γ represents the probability density value corresponding to the first local minimum, that is, the KDE threshold; Among them, Minimas is the local minimum index; the formula is: where the argrelextrema function returns the positions of the local minima in the array; the np.less parameter specifies to find the local minimum.