A method for concept drift detection and adaptation and IoT security system

By using histogram representation and Self-adaption EMD similarity measurement methods in IoT systems, the performance degradation caused by concept drift in IoT systems is solved, and a higher IoT secure traffic recognition accuracy is achieved.

CN115412337BActive Publication Date: 2025-05-06JIANGSU POLICE INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211033769.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-26
Publication Date
2025-05-06
Estimated Expiration
2042-08-26

AI Technical Summary

Technical Problem

The conceptual drift of data exists in IoT systems, resulting in a degradation of system performance and it is difficult for the existing technology to effectively process and improve performance, especially in a label-free actual environment.

Method used

A method for concept drift detection and adaptation is proposed. Through data preparation, training model candidates, model selection and other steps, the similarity between data sets is characterized by histograms, and the similarity measurement method of Self-adaption EMD is used to select the best decision model from the candidate models.

Benefits of technology

Effectively detect and handle concept drifts, improve system performance, and improve the accuracy of IoT secure traffic recognition, suitable for real-time environments without labels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115412337B_ABST
    Figure CN115412337B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for concept drift detection and adaptation and an IoT security system, which relates to the field of IoT technology, including four steps: data preparation, training model candidates, model selection, and prediction output. The present invention aims at security issues such as IoT attacks and malicious traffic identification, and proposes a method for concept drift detection and adaptation and a corresponding IoT security system. By verifying the effectiveness of IoT traffic on data drift detection, the performance of the AI ​​model is evaluated, so as to select the best AI model. Through comparative experiments, the feasibility of the framework and adaptive method in practice is verified, as well as its role in improving the performance of IoT security traffic identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of Internet of Things, and in particular relates to a method for concept drift detection and adaptation and an IoT security system. Background Art

[0002] Traditional security issues manifest themselves uniquely in IoT systems, and the inherent characteristics of IoT systems also introduce new types of threats. Intrusion detection based on traditional labeled classification machine learning is a "closed-world" problem. However, real-world attacks are unlabeled, "open-world" problems, attracting increasing attention from scholars. On the one hand, the wide variety of IoT devices, with their potential for diverse security vulnerabilities and the emergence of unknown attack methods and virus samples, makes identifying malicious IoT traffic a typical "open-world" problem. On the other hand, the large number of IoT devices, the vast amount of data, and the ever-increasing number of attack methods make data anomaly detection particularly important in IoT security. Concept drift in IoT security is highly susceptible to data drift, which can significantly harm models and systems.

[0003] Concept drift describes the unpredictable changes in the distribution of data streams over time. However, the emergence of concept drift inevitably leads to a decline in system performance. There are few research results on how to deal with and ultimately improve system performance. Existing technologies have also proposed some architectures for dealing with concept drift to improve system performance, but generally speaking, they are not applicable to actual unlabeled environments. Summary of the Invention

[0004] The object of the present invention is to provide a method for concept drift detection and adaptation and an IoT security system to address the above-mentioned defects caused by the prior art.

[0005] The present invention discloses a method for concept drift detection and adaptation, comprising the following steps:

[0006] S1. Data Preparation

[0007] Prepare the IoT dataset DataSet for model training, perform replacement sampling on the dataset DataSet, and obtain M sub-training sets DataSet1, DataSet2, ..., DataSet M ;

[0008] S2. Training model candidates

[0009] In the model preparation stage, a classifier is trained as a candidate model on each training set, and a histogram is used to characterize the feature distribution of the training set, while recording the importance of each feature to the model;

[0010] S3. Model selection

[0011] In the model selection stage, a histogram is used to characterize the feature distribution of the data to be predicted. Then, based on the similarity of the histograms, the histogram representations are compared among the candidate models to find the most similar training set, and the model trained on the training set is used as the decision model.

[0012] Preferably, in step S1, a bagging sampling method is adopted, different positive and negative sample ratios are used, and 5% of the original data samples are selected.

[0013] Preferably, the specific method in step S2 is as follows:

[0014] Training baseline model: Use M training sets and N features. On each training set, generate a histogram matrix H, train a classifier E and corresponding weight sequence K, and get a total of M histograms (H1, H2, ... H M ), M classifiers (E1, E2, … E M ) and M weight sequences (K1, K2, ... K M );

[0015] The histogram H of each training set is generated as follows:

[0016] For a certain feature X, suppose the feature value of X is a sequence {x1,x2,…x T}, T represents the length of the eigenvalue sequence, i represents the position in the eigenvalue sequence, and the elements in the eigenvalue sequence are divided into several ranges according to their size, each range is called a bin, h j Indicates the number of times the eigenvalue appears in the jth bin, h j The calculation method is shown in formula (1):

[0017]

[0018] Among them I i is the indicator function, and the calculation method is shown in formula (2):

[0019]

[0020] The histogram of feature X can be calculated by formulas (1) and (2): h = {h1, h2, ..., h C}, using the same method, calculate the histogram for each feature and obtain the histogram matrix h' as ​​shown in formula (3), where C is the number of bins in the histogram and N is the number of features:

[0021]

[0022] Convert the histogram matrix h' into one dimension and obtain the histogram H of the training set as shown in formula (4):

[0023] H={h 1,1 ,h 1,2 ,…,h 1,C ,h 2,1 ,h 2,2 ,…,h 2,C ,…,h N,1 ,h N,2 ,…,h N,C} (4)

[0024] The method for generating the classifier E and weight sequence K on each training set is as follows:

[0025] Select a classifier model and use the samples in the training set to train the model. The classifier is obtained through training. During the training, the importance of each feature to the model is recorded and the sequence {k1, k2, ..., k N}, N is the number of features, and the sequence {k1, k2, …, k N} is the weight sequence K.

[0026] Preferably, the self-adaption EMD similarity measurement method is adopted in step S3. This method first uses the histogram generation method in step 3 to generate a histogram H' for the data to be predicted, and then calculates the similarity between H' and each training set histogram H respectively.

[0027] Preferably, in step S3, let H=P, H'=Q, that is, P and Q represent the histogram of the training sample and the data to be predicted respectively, P i and Q i Respectively represent a bin value of their respective histograms, Indicates P i The weight of Indicates Q i The weight of is shown in formulas (5) and (6):

[0028]

[0029] Define the flow matrix F as shown in formula (7), where f ij Indicates that from P i to Q iThe amount of flow;

[0030] F=[f ij ] (7)

[0031] Definition d ij For P i to Q i The cost function is shown in formula (8), where K represents the weight sequence, A represents the number of bins in the training set histogram, B represents the number of bins in the histogram of the data to be predicted, and g(i) is the value of the feature corresponding to the i-th element in the training set histogram in the weight sequence K;

[0032]

[0033] Through linear programming, the value of the flow matrix F is calculated when formula (8) takes the minimum value. The calculation method is shown in formula (9), where: is the cost function, st represents the constraint condition;

[0034]

[0035] After using formula (9) to calculate the flow matrix F, use formula (10) to calculate the distance between the histogram Q of the data to be predicted and the histogram P of the training set, and use this distance to represent the similarity between the histograms Q and P. The smaller the distance, the more similar the two histograms are.

[0036]

[0037] Preferably, each data set in step S1 includes normal traffic and abnormal traffic.

[0038] The present invention also discloses an IoT security system for concept drift detection and adaptation, which includes four parts: data preparation module, training model candidate module, model selection module and prediction output module.

[0039] The advantages of the present invention are:

[0040] 1. This paper addresses security issues such as IoT attacks and malicious traffic identification. This paper proposes an IoT security system for concept drift detection and adaptation. By validating the effectiveness of IoT traffic in detecting data drift, the performance of AI models is evaluated, allowing the optimal AI model to be selected.

[0041] 2. The IoT security system proposed in this invention first performs data preparation, using two basic datasets, NF-ToN-IoT and NF-ToN-IoT-V2. Through preprocessing, the labels in the datasets are converted into two categories: normal traffic and abnormal traffic. A set of training sets and validation sets with different distributions are sampled from the NF-ToN-IoT-V2 dataset, and a set of test sets with different distributions are sampled from the NF-ToN-IoT dataset. In the model preparation stage, a batch of binary classification models are trained as candidate models on each training set, and the feature distribution of the training set is characterized using a histogram. In the model selection stage, the feature distribution of the test set is characterized using a histogram. Then, based on the similarity of the histograms, the most similar training set is found, and the model trained on this training set is used as the decision model. Through experiments, various similarity measurement methods were compared, and it was found that the similarity measurement using Self-adaptation EMD is more effective, and the best decision model can be selected from the candidate models. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A histogram is used to represent the similarity between data sets.

[0043] Figure 2 Comparison of the performance of the security framework of the present invention with the average baseline model in the labeled experimental scenario.

[0044] Figure 3 For real-world scenarios without labels, the performance of the proposed security framework is compared with the average baseline model.

[0045] Figure 4 This is the IoT security system model of the present invention. DETAILED DESCRIPTION

[0046] In order to make the technical means, creative features, objectives and effects achieved by the present invention easier to understand, the present invention is further described below in conjunction with specific implementation methods.

[0047] like Figures 1 to 4 As shown, a method for concept drift detection and adaptation includes the following steps:

[0048] S1. Data Preparation

[0049] Prepare the IoT dataset DataSet for model training as follows:

[0050] The bagging sampling method is used to sample the dataset DataSet with replacement using different positive and negative sample ratios. After sampling, M sub-training sets DataSet1, DataSet2, ..., DataSet are obtained. M .

[0051] For example, using traffic data for IoT anomaly detection can be done using two basic datasets: NF-ToN-IoT and NF-ToN-IoT-V2. Since IoT data streams typically generate large amounts of data continuously, using all data samples for learning model development is often infeasible and unnecessary. Therefore, effective data sampling methods should be employed to select highly representative data samples.

[0052] Prepare two IoT datasets A and B (each dataset contains normal traffic and abnormal traffic)

[0053] A: used to train the NF-ToN-IoT-V2 model, with 16,940,496 samples.

[0054] B: used to test the NF-ToN-IoT model, with 1,379,274 samples;

[0055] Using a bagging sampling method, we sampled two IoT datasets, A and B, with replacement, generating 812,570 bootstrap pseudo samples for each dataset. Using varying ratios of positive and negative samples (positive sample ratios ranging from 0.1, 0.2, to 0.9), we performed weighted random sampling from dataset A to generate nine subsets, A1, A2, …, A9 (each containing 1 / 9 the number of samples in dataset A).

[0056] Due to the large size of the IoT traffic dataset, a 5% sample of the original data was selected using a bagging sampling method to evaluate the proposed framework. The size of the resulting subset can vary, depending on the rate of IoT data generation and the computational power of the server machine. Compared to other sampling methods, bagging sampling primarily focuses on reducing variance, making it more effective on learners that are susceptible to sample perturbations, such as unpruned decision trees and neural networks.

[0057] S2. Training model candidates

[0058] In the model preparation stage, a batch of binary classification models are trained as candidate models on each training set, and a histogram is used to characterize the feature distribution of the training set; a weight sequence is used to characterize the importance of each feature in the model.

[0059] Training baseline model: Use M training sets and N features. On each training set, generate a histogram matrix H, train a classifier E and corresponding weight sequence K, and get a total of M histograms (H1, H2, ... H M ), M classifiers (E1, E2, … E M ) and M weight sequences (K1, K2, ... K M )

[0060] The histogram H of each training set is generated as follows:

[0061] For a certain feature X, suppose the feature value of X is a sequence {x1,x2,…x T}, T represents the length of the eigenvalue sequence, i represents the position in the eigenvalue sequence, and the elements in the eigenvalue sequence are divided into several ranges according to their size, each range is called a bin, h j Indicates the number of times the eigenvalue appears in the jth bin, h j The calculation method of is shown in formula (1).

[0062]

[0063] Among them I i is the indicator function, and its calculation method is shown in formula (2).

[0064]

[0065] The histogram of feature X can be calculated by formulas (1) and (2): h = {h1, h2, ..., h C}, using the same method, the histogram is calculated for each feature to obtain the histogram matrix h' as ​​shown in formula (3), where C is the number of bins in the histogram and N is the number of features.

[0066]

[0067] Convert the histogram matrix h' to one dimension and get the histogram H of the training set as shown in formula (4)

[0068] H={h 1,1 ,h 1,2 ,…,h 1,C ,h 2,1 ,h 2,2 ,…,h 2,C ,…,h N,1 ,h N,2 ,…,h N,C} (4)

[0069] The method for generating the classifier E and weight sequence K on each training set is as follows:

[0070] Select a classifier model and use the samples in the training set to train the model. The classifier is obtained through training. During the training, the importance of each feature to the model is recorded and the sequence {k1, k2, ..., k N} (N is the number of features), the sequence {k1, k2, ..., k N} is the weight sequence K.

[0071] To train a baseline model, you can select several features, for example, train multiple classifiers M1-M9, and generate histograms for D1-D9 (datasets). Based on the feature distribution of the dataset, a histogram is generated for each feature with equal width.

[0072] S3. Model selection

[0073] When processing features in the data processing stage, it may happen that two features are not of the same magnitude. If linear operations are used, it will easily lead to imbalance. This paper uses the max-min method to normalize the features.

[0074] Max-min performs a linear transformation on the data, mapping the X value to the range [0,1]. This method can be used for normalization when the feature data is relatively scattered or linear, and there are not many outliers. The formula is as follows:

[0075]

[0076] Among them, X and X norm Represent the features before and after normalization, X max and X min represent the maximum and minimum values ​​of the feature respectively.

[0077] During the model selection phase, a histogram is used to characterize the feature distribution of the test set. Then, based on the similarity of the histograms, the most similar training set is found and the model trained on that training set is used as the decision model. Experiments comparing various similarity metrics revealed that the self-adaptation EMD similarity metric performs better, allowing the optimal decision model to be selected from the candidate models.

[0078] EMD (Earth Mover's Distance) is an image similarity metric proposed in the 2000 IJCV article "The Earth Mover's Distance as a Metric for Image Retrieval." The EMD concept was originally used for image retrieval. Due to its various advantages, it has since been adopted for other similarity metrics. Considering that different features have varying impacts on model output, the self-adaptation EMD adds a weighting factor to the EMD distance to calculate the degree of similarity between two sets of data distributions.

[0079] The self-adaptation EMD similarity measurement method is used. This method first uses the histogram generation method to generate a histogram H' for the data to be predicted, and then calculates the similarity between H' and each training set histogram H. The calculation method is as follows:

[0080] Define P and Q to represent the histogram of training samples and data to be predicted, respectively. i and Q i Respectively represent a bin value of their respective histograms, Indicates P i The weight of Indicates Q i The weights of are shown in formulas (5) and (6).

[0081]

[0082] The flow matrix F is defined as shown in formula (7), where f ij Indicates that from P i to Q i The amount of flow.

[0083] F=[f ij ] (7)

[0084] Definition d ij For P i to Q i The cost function is shown in formula (8). Where K represents the weight sequence, A represents the number of bins in the training set histogram, B represents the number of bins in the histogram of the data to be predicted, and g(i) is the value of the feature corresponding to the i-th element in the training set histogram in the weight sequence K.

[0085]

[0086] Through linear programming, the value of the flow matrix F is calculated when formula (8) takes the minimum value. The calculation method is shown in formula (9), where, is the cost function, and st represents the constraint condition.

[0087]

[0088] After using formula (9) to calculate the flow matrix F, use formula (10) to calculate the distance between the histogram H' of the data to be predicted and the histogram H of the training set, and use this distance to represent the similarity between the histograms H' and H. The smaller the distance, the more similar the two histograms are.

[0089]

[0090] S4. Prediction output

[0091] By using the similarity histogram representation comparison of Self-adaption EMD, the most similar training set is found, and the model trained on this training set is used as the decision model, thereby overcoming the concept drift problem, improving system performance, and enhancing the accuracy of IoT security traffic identification.

[0092] Experimental equipment and preparation

[0093] The hardware used in the experiment includes: x86_64 architecture, 20 cores, 128GB memory, and a 256GB hard disk; the software includes: scipy (1.5.2), h2o (3.34.0.3), scikit-learn (1.0.1), pandas (1.2.1), and matplotlib (3.3.2), representing the mainstream server configuration for big data analysis of IoT data.

[0094] Two datasets were used to evaluate the proposed framework. The first, the NF-ToN-IoT dataset, utilizes publicly available pcaps to generate its NetFlow records, resulting in a NetFlow-based IoT network dataset. The total number of data flows is 1,379,274, of which 1,108,995 (80.4%) are attack samples and 270,279 (19.6%) are benign samples. The second, the NF-ToN-IoT-V2 dataset, contains a total of 16,940,496 data flows, of which 10,841,027 (63.99%) are attack samples and 6,099,469 (36.01%) are benign samples. Since both datasets are imbalanced, we use four metrics: precision, accuracy, recall, and f1 score to evaluate the overall performance of the proposed framework.

[0095] Experimental verification

[0096] Use the positive / negative sample ratio of 1:9, 2:8, 3:7…9:1, and generate 9 training sets (abnormal traffic is ddos). Use histogram + EMD to describe the similarity between the 9 data sets. Figure 1 As shown, the smaller the value, the lighter the color, and the more similar it is.

[0097] Figure 1 It intuitively represents the similarity between datasets. Table 1 numerically verifies that the similarity measurement method is effective. The closer the positive / negative data ratio is, the more similar the results are.

[0098] Table 1. Similarity between datasets using histogram data.

[0099]

[0100] Comparative performance improvement

[0101] 1) Tagged scenes

[0102] For labeled experimental scenarios, such as Figure 2 As shown in the figure, compared with the average baseline model performance, the proposed CDDAM framework can improve the accuracy by 3.2 percentage points, the precision by 1.7 percentage points and the F1 score by 6.6 percentage points, with the best improvement in recall, which can increase by 9.8 percentage points.

[0103] As shown in Table 2, the proposed CDDAM framework is compared with traditional models such as RF, GLM, GBM, DEEPLEARN, CNN, and BAYES. Although some of the models perform worse than RF and DEEPLEARN, this is due to overfitting during training, making them unsuitable for real-world open-world problems. On the other hand, while some of the models' performance is not optimal, the proposed CDDAM framework outperforms all six of the aforementioned models on average.

[0104] Table 2 Improvement of each model

[0105] accuracy precision recall f1_score rf -0.00012769 -0.00347093 -0.00141618 -0.00242901 glm 0.12639792 -0.00952783 0.25691671 0.15286328 gbm 0.02667125 -0.00116768 0.05314280 0.02668254 deeplearn -0.00212379 -0.01275931 -0.00141750 -0.00715475 cnn 0.02502224 0.13489374 0.27278561 0.22206910 bayes 0.01839170 -0.00731969 0.00940429 0.00231622 average 0.03237194 0.01677472 0.09823595 0.06572456

[0106] Table 3 shows the overall performance comparison of the proposed CDDAM framework with traditional models such as RF, GLM, GBM, DEEPLEARN, CNN and BAYES in the labeled scenario.

[0107] Table 3 Overall performance improvement (overall graph) - labeled scenario

[0108]

[0109] 2) Unlabeled real environment

[0110] For real scenes without labels, such as Figure 3 As shown in the figure, compared with the average baseline model performance, the proposed CDDAM framework can improve the accuracy by 2.9 percentage points and the F1 score by 4.6 percentage points. The best improvement is in the recall rate, which can increase by 7.3 percentage points, and the improvement in precision performance is not obvious.

[0111] As shown in Table 4, the proposed CDDAM framework is compared with traditional models such as RF, GLM, GBM, DEEPLEARN, CNN, and BAYES. Except for a small improvement in accuracy, all other performance improvements are significant. Similarly, due to overfitting in the training of RF and DEEPLEARN, the improvement is not significant, but these two models are actually unusable in real-world, unlabeled scenarios. On the other hand, the average performance of the proposed CDDAM framework is significantly improved compared to GLM, GBM, CNN, and BAYES. The average performance of the proposed CDDAM framework is better than all six models mentioned above.

[0112] Table 4 Improvement of each model

[0113] accuracy precision recall f1_score rf 0.00113408 -0.00321378 0.00189657 -0.00051347 glm 0.07158803 -0.02927362 0.15969995 0.07537674 gbm 0.03230318 -0.00197413 0.06375612 0.03277711 deeplearn 0.00070333 0.00302191 0.00155628 0.00223637 cnn 0.04232198 0.04145876 0.16595822 0.14079421 bayes 0.02721373 -0.00422575 0.05027182 0.02765012 average 0.02921072 0.00096556 0.07385649 0.04638684

[0114] Table 5 is the overall performance comparison of the proposed CDDAM framework with traditional models such as RF, GLM, GBM, DEEPLEARN, CNN and BAYES in the unlabeled scenario.

[0115] Table 5 Overall performance improvement (overall graph) - unlabeled scenario

[0116]

[0117] It is understood from common technical knowledge that the present invention may be implemented by other embodiments that do not depart from its spirit or essential features. Therefore, the embodiments disclosed above are, in all respects, merely illustrative and not exclusive. All modifications within the scope of the present invention or equivalent to the scope of the present invention are intended to be encompassed by the present invention.

Claims

1. A method for concept drift detection and adaptation, characterized in that: The steps include: S1. Data preparation Prepare IoT dataset DataSet for model training, perform replacement sampling on the dataset DataSet, and obtain M sub-training sets DataSet1, DataSet2, ..., DataSet M ; S2. Training model candidates In the model preparation stage, a classifier is trained as a candidate model on each training set, and a histogram is used to characterize the feature distribution of the training set, while recording the importance of each feature to the model; S3. Model selection In the model selection stage, a histogram is used to characterize the feature distribution of the data to be predicted. Then, based on the similarity of the histograms, the histogram representations are compared in the candidate models to find the most similar training set, and the model trained on the training set is used as the decision model. The step S3 adopts a self-adaption EMD similarity measurement method, which first uses the histogram generation method in step S3 to generate a histogram H' for the data to be predicted, and then calculates the similarity between H' and each training set histogram H respectively; In step S3, let H = P, H' = Q, that is, P and Q represent the histogram of the training sample and the data to be predicted respectively, P i and Q i Respectively represent a bin value of their respective histograms, Indicates P i The weight of Indicates Q i The weight of is shown in formulas (5) and (6): Define the flow matrix F as shown in formula (7), where f ij Indicates that from P i To Q i The number of flows; F=[f ij ] (7) Definition ij For P i To Q i The cost function is shown in formula (8), where K represents the weight sequence, A represents the number of bins in the training set histogram, B represents the number of bins in the data to be predicted histogram, and g(i) is the value of the feature corresponding to the i-th element in the training set histogram in the weight sequence K; Through linear programming, the value of the flow matrix F is calculated when formula (8) takes the minimum value. The calculation method is shown in formula (9), where: is the cost function, st represents the constraint condition; After using formula (9) to calculate the flow matrix F, use formula (10) to calculate the distance between the histogram Q of the data to be predicted and the histogram P of the training set, and use this distance to represent the similarity between the histograms Q and P. The smaller the distance, the more similar the two histograms are.

2. A method for concept drift detection and adaptation according to claim 1, characterized in that: In step S1, the bagging sampling method is adopted, and different positive and negative sample ratios are used to select 5% of the original data samples.

3. A method for concept drift detection and adaptation according to claim 1, characterized in that: The specific method in step S2 is as follows: Training baseline model: Use M training sets and N features. On each training set, generate a histogram matrix H, train a classifier E and corresponding weight sequence K, and get a total of M histograms (H1, H2, ... H M ), M classifiers (E1, E2, …E M ) and M weight sequences (K1,K2,…K M ); The histogram H of each training set is generated as follows: For a certain feature X, assume that the feature value of X is a sequence {x1,x2,…x T }, T represents the length of the eigenvalue sequence, i represents the position in the eigenvalue sequence, and the elements in the eigenvalue sequence are divided into several ranges according to their size, each range is called a bin, h j Indicates the number of times the eigenvalue appears in the jth bin, h j The calculation method is shown in formula (1): Among them I i is the indicator function, and the calculation method is shown in formula (2): The histogram of feature X can be calculated by formulas (1) and (2): h = {h1, h2, …, h C }, using the same method, calculate the histogram for each feature and get the histogram matrix h' as ​​shown in formula (3), where C is the number of bins in the histogram and N is the number of features: Convert the histogram matrix h' to one dimension and obtain the histogram H of the training set as shown in formula (4): H={h 1,1 ,h 1,2 ,…,h 1,C ,h 2,1 ,h 2,2 ,…,h 2,C ,…,h N,1 ,h N,2 ,…,h N,C } (4) The method for generating the classifier E and weight sequence K on each training set is as follows: Select a classifier model and use the samples in the training set to train the model. The classifier is obtained through training. During the training, the importance of each feature to the model is recorded and the sequence {k1, k2, …, k N }, N is the number of features, and the sequence {k1, k2, …, k N } is the weight sequence K.

4. A method for concept drift detection and adaptation according to claim 1, characterized in that: In step S1, each data set includes normal traffic and abnormal traffic.

5. An IoT security system obtained according to the method for concept drift detection and adaptation according to claim 1, characterized in that: It includes four parts: data preparation module, training model candidate module, model selection module and prediction output module.

Citation Information

Patent Citations

  • Detection method capable of processing reappearance concepts

    CN108171251A

  • Double-window concept drift detection method based on sample distribution statistical test

    CN110717543A