A Method for Feature Extraction and Selection of Internet of Things Device Recognition Based on Improved Honey Badger Algorithm

By improving the honey badger algorithm to perform feature extraction and feature selection in the IoT device recognition system, the problems of inefficiency and high computing overhead caused by high-dimensional features are solved, and efficient IoT device classification is achieved.

CN115660025BActive Publication Date: 2025-06-24JILIN UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211307323.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2025-06-24
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

The existing IoT device identification system has difficulties caused by high-dimensional features in the process of feature extraction and feature selection, resulting in inefficiency and high computing overhead.

Method used

The improved honey badger algorithm is adopted to capture traffic data from the Internet of Things environment, perform standardized preprocessing, and construct a multi-objective joint feature selection objective function. The improved honey badger algorithm is used to solve the feature subset and output the optimal feature subset.

Benefits of technology

Effectively reduce the dimension of feature subsets, improve the classification efficiency of IoT devices, reduce the calculation overhead and running time of classifiers, and achieve a classification accuracy of 95%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115660025B_ABST
    Figure CN115660025B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for extracting and selecting identification features of Internet of Things devices based on an improved honey badger algorithm, including: Step 1, capturing traffic data from an Internet of Things environment gateway and extracting Internet of Things traffic feature data; Step 2, performing standardized preprocessing on the extracted feature data; Step 3, constructing an objective function for multi-objective joint feature selection and evaluating the feature subset using the objective function; Step 4, solving the feature subset through the improved honey badger algorithm and outputting the optimal feature subset. It can extract features from the real Internet of Things traffic environment for selection and classification, effectively reduce the dimension of the feature subset, improve the classification efficiency of Internet of Things devices, reduce the computational overhead of the classifier, and reduce the running time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for extracting and selecting identification features of Internet of Things devices based on an improved honey badger algorithm, and belongs to the field of Internet of Things device identification. Background Art

[0002] With the rapid growth of the scale of the Internet of Things, various network security problems have become complex and diverse. Attackers can use the vulnerabilities of one device model to harm thousands of devices of the same type. In addition, due to the generally lower computing resources configured for Internet of Things devices, they are more vulnerable than ordinary computers and are more likely to suffer large-scale network attacks. Device identification is an important means to detect and prevent these security problems. In recent years, research on Internet of Things device identification systems has been continuously proposed. They usually extract features from network traffic based on machine learning methods and select a part of the features for classification. However, in this process, feature extraction and feature selection are often the short boards and difficulties of many studies. Also, due to the high-dimensional characteristics of network traffic features, therefore, developing a feature extraction and feature selection method for device identification can effectively overcome the defects in the above technologies and is more conducive to the development of Internet of Things device identification research. Summary of the Invention

[0003] The present invention designs and develops a method for extracting and selecting identification features of Internet of Things devices based on an improved honey badger algorithm, which can extract features from the real Internet of Things traffic environment for selection and classification, can effectively reduce the dimension of the feature subset, improve the classification efficiency of Internet of Things devices, reduce the computational overhead of the classifier, and reduce the running time.

[0004] The technical solution provided by the present invention is as follows:

[0005] A method for extracting and selecting identification features of Internet of Things devices based on an improved honey badger algorithm, comprising:

[0006] Step 1: Capture traffic data from the Internet of Things environment gateway and extract Internet of Things traffic feature data;

[0007] Step 2: Perform standardized preprocessing on the extracted feature data;

[0008] Step 3: Construct an objective function for multi-objective joint feature selection and evaluate the feature subset using the objective function;

[0009] Step 4: Solve the feature subset through the improved honey badger algorithm and output the optimal feature subset.

[0010] Preferably, in the second step, the standardization formula for the feature data is:

[0011]

[0012] where y i,j is the j-th eigenvalue of the i-th data, and y max is the maximum value of the j-th feature, and y min is the minimum value of the j-th feature.

[0013] Preferably, the formula of the objective function in the third step is:

[0014]

[0015]

[0016] where fitness is the fitness, ACC is the accuracy of the current model on the test set, num_feat is the number of features selected by the current search individual, max_feat is the total number of features, TP is the number of samples predicted as positive samples by the classifier, TN is the number of samples predicted as negative samples and actually being negative samples, and FN is the number of samples predicted as negative samples and actually being positive samples.

[0017] Preferably, the fourth step includes:

[0018] Step 1: Initialize the population through the Sine chaotic mapping and the population filtering mechanism;

[0019] Step 2: Introduce a sub-population mechanism, divide the current population into two sub-populations, and select the optimal solutions of each sub-population, which are respectively defined as the optimal solution and the sub-optimal solution of the current algorithm, and use the optimal solution and the sub-optimal solution to guide the position update of the two populations respectively;

[0020] Step 3: Perform binary mapping on the position vectors of the discrete solution space of the individuals in the population;

[0021] Step 4: Merge the sub-populations and output the optimal feature combination;

[0022] When the number of iterations does not meet the termination condition, repeat steps 2-4.

[0023] Preferably, the first step includes:

[0024] Use the Sine chaotic mapping to generate an initial population X with twice the number of individuals origin , and the Sine chaotic mapping formula includes:

[0025] h i+1 = μ × sin(π × h i );

[0026] X i = lb i + h i × (ubi +lb i );

[0027] Wherein, h i is the i-th chaotic number generated, μ is a constant, which is 0.99, lb i is the lower limit of the i-th solution, ub i is the upper limit of the i-th solution, X i is the i-th initial solution generated;

[0028] Calculate the fitness of individuals in the population through the objective function and sort them;

[0029] Take the first half of the individuals of X origin to form the population X.

[0030] Preferably, the step 2 includes:

[0031] Update the odor intensity factor I

[0032]

[0033] S = (X m - X m+1 ) 2 ;

[0034] d m = X best - X m ;

[0035] Wherein, I m is the odor intensity of the prey for the m-th honey badger individual, S is the concentration intensity, d m is the distance between the prey and the m-th honey badger;

[0036] Perform position update,

[0037]

[0038] Wherein, X new is the solution after position update, X best is the optimal solution of the population so far, F is the direction vector, taking -1 or 1, β is the ability of the honey badger to obtain food, the value is 6, I is the odor intensity factor, d i is the distance between the prey and the i-th honey badger individual, α is the balance factor, r3, r4, r5 are random numbers on [0, 1] respectively, Levy(λ) is the random step size generated by the Lévy distribution, X A , X B are two random solutions in the current population respectively.

[0039] Preferably, the method for generating the random step size subject to the Lévy distribution includes:

[0040]

[0041] Wherein, S is the generated random step size, u ~ N(0, σ 2 ), v ~ N(0, 1), u follows a normal distribution with a mathematical expectation of 0 and a variance of σ 2 , v follows a standard normal distribution with a mathematical expectation μ = 0 and a variance σ = 1, and η is a constant, taking 1.5;

[0042]

[0043] Wherein, Γ() is the gamma function, and η is a constant, taking 1.5.

[0044] Preferably, the step 3 includes:

[0045] Performing binary mapping on the position vectors in the discrete solution space of the individuals in the population, and the mapping function is:

[0046]

[0047] Wherein, x binary is the solution after binary conversion, x is the solution in the continuous space, and thres is the threshold, taking 0.5.

[0048] Advantages of the present invention:

[0049] A method for feature extraction and feature classification of Internet of Things device identification proposed by the present invention can extract features from the real Internet of Things traffic environment for selection and classification, without relying on existing datasets in the past. The extracted dataset can achieve a classification accuracy of 95%. The improved honey badger algorithm for Internet of Things device identification feature selection has better performance than the original algorithm and is superior to other similar algorithms. The accuracy and fitness value have both increased, which can effectively reduce the dimension of the feature subset, reduce the computational overhead of the classifier, reduce the running time, and has a wide application prospect in the field of Internet of Things device identification. Description of the Drawings

[0050] Figure 1 It is a flowchart of the method for Internet of Things device identification feature extraction and selection based on the improved honey badger algorithm of the present invention.

[0051] Figure 2 It is a flowchart of the improved binary honey badger algorithm for solving the optimal feature subset of the present invention. Detailed Embodiments

[0052] The following further describes the present invention in detail with reference to the drawings, so that those skilled in the art can implement it according to the description in the specification.

[0053] AsFigure 1-2 As shown in Figure 1-2 , the present invention provides a method for extracting and selecting identification features of Internet of Things devices based on an improved honey badger algorithm, including:

[0054] Step 1: Capture traffic data from the Internet of Things environment gateway and extract Internet of Things traffic feature data;

[0055] Use the network packet tool wireshark to capture pcap or pcapng data files from the gateway of the Internet of Things environment;

[0056] Use the scapy module of the python language to parse information such as the protocol headers and data packets of each data packet. The included protocols include Ethernet, LLC, EAPOL, IP, ICMP, TCP, UDP, BOOTP, DNS, NTP, TLS, SSL. In addition, it also includes traffic features such as packet_size, payload_bytes, and protocols, and a total of 111-dimensional feature data is obtained;

[0057] Step 2: Perform standardized preprocessing on the extracted feature data;

[0058] Perform data normalization operations of numericalization, duplicate removal, and missing value filling on the obtained data, and perform data standardized preprocessing. The data standardization formula used is as follows:

[0059]

[0060] where y i,j is the j-th feature value of the i-th data, y max is the maximum value of the j-th feature, y min is the minimum value of the j-th feature;

[0061] Step 3: Construct an objective function for multi-objective joint feature selection and evaluate the feature subset using the objective function;

[0062] The formula of the objective function is:

[0063]

[0064]

[0065] where fitness is the fitness, ACC is the accuracy rate of the current model on the test set, num_feat is the number of features selected by the current search individual, max_feat is the total number of features, TP is the number of samples predicted as positive samples by the classifier, TN is the number of samples predicted as negative samples that are actually negative samples, and FN is the number of samples predicted as negative samples that are actually positive samples.

[0066] Step 4: Solve the feature subset by the improved honey badger algorithm and output the optimal feature subset, including:

[0067] Step 1: Initialize the population using the Sine chaotic map and obtain the first half of the better solutions using the population filtering mechanism. The Sine chaotic map and the population filtering mechanism include:

[0068] Generate an initial population X with a size twice the number of individuals using the Sine chaotic map origin , and the Sine chaotic map formula is:

[0069] h i+1 = μ × sin(π × h i );

[0070] X i = lb i + h i × (ub i + lb i );

[0071] In the formula, h i is the i-th generated chaotic number, μ is a constant, which is 0.99, lb i is the lower limit of the i-th solution, and ub i is the upper limit of the i-th solution. X i is the i-th generated initial solution;

[0072] Calculate the fitness of individuals in the population through the objective function and sort them;

[0073] Select the first half of the individuals in X origin to form the population X.

[0074] Update the balance factor α, and the formula is as follows:

[0075]

[0076] In the formula, t max is the maximum number of iterations, and t is the current number of iterations;

[0077] Step 2: Introduce a sub-population mechanism, divide the current population into two sub-populations, and select the optimal solution of each sub-population respectively, which are defined as the optimal solution and the sub-optimal solution of the current algorithm. Use the optimal solution and the sub-optimal solution to guide the position update of the two populations respectively. The specific process includes:

[0078] Calculate the fitness values of individuals in the population using the objective function and sort them;

[0079] Divide the sorted population into two sub-populations according to odd and even indices;

[0080] The best fitness values in each sub-population are the optimal and sub-optimal solutions of the corresponding algorithms, respectively.

[0081] Update the odor intensity factor I

[0082]

[0083] S = (X m - X m+1 ) 2 ;

[0084] d m = X best - X m ;

[0085] In the formula, I m is the odor intensity of the prey for the m-th honey badger individual, S is the concentration intensity, and d m is the distance between the prey and the m-th honey badger;

[0086] Update the positions of the two populations using the following formula respectively.

[0087]

[0088] In the formula, X new is the solution after position update, X best is the optimal solution of this population so far, F is the direction vector, taking -1 or 1, β is the ability of the honey badger to obtain food, with a value of 6, I is the odor intensity factor, d i is the distance between the prey and the i-th honey badger individual, α is the balance factor, r3, r4, and r5 are random numbers on [0, 1] respectively, Levy(λ) is the random step size generated by the Lévy distribution, X A , X B are two random solutions in the current population respectively.

[0089] Introduce Lévy flight in the formula X new = X best + Levy(λ) × α × (Levy(λ) × X A - X B ) to enhance the local optimization ability. The method for generating a random step size that follows the Lévy distribution includes:

[0090]

[0091] In the formula, S is the generated random step size, u ~ N(0, σ 2 ), v ~ N(0, 1), u follows a normal distribution with a mathematical expectation of 0 and a variance of σ 2 , v follows a standard normal distribution with a mathematical expectation μ = 0 and a variance σ = 1, and η is a constant, taking 1.5;

[0092]

[0093] Wherein, Γ() is the gamma function, η is a constant, taking 1.5.

[0094] Step 3: Perform binary mapping on the position vectors in the discrete space of the individuals in the population. The mapping function is as follows:

[0095]

[0096] Wherein, x binary is the solution after binaryization, x is the solution in the continuous space, and thres is the threshold, taking 0.5.

[0097] Step 4: Combine the current two sub-populations into one population and obtain the optimal solution of the current algorithm

[0098] Judge whether the iteration termination condition is satisfied. When it is not satisfied, repeat Steps 2-4;

[0099] Output the optimal feature subset.

[0100] Although the embodiments of the present invention have been disclosed as above, they are not limited to only the applications listed in the description and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily achieved. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated examples here.

Claims

1. An identification feature extraction and selection method for Internet of Things devices based on an improved honey badger algorithm, characterized in that, Including: Step 1: Capture traffic data from the Internet of Things environment gateway and extract Internet of Things traffic feature data; Step 2: Perform standardized preprocessing on the extracted feature data; Step 3: Construct an objective function for multi-objective joint feature selection and evaluate the feature subset using the objective function; The formula of the objective function is: Where fitness is the fitness, ACC is the accuracy of the current model on the test set, num_feat is the number of features selected by the current search individual, max_feat is the total number of features, TP is the number of samples predicted as positive samples by the classifier, TN is the number of samples predicted as negative samples and actually negative samples by the classifier, and FN is the number of samples predicted as negative samples and actually positive samples by the classifier; Step 4: Solve the feature subset through an improved honey badger algorithm and output the optimal feature subset, including: Step 1: Initialize the population through Sine chaotic mapping and population filtering mechanism; Step 2: Introduce a sub-population mechanism, divide the current population into two sub-populations, and select the optimal solution of each sub-population respectively, which are defined as the optimal solution and the sub-optimal solution of the current algorithm respectively, and use the optimal solution and the sub-optimal solution to guide the position update of the two populations respectively; Step 3: Perform binary mapping on the position vector of the discrete solution space of the individuals in the population; Step 4: Merge the sub-populations and output the optimal feature combination; When the number of iterations does not meet the termination condition, repeat steps 2-4.

2. The method for extracting and selecting identification features of Internet of Things devices based on the improved honey badger algorithm according to claim 1, characterized in that In step 2, the formula for standardizing the feature data is: where y i,j is the j-th eigenvalue of the i-th data, y max is the maximum value of the j-th feature, y min is the minimum value of the j-th feature.

3. The method for extracting and selecting the identification features of Internet of Things devices based on the improved honey badger algorithm according to claim 2, wherein Step 1 includes: Generate an initial population X with twice the number of individuals using the Sine chaotic map origin , and the Sine chaotic map formula includes: h i+1 = μ × sin(π × h i ); X i = lb i + h i × (ub i + lb i ); where h i is the i-th generated chaotic number, μ is a constant, which is 0.99, lb i is the lower bound of the i-th solution, ub i is the upper bound of the i-th solution, X i is to generate the i-th initial solution; Calculate the fitness of the individuals in the population through the objective function and sort them; Select the first half of the individuals of X origin to form population X.

4. The method for extracting and selecting identification features of Internet of Things devices based on the improved honey badger algorithm according to claim 3, characterized in that Step 2 includes: Update the odor intensity factor I S = (X m - X m+1 ) 2 ; d m = X best - X m ; Where I m is the odor intensity of the prey for the m-th honey badger individual, S is the concentration intensity, and d m is the distance between the prey and the m-th honey badger; Perform position update, where X new is the solution after position update, X best is the optimal solution of the population so far, F is the direction vector, taking -1 or 1, β is the ability of the honey badger to obtain food, with a value of 6, I is the odor intensity factor, d i is the distance between the prey and the i-th honey badger individual, α is the balance factor, r3, r4, and r5 are random numbers on [0, 1] respectively, Levy(λ) is the random step size generated by the Lévy distribution, X A , X B are two random solutions in the current population respectively.

5. The method for extracting and selecting identification features of Internet of Things devices based on the improved honey badger algorithm according to claim 4, characterized in that The method for generating a random step size obeying the Lévy distribution includes: where S is the generated random step size, u ~ N(0, σ 2 ), v ~ N(0, 1), u follows a normal distribution with a mathematical expectation of 0 and a variance of σ 2 , v follows a standard normal distribution with a mathematical expectation μ = 0 and a variance σ = 1, and η is a constant, taking 1.5; In the formula, Γ() is the gamma function, η is a constant, taking 1.

5.

6. The method for extracting and selecting the identification features of the Internet of Things devices based on the improved honey badger algorithm according to claim 5, wherein Step 3 includes: Perform binary mapping on the position vector of the discrete solution space of the individuals in the population, and the mapping function is: where x binary is the solution after binarization, x is the solution in the continuous space, and thres is the threshold, taking 0.5.

Citation Information

Patent Citations

  • Low-cost and high-accuracy feature selection method

    CN114547405A

  • Differential evolution-based feature selection

    WO2014195782A2