Internet of Things malicious traffic detection method and device, and storage medium
By converting malicious IoT traffic data into grayscale images and using the space-time capsule Transformer model to extract spatial and temporal features, the existing detection methods are solved inadequate detection problems in the face of complex traffic patterns and new variant attacks, and efficient and reliable real-time detection is achieved.
Patent Information
- Application Number
- CN202510782647.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing malicious IoT traffic detection methods are insufficient in the face of complex and variable traffic patterns and new variant attacks, and are difficult to achieve real-time detection with low power consumption and low latency.
The text imaging method is used to convert malicious traffic data into grayscale images, and combined with local binary mode features, global image features and high-order local automatic correction features, a malicious traffic feature database is built, and the detection model is trained using the space-time capsule Transformer, spatial and time series features are extracted, and classification training is performed.
It improves the adaptability and detection accuracy of variant malicious traffic, reduces the computational complexity, enhances the real-time detection capabilities on IoT devices, and achieves efficient and reliable malicious traffic detection.
Smart Images

Figure CN120498829A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet of Things security technology, and in particular to an Internet of Things malicious traffic detection method, device and storage medium. Background Art
[0002] With the rapid development of IoT technology, an increasing number of smart devices are connecting to the internet, encompassing devices in diverse fields such as smart homes, smart industry, and smart healthcare. These devices connect via wireless communication protocols, enabling remote monitoring, automated control, and data sharing, significantly improving productivity and enhancing convenience. However, the widespread use of IoT devices also presents serious security risks. Malicious attackers exploit security vulnerabilities in IoT devices to launch large-scale malicious traffic attacks, causing device failures, data leaks, and even complete network disruptions. Common malicious IoT traffic attacks include distributed denial of service attacks, malware distribution, data theft, and network hijacking. These attacks can not only misuse system resources but also damage critical infrastructure, seriously threatening the security and stability of IoT networks.
[0003] Currently, detection methods for malicious IoT traffic primarily include rule-based, statistical, and machine learning methods. Rule-based methods rely on predefined traffic signature libraries, such as deep packet inspection and intrusion detection systems, to determine attack activity by comparing the signatures of known malicious traffic. While this approach offers high accuracy in detecting known attacks, it is less effective against new variants of malicious traffic. Furthermore, the signature library requires constant updating, resulting in high maintenance costs. Statistical methods analyze behavioral characteristics of traffic data, such as packet size and transmission interval, and utilize statistical models to determine traffic anomalies. While these methods can detect certain anomalous patterns, they are susceptible to noise, have a high false positive rate, and are less effective for stealthy attacks, such as low-rate DDoS attacks. Furthermore, machine learning methods utilize classification algorithms trained on traffic data to automatically distinguish between normal and malicious traffic, reducing reliance on manual feature engineering and improving detection efficiency. However, traditional machine learning methods still rely on manually constructed features, making them difficult to adapt to the complex and changing traffic patterns of IoT devices, resulting in insufficient generalization capabilities.
[0004] In recent years, deep learning technologies, such as convolutional neural networks, recurrent neural networks, and long short-term memory networks, have been widely used in malicious traffic detection. These methods can automatically extract features from traffic data, improve detection accuracy, and enhance adaptability to new malicious traffic variants. However, existing deep learning methods still face several challenges in the IoT environment. Existing methods often lack the ability to jointly model the spatial and temporal characteristics of traffic data. Most methods focus solely on temporal patterns or isolated spatial features. However, malicious IoT traffic often exhibits complex spatiotemporal correlations, and existing methods are unable to fully utilize this information, resulting in insufficient detection capabilities. Furthermore, when faced with new malicious traffic variants, deep learning models trained based on static features struggle to adapt to the rapid changes in attack methods. Detection accuracy decreases over time, impacting practical application effectiveness. Furthermore, deep learning models are generally computationally complex, making them difficult to run efficiently on IoT devices with limited computing resources. This hinders the ability to achieve low-power, low-latency, and real-time detection.
[0005] Therefore, how to provide a method, device and storage medium for detecting malicious traffic in the Internet of Things is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention
[0006] One purpose of the present invention is to propose a method, device and storage medium for detecting malicious traffic in the Internet of Things. The present invention describes in detail a spatial feature extraction method based on a capsule network and a time series modeling algorithm based on a spatiotemporal attention mechanism, and constructs a malicious traffic classification model suitable for the Internet of Things environment. The model has the advantages of high detection accuracy, strong adaptability to variant malicious traffic and good real-time performance.
[0007] According to an embodiment of the present invention, a method for detecting malicious traffic in the Internet of Things includes the following steps:
[0008] S1. Collect malicious traffic data from IoT devices, classify them, and build a malicious traffic sample library.
[0009] S2. Use text imaging method to convert malicious traffic data into grayscale image, and perform normalization, noise reduction and feature enhancement to obtain processed grayscale image;
[0010] S3. Based on the processed grayscale image, extract local binary pattern features, global image features, and high-order local automatic correction features, screen highly relevant features, and establish a malicious traffic feature database;
[0011] S4. Based on the malicious traffic feature database, the space-time capsule Transformer is used to train a malicious traffic detection model, extract the spatial features of the traffic, and combine the space-time attention mechanism to model the traffic time series. Classification training is performed to generate a malicious traffic classification model.
[0012] S5. Collect real-time traffic data, convert it into grayscale images, extract features, use the trained classification model to extract spatial and temporal features, calculate anomaly scores, and determine traffic categories.
[0013] S6. Send the detection results to the command and control server, trigger an alarm signal, and send response instructions to the attacked device, including blocking the connection, isolating the device, or adjusting the security policy.
[0014] Optionally, the S2 specifically includes:
[0015] S21. Read the byte data of the target traffic data file, one byte at a time, and store them into a one-dimensional array in the order in which the data is read;
[0016] S22. Determine the width of the grayscale image:
[0017]
[0018] Where N is the total number of bytes read, W is the width of the grayscale image, Indicates floor operation;
[0019] S23. Convert the data dimension, convert the one-dimensional array into a two-dimensional array, and calculate the height of the grayscale image:
[0020]
[0021] Where H is the height of the grayscale image. If N is not divisible by W, zero padding is used to make the matrix dimension meet W×H.
[0022] S24. Generate a grayscale image and calculate pixel values:
[0023] G ij =b k ;
[0024] Among them, G ij is the pixel value of the i-th row and j-th column in the grayscale image matrix, b k is the kth byte in the traffic data, with pixel values ranging from 0 to 255;
[0025] S25, normalizing the generated grayscale image, adjusting the image size using a nearest neighbor interpolation algorithm, normalizing the image to a square, and ensuring that the normalized width and height are equal;
[0026] S26. Store the normalized grayscale image and output it to the malicious traffic feature extraction module.
[0027] Optionally, the S3 specifically includes:
[0028] S31. Extract the local binary pattern features of the grayscale image and calculate the local binary pattern values of the pixels:
[0029] LBP ij =∑s(P k -G ij )·2 k ;
[0030] Among them, LBP ij is the local binary pattern value of the pixel in the i-th row and j-th column in the grayscale image matrix, G ij is the gray value of the pixel, P k are the eight neighboring pixel values of the pixel point, k is the index of the neighboring pixel, ranging from 0 to 7, and s(x) is the binarization function, which is calculated as follows:
[0031]
[0032] S32, calculating the global Gist features of the grayscale image, filtering the image using a preset filter group, and performing histogram statistics on the filtered results to form a feature vector;
[0033] S33, calculating high-order local automatic correction features, performing correlation calculation on pixels of the image at different offset distances, counting the autocorrelation values, and forming a feature vector;
[0034] S34, fuses local binary pattern features, global Gist features, and high-order local automatic correction features to form a comprehensive feature vector of malicious traffic;
[0035] S35. Calculate the feature variance of normal traffic and malicious traffic, assuming a feature F k The variances in the normal traffic dataset and the malicious traffic dataset are and Calculate the evaluation value of the feature:
[0036]
[0037] Among them, V k is feature F k The evaluation value of is the feature F in the normal traffic data set k The variance of is the feature F in the malicious traffic dataset k variance;
[0038] S36. Sort the features according to the evaluation values, and select features with higher evaluation values to construct a malicious traffic feature database;
[0039] S37. Store the filtered feature data and use it for training a malicious traffic detection model.
[0040] Optionally, the S4 specifically includes:
[0041] S41. Label the normal traffic data and the malicious traffic data to construct a normal traffic dataset and a malicious traffic dataset, wherein the normal traffic dataset includes multiple normal traffic samples, and the malicious traffic dataset includes multiple malicious traffic samples, and assign a corresponding category label to each sample;
[0042] S42. Use the space-time capsule transformer to extract traffic features. First, the grayscale image of the input traffic sample is transformed through the capsule network to obtain the spatial feature representation of the traffic data. Then, the space-time transformer is used to model the traffic time series information. The spatial and time series features of the traffic sample are combined to generate the final traffic feature representation.
[0043] S43, feature mapping is performed on the extracted traffic features and input into the classification network. The classification network calculates the classification output result based on the extracted spatial features and temporal features, determines whether the traffic sample is malicious traffic, and outputs the detection result;
[0044] S44. Calculate the classification error. Let the true category label be y and the output probability of the classification network be The loss function is calculated using cross entropy loss:
[0045]
[0046] Among them, L loss function, y represents the true category label of the sample, Represents the sample category probability calculated by the classification network;
[0047] S45. Optimize the model parameters based on the calculated loss function, set the learning rate to α, and use the gradient descent algorithm to update the model parameters:
[0048]
[0049] Among them, W′ is the updated model parameter, W is the model parameter, is the gradient of the loss function with respect to the weight, and the learning rate α is used to control the update step size;
[0050] S46: Test the trained model to evaluate the detection accuracy of the model. If the detection accuracy reaches a set threshold, store the trained model parameters. Otherwise, continue training until the model meets the set accuracy requirements.
[0051] S47. Store the trained model parameters and use them for real-time malicious traffic detection. Use the stored model parameters to detect real-time traffic data, identify whether there is malicious traffic, and output the detection results.
[0052] Optionally, the S42 specifically includes:
[0053] S421. Perform feature mapping on the grayscale image of the traffic sample. Let the input grayscale image matrix be X. Obtain the initial feature map through multi-layer convolution transformation:
[0054] U i =σ(W i *X+b i );
[0055] Among them, U i is the initial feature map, W i is the convolution kernel weight, * represents the convolution operation, b i is the bias term, σ(·) is the nonlinear activation function;
[0056] S422, build a capsule network to represent the features at a high level, and set the initial capsule feature to be V i , use affine transformation to calculate the candidate capsule activation vector:
[0057]
[0058] Among them, W ij is the affine transformation matrix, b ij is the bias term, is the candidate capsule vector, U i is the input feature map;
[0059] S423, use the dynamic routing mechanism to calculate the coupling coefficient between capsules, and set the initial weight from capsule i to capsule j to be c ij , the coupling distribution is calculated by softmax normalization:
[0060]
[0061] Where t is the number of routing iterations, b ij is the coupling weight from capsule i to capsule j, b ik is the coupling weight from capsule i to capsule k, V j is the target capsule vector, is the coupling weight calculated for the previous round t, is the coupling weight after iteration t+1 rounds, is the candidate affine transformation vector of input capsule i;
[0062] S424. Perform nonlinear normalization on the updated capsule vector and calculate the final capsule output:
[0063]
[0064]
[0065] Among them, s j is the capsule input after weighted summation, v j For the final capsule output, the squash function is used to ensure the shrinkage and stability of the vector norm;
[0066] S425, based on the spatiotemporal Transformer, temporal feature modeling is performed, and the input is the capsule vector sequence V at time step t. t , calculate the self-attention weight:
[0067]
[0068] Among them, Q t is the query matrix, K t is the bond matrix, d k is the feature dimension, A t is the feature representation of spatiotemporal attention calculation, and softmax(·) is the attention score normalization function;
[0069] S426, use the multi-head attention mechanism to optimize the temporal features, and set the multi-head attention output M t The calculation is as follows:
[0070]
[0071] Among them, H is the number of attention heads, is the projection weight of the query matrix, is the projection weight of the bond matrix, is the projection weight of the value matrix, α h is the attention coefficient, M t is the temporal feature calculated by multi-head attention, is the input key matrix, V t is the input value matrix;
[0072] S427, fusion capsule network extracted spatial features V j and the temporal features M calculated by the spatiotemporal Transformer t , calculate the final fusion features:
[0073] Z=LayerNorm(V j +γM t );
[0074] Among them, Z is the fused feature vector, V j is the spatial feature extracted by the capsule network, M t is the temporal feature calculated by Transformer, γ is the trainable weight parameter, and LayerNorm performs feature normalization;
[0075] S428. Perform high-dimensional feature transformation on the fused feature vector Z and calculate the final feature using nonlinear mapping:
[0076] F=σ(W f Z+b f )+λtanh(W g Z+b g );
[0077] Among them, F is the high-dimensional feature after the final conversion, W f is the linear transformation matrix, W g is the nonlinear transformation matrix, b f and b g is the corresponding bias term, σ(·) is the sigmoid activation function, and tanh(·) is the hyperbolic tangent activation function;
[0078] S429, input feature F for final classification, and use Softmax to calculate the classification probability:
[0079]
[0080] Among them, P is the probability distribution after classification, and W T is the weight matrix of the classification layer, b is the bias term of the classification layer, and c is the category index.
[0081] Optionally, the S5 specifically includes:
[0082] S51. Collect real-time traffic data, divide it according to the set time window, obtain traffic data segments, and perform normalization processing;
[0083] S52. Use the space-time capsule transformer to extract spatial and temporal features, input the traffic data into the capsule network to calculate the spatial features, and perform time series modeling through the space-time capsule transformer to generate a fusion feature vector;
[0084] S53. Calculate the anomaly score. Let the mean eigenvector of the normal traffic sample be Calculate the anomaly score S:
[0085]
[0086] Among them, S is the anomaly score, Z is the fusion feature vector of the traffic sample to be detected, is the characteristic mean of normal traffic samples;
[0087] S54. Set an anomaly score threshold T and determine the traffic category based on the anomaly score S. The determination rules are as follows:
[0088]
[0089] S55. Output the traffic category detection result. If the traffic is determined to be malicious traffic, send an alarm signal to the command and control server, and execute blocking, isolation or adjust the security policy.
[0090] An IoT malicious traffic detection device according to an embodiment of the present invention includes:
[0091] The data acquisition module is used to collect network traffic data of IoT devices, convert the traffic data format, and extract traffic feature data;
[0092] The real-time detection module is used to input the real-time collected traffic data into the trained malicious traffic detection model and analyze the traffic data to detect whether there is malicious traffic;
[0093] A result sending module, used to send the detection result to the command and control server when malicious traffic is detected;
[0094] The alarm and response module is used to receive alarm signals returned by the command and control server and send response instructions to the attacked IoT devices to perform operations such as blocking connections, isolating devices, or adjusting security policies.
[0095] A computer-readable storage medium according to an embodiment of the present invention is characterized in that a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the processor is enabled to execute the method for detecting malicious traffic in the Internet of Things according to any one of claims 1 to 6.
[0096] The beneficial effects of the present invention are:
[0097] This invention provides a method for detecting malicious traffic in the Internet of Things (IoT). Addressing the challenges faced by existing malicious traffic detection methods in IoT environments, such as high heterogeneity, limited computing resources, and insufficient detection capabilities for new attack variants, this paper proposes a traffic analysis method based on a space-time capsule transformer. Using text imaging technology, malicious traffic data is converted into grayscale images. This is then combined with a capsule network to extract spatial features. The spatiotemporal attention mechanism is then used to model time series information. This allows the detection model to fully capture the global structural information and local pattern variations of malicious traffic, improving its adaptability to variant malicious traffic. Furthermore, the invention constructs a malicious traffic feature database and selects highly relevant features during training, enabling the model to extract more representative traffic features, thereby improving detection accuracy. Furthermore, the space-time capsule transformer architecture employed effectively reduces redundant computational overhead. Compared to traditional deep learning methods, this method improves the deployment efficiency and real-time performance of the model on IoT devices while maintaining detection performance. This invention enables efficient detection of malicious traffic in IoT environments, possesses strong generalization and real-time responsiveness, and provides a reliable technical solution for IoT security. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0099] Figure 1 This is a flow chart of a method for detecting malicious traffic in the Internet of Things proposed by the present invention;
[0100] Figure 2 This is a schematic diagram of the structure of the space-time capsule Transformer for time series modeling in the present invention. DETAILED DESCRIPTION
[0101] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0102] refer to Figure 1-2 , a method for detecting malicious traffic in the Internet of Things, comprising the following steps:
[0103] S1. Collect malicious traffic data from IoT devices, classify them, and build a malicious traffic sample library.
[0104] S2. Use text imaging method to convert malicious traffic data into grayscale image, and perform normalization, noise reduction and feature enhancement to obtain processed grayscale image;
[0105] S3. Based on the processed grayscale image, extract local binary pattern features, global image features, and high-order local automatic correction features, screen highly relevant features, and establish a malicious traffic feature database;
[0106] S4. Based on the malicious traffic feature database, the space-time capsule Transformer is used to train a malicious traffic detection model, extract the spatial features of the traffic, and combine the space-time attention mechanism to model the traffic time series. Classification training is performed to generate a malicious traffic classification model.
[0107] S5. Collect real-time traffic data, convert it into grayscale images, extract features, use the trained classification model to extract spatial and temporal features, calculate anomaly scores, and determine traffic categories.
[0108] S6. Send the detection results to the command and control server, trigger an alarm signal, and send response instructions to the attacked device, including blocking the connection, isolating the device, or adjusting the security policy.
[0109] This invention optimizes the entire chain from data collection, feature extraction, model training, to real-time detection and response by constructing a complete IoT malicious traffic detection process. First, a text imaging method is used to convert traffic data into grayscale images. This is combined with normalization, noise reduction, and feature enhancement techniques to improve data stability and usability. Subsequently, key features of malicious traffic are extracted using local binary pattern features, global image features, and high-order local autocorrection features. A high-quality malicious traffic feature database is constructed, enhancing the ability to identify different types of malicious traffic. By training the malicious traffic detection model using a space-time capsule transformer, a deep fusion of spatial and temporal traffic features is achieved, improving detection accuracy and adaptability to variant malicious traffic. During the real-time detection phase, the invention utilizes the trained classification model to extract features and calculate anomaly scores on the collected traffic data, ensuring efficient and real-time detection. Finally, the detection results are transmitted to a command and control server, triggering an alarm and implementing appropriate security policies. This enables IoT devices to quickly respond to malicious attacks and mitigate security risks. This invention improves the accuracy, adaptability, and real-time performance of IoT malicious traffic detection, providing an efficient and reliable technical solution for network security in IoT environments.
[0110] In this embodiment, S2 specifically includes:
[0111] S21. Read the byte data of the target traffic data file, one byte at a time, and store them into a one-dimensional array in the order in which the data is read;
[0112] S22. Determine the width of the grayscale image:
[0113]
[0114] Where N is the total number of bytes read, W is the width of the grayscale image, Indicates floor operation;
[0115] S23. Convert the data dimension, convert the one-dimensional array into a two-dimensional array, and calculate the height of the grayscale image:
[0116]
[0117] Where H is the height of the grayscale image. If N is not divisible by W, zero padding is used to make the matrix dimension meet W×H.
[0118] S24. Generate a grayscale image and calculate pixel values:
[0119] G ij =b k ;
[0120] Among them, G ij is the pixel value of the i-th row and j-th column in the grayscale image matrix, b k is the kth byte in the traffic data, with pixel values ranging from 0 to 255;
[0121] S25, normalizing the generated grayscale image, adjusting the image size using a nearest neighbor interpolation algorithm, normalizing the image to a square, and ensuring that the normalized width and height are equal;
[0122] S26. Store the normalized grayscale image and output it to the malicious traffic feature extraction module.
[0123] The present invention improves the distinguishability of malicious traffic features and the adaptability of model training by improving the processing method of traffic data. First, by reading the traffic data byte by byte and storing it in a one-dimensional array in the order of the data, the integrity of the data is ensured, while avoiding the loss of features caused by block reading. When determining the width of the grayscale image, the ratio of the total number of bytes to the fixed width is rounded down to ensure the stability of the image construction, and the matrix dimension is adjusted by zero padding to make the grayscale image structure after data conversion more regular, which is conducive to subsequent feature extraction. When calculating the pixel value, the byte value of the traffic data is directly mapped to the grayscale range of 0 to 255, so that the traffic data visually presents a stable distribution feature, which helps the machine learning model to mine potential patterns. In addition, the nearest neighbor interpolation method is used to normalize the grayscale image so that the image is unified into a square, improve data consistency, reduce the impact of size changes, and optimize the input format of the model, so that the subsequent malicious traffic feature extraction is more efficient. The preprocessing method of the present invention effectively reduces the computational complexity caused by inconsistent data dimensions, enhances the spatial distribution stability of features, and improves the accuracy and generalization ability of malicious traffic detection.
[0124] In this embodiment, S3 specifically includes:
[0125] S31. Extract the local binary pattern features of the grayscale image and calculate the local binary pattern values of the pixels:
[0126] LBP ij =∑s(P k -G ij )·2 k ;
[0127] Among them, LBP ij is the local binary pattern value of the pixel in the i-th row and j-th column in the grayscale image matrix, G ij is the gray value of the pixel, P k are the eight neighboring pixel values of the pixel point, k is the index of the neighboring pixel, ranging from 0 to 7, and s(x) is the binarization function, which is calculated as follows:
[0128]
[0129] S32, calculating the global Gist features of the grayscale image, filtering the image using a preset filter group, and performing histogram statistics on the filtered results to form a feature vector;
[0130] S33, calculating high-order local automatic correction features, performing correlation calculation on pixels of the image at different offset distances, counting the autocorrelation values, and forming a feature vector;
[0131] S34, fuses local binary pattern features, global Gist features, and high-order local automatic correction features to form a comprehensive feature vector of malicious traffic;
[0132] S35. Calculate the feature variance of normal traffic and malicious traffic, assuming a feature F k The variances in the normal traffic dataset and the malicious traffic dataset are and Calculate the evaluation value of the feature:
[0133]
[0134] Among them, V k is feature F k The evaluation value of is the feature F in the normal traffic data set k The variance of is the feature F in the malicious traffic dataset k variance;
[0135] S36. Sort the features according to the evaluation values, and select features with higher evaluation values to construct a malicious traffic feature database;
[0136] S37. Store the filtered feature data and use it for training a malicious traffic detection model.
[0137] The present invention comprehensively extracts the spatial structure information of malicious traffic by calculating local binary pattern features, global Gist features, and high-order local automatic correction features, and integrates multiple features to construct a comprehensive feature vector, thereby improving the ability to distinguish traffic data. By calculating the feature variance of normal traffic and malicious traffic, evaluating the importance of features, and screening high-discrimination features to construct a malicious traffic feature database, the model can learn traffic patterns more accurately. This method effectively reduces redundant features, improves model training efficiency and detection accuracy, makes malicious traffic detection more efficient and robust, and improves adaptability to new variants of malicious traffic.
[0138] In this embodiment, the S4 specifically includes:
[0139] S41. Label the normal traffic data and the malicious traffic data to construct a normal traffic dataset and a malicious traffic dataset, wherein the normal traffic dataset includes multiple normal traffic samples, and the malicious traffic dataset includes multiple malicious traffic samples, and assign a corresponding category label to each sample;
[0140] S42. Use the space-time capsule transformer to extract traffic features. First, the grayscale image of the input traffic sample is transformed through the capsule network to obtain the spatial feature representation of the traffic data. Then, the space-time transformer is used to model the traffic time series information. The spatial and time series features of the traffic sample are combined to generate the final traffic feature representation.
[0141] S43, feature mapping is performed on the extracted traffic features and input into the classification network. The classification network calculates the classification output result based on the extracted spatial features and temporal features, determines whether the traffic sample is malicious traffic, and outputs the detection result;
[0142] S44. Calculate the classification error. Let the true category label be y and the output probability of the classification network be The loss function is calculated using cross entropy loss:
[0143]
[0144] Among them, L loss function, y represents the true category label of the sample, Represents the sample category probability calculated by the classification network;
[0145] S45. Optimize the model parameters based on the calculated loss function, set the learning rate to α, and use the gradient descent algorithm to update the model parameters:
[0146]
[0147] Among them, W′ is the updated model parameter, W is the model parameter, is the gradient of the loss function with respect to the weight, and the learning rate α is used to control the update step size;
[0148] S46: Test the trained model to evaluate the detection accuracy of the model. If the detection accuracy reaches a set threshold, store the trained model parameters. Otherwise, continue training until the model meets the set accuracy requirements.
[0149] S47. Store the trained model parameters and use them for real-time malicious traffic detection. Use the stored model parameters to detect real-time traffic data, identify whether there is malicious traffic, and output the detection results.
[0150] The present invention constructs a training data set containing normal traffic and malicious traffic, uses the space-time capsule transformer to extract features, combines the spatial feature modeling capabilities of the capsule network and the temporal modeling capabilities of the transformer, generates an efficient traffic feature representation, and improves the detection capability of malicious traffic. In the classification stage, the classification network uses the extracted spatial features and temporal features for classification, and optimizes the model parameters by calculating the cross entropy loss. The gradient descent algorithm is used to iteratively update the weights to improve the model convergence speed and classification accuracy. During the training process, the model is evaluated by setting a detection accuracy threshold to ensure that the model stores parameters after achieving optimal performance for real-time malicious traffic detection. By fusing spatial features and temporal features, the present invention enables the model to accurately identify variant malicious traffic, while optimizing the training process, improving the generalization capability and detection efficiency of the model, and providing an efficient and scalable malicious traffic detection solution for the security of the Internet of Things.
[0151] In this embodiment, the S42 specifically includes:
[0152] S421. Perform feature mapping on the grayscale image of the traffic sample. Let the input grayscale image matrix be X. Obtain the initial feature map through multi-layer convolution transformation:
[0153] U i =σ(W i *X+b i );
[0154] Among them, U i is the initial feature map, W i is the convolution kernel weight, * represents the convolution operation, b i is the bias term, σ(·) is the nonlinear activation function;
[0155] S422, build a capsule network to represent the features at a high level, and set the initial capsule feature to be V i , use affine transformation to calculate the candidate capsule activation vector:
[0156]
[0157] Among them, W ij is the affine transformation matrix, b ij is the bias term, is the candidate capsule vector, U i is the input feature map;
[0158] S423, use the dynamic routing mechanism to calculate the coupling coefficient between capsules, and set the initial weight from capsule i to capsule j to be c ij , the coupling distribution is calculated by softmax normalization:
[0159]
[0160] Where t is the number of routing iterations, b ij is the coupling weight from capsule i to capsule j, b ik is the coupling weight from capsule i to capsule k, V j is the target capsule vector, is the coupling weight calculated for the previous round t, is the coupling weight after iteration t+1 rounds, is the candidate affine transformation vector of input capsule i;
[0161] S424. Perform nonlinear normalization on the updated capsule vector and calculate the final capsule output:
[0162]
[0163] Among them, s j is the capsule input after weighted summation, v j For the final capsule output, the squash function is used to ensure the shrinkage and stability of the vector norm;
[0164] S425, based on the spatiotemporal Transformer, temporal feature modeling is performed, and the input is the capsule vector sequence V at time step t. t , calculate the self-attention weight:
[0165]
[0166] Among them, Q t is the query matrix, K t is the bond matrix, d k is the feature dimension, A t is the feature representation of spatiotemporal attention calculation, and softmax(·) is the attention score normalization function;
[0167] S426, use the multi-head attention mechanism to optimize the temporal features, and set the multi-head attention output M t The calculation is as follows:
[0168]
[0169] Among them, H is the number of attention heads, is the projection weight of the query matrix, is the projection weight of the bond matrix, is the projection weight of the value matrix, α h is the attention coefficient, M t is the temporal feature calculated by multi-head attention, is the input key matrix, V t is the input value matrix;
[0170] S427, fusion capsule network extracted spatial features V j and the temporal features M calculated by the spatiotemporal Transformer t , calculate the final fusion features:
[0171] Z=LayerNorm(V j +γM t );
[0172] Among them, Z is the fused feature vector, V j is the spatial feature extracted by the capsule network, M t is the temporal feature calculated by Transformer, γ is the trainable weight parameter, and LayerNorm performs feature normalization;
[0173] S428. Perform high-dimensional feature transformation on the fused feature vector Z and calculate the final feature using nonlinear mapping:
[0174] F=σ(W f Z+b f )+λtanh(W g Z+b g );
[0175] Among them, F is the high-dimensional feature after the final conversion, W f is the linear transformation matrix, W g is the nonlinear transformation matrix, b f and b g is the corresponding bias term, σ(·) is the sigmoid activation function, and tanh(·) is the hyperbolic tangent activation function;
[0176] S429, input feature F for final classification, and use Softmax to calculate the classification probability:
[0177]
[0178] Among them, P is the probability distribution after classification, and W T is the weight matrix of the classification layer, b is the bias term of the classification layer, and c is the category index.
[0179] The present invention constructs a space-time capsule transformer to extract features from grayscale images of traffic samples, combining spatial features and time series information to achieve accurate malicious traffic detection. First, a multi-layer convolution transform is used to obtain the initial feature map, and a capsule network is used for high-order feature representation. A dynamic routing mechanism is used to optimize information transfer between capsules, ensuring that features can be adaptively adjusted and improving the model's ability to perceive variant malicious traffic. Subsequently, the space-time transformer is used to calculate temporal features, and a self-attention mechanism is used to capture the pattern of traffic changes over time. Multi-head attention is used to optimize temporal information, enabling the model to more accurately characterize complex traffic behavior. Finally, the spatial features of the capsule network and the temporal features calculated by the transformer are integrated, and nonlinear mapping is used for high-dimensional feature transformation. Softmax is used to calculate classification probabilities and determine traffic categories. The present invention improves the accuracy and robustness of malicious traffic detection by deeply integrating spatial and temporal information, while optimizing the computational structure, enabling the model to have higher computational efficiency and real-time performance while ensuring detection results.
[0180] In this embodiment, the S5 specifically includes:
[0181] S51. Collect real-time traffic data, divide it according to the set time window, obtain traffic data segments, and perform normalization processing;
[0182] S52. Use the space-time capsule transformer to extract spatial and temporal features, input the traffic data into the capsule network to calculate the spatial features, and perform time series modeling through the space-time capsule transformer to generate a fusion feature vector;
[0183] S53. Calculate the anomaly score. Let the mean eigenvector of the normal traffic sample be Calculate the anomaly score S:
[0184]
[0185] Among them, S is the anomaly score, Z is the fusion feature vector of the traffic sample to be detected, is the characteristic mean of normal traffic samples;
[0186] S54. Set an anomaly score threshold T and determine the traffic category based on the anomaly score S. The determination rules are as follows:
[0187]
[0188] S55. Output the traffic category detection result. If the traffic is determined to be malicious traffic, send an alarm signal to the command and control server, and execute blocking, isolation or adjust the security policy.
[0189] This invention achieves efficient malicious traffic detection in the Internet of Things (IoT) through real-time traffic data collection, feature extraction, anomaly score calculation, and alarm response. First, real-time traffic data is divided into predefined time windows and normalized to ensure consistent input data formats. Subsequently, a space-time capsule transformer is used to extract spatial and temporal features. The spatial features of the traffic data are calculated using a capsule network, and combined with time series modeling, a fused feature vector is generated to enhance the model's ability to identify malicious traffic. Anomaly scores are calculated based on the mean feature vector of normal traffic samples to measure the degree of deviation from the detected traffic, thereby distinguishing normal from malicious traffic. Setting anomaly score thresholds allows for precise traffic classification, ensuring robustness and generalization of detection. If a detection result identifies malicious traffic, the system immediately sends an alarm to the command and control server and implements appropriate security policies, such as blocking connections, isolating devices, or adjusting security policies, to enhance the security of the IoT environment. This invention, through its efficient real-time detection mechanism, improves the accuracy and response speed of malicious traffic identification, enabling the system to promptly detect and respond to potential security threats.
[0190] An IoT malicious traffic detection device, comprising:
[0191] The data acquisition module is used to collect network traffic data of IoT devices, convert the traffic data format, and extract traffic feature data;
[0192] The real-time detection module is used to input the real-time collected traffic data into the trained malicious traffic detection model and analyze the traffic data to detect whether there is malicious traffic;
[0193] A result sending module, used to send the detection result to the command and control server when malicious traffic is detected;
[0194] The alarm and response module is used to receive alarm signals returned by the command and control server and send response instructions to the attacked IoT devices to perform operations such as blocking connections, isolating devices, or adjusting security policies.
[0195] A computer-readable storage medium, characterized in that a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the processor is enabled to execute the method for detecting malicious traffic in the Internet of Things according to any one of claims 1 to 6.
[0196] Example 1:
[0197] In order to verify the feasibility of the present invention in detecting malicious traffic in the Internet of Things, the present invention was deployed on the Internet of Things security protection platform of a smart city management system to improve the accuracy and real-time performance of malicious traffic detection, and reduce the false alarm rate and detection delay. This smart city platform manages a large number of Internet of Things devices, including smart cameras, smart street lights, environmental sensors, intelligent traffic signal control systems, etc. Its network environment is complex and the types of devices are diverse. It faces the threat of malicious behaviors such as large-scale DDoS attacks, port scanning, malware propagation, and traffic spoofing. However, traditional rule-matching-based detection methods have high false alarm rates and long detection delays when facing new variant attacks, making it difficult to meet the high security requirements of the Internet of Things environment.
[0198] In this embodiment, the malicious traffic detection method of the present invention is integrated into the security system of the smart city management platform to improve detection accuracy and response speed. First, the system collects network traffic data of IoT devices in real time and divides it according to the set time window to form traffic data segments. Subsequently, the traffic data is converted into grayscale images using a text imaging method, and normalized, denoised and enhanced to remove redundant information and improve data stability. The system further uses a capsule network to extract the spatial features of the traffic, and combines it with the spatiotemporal Transformer for time series modeling, thereby capturing the spatial patterns and behavioral change trends of malicious traffic. The trained classification model can accurately identify different types of malicious traffic, calculate anomaly scores, determine traffic categories based on preset thresholds, and finally send detection results to the command and control server, and execute corresponding security policies, such as blocking connections, isolating devices or adjusting security policies.
[0199] To evaluate the detection effectiveness of the present invention, a comparative test was conducted using the proposed method and traditional intrusion detection systems on multiple key metrics, including detection accuracy, false alarm rate, detection latency, and abnormal behavior of IoT devices. The experiment lasted 30 days and covered 2,000 smart cameras, 5,000 smart streetlights, 3,000 smart traffic signal control devices, and 1,000 environmental sensors.
[0200] The test results show that under the traditional IDS solution, the false alarm rate is high, resulting in an abnormal offline rate of 10.5% for smart cameras, an abnormal restart rate of 8.7% for smart street lights, and a data loss rate of 7.2% for environmental sensors. In contrast, the false alarm rate of the method of the present invention is greatly reduced, with the abnormal offline rate of smart cameras reduced to 2.1%, the abnormal restart rate of smart street lights reduced to 1.5%, and the data loss rate of environmental sensors reduced to 1.3%. In addition, in terms of traffic detection speed, the present invention only takes 0.5 seconds to process 1,000 traffic samples on average, while the traditional IDS takes 2.7 seconds, and the overall detection efficiency is improved by 4.4 times. In terms of variant malicious traffic detection, the detection accuracy of the present invention remains at 92.5%, an increase of 16.9% compared to the traditional IDS.
[0201] The deployment of the present invention greatly improves the security of smart city IoT devices, reduces the impact of malicious traffic, and significantly reduces the false alarm rate and detection delay, thereby improving detection efficiency and system stability.
[0202] Table 1 Comparative data of IoT malicious traffic detection experiments
[0203]
[0204] As can be seen from the data in Table 1, the present invention significantly improves on several key detection metrics compared to traditional IDS. The improved detection accuracy demonstrates the model's ability to accurately identify malicious traffic, while the enhanced detection capabilities for variant traffic demonstrate the system's adaptability. The reduced false alarm rate significantly mitigates the impact of misjudgments, while the faster detection speed enables the system to rapidly respond to network attacks. Furthermore, the reduced anomaly rate of IoT devices further demonstrates the present invention's ability to effectively mitigate the impact of malicious traffic and improve the security and stability of smart city IoT environments.
[0205] In summary, the present invention uses the space-time capsule Transformer combined with the text imaging method to achieve efficient and accurate malicious traffic detection, enhance the system's real-time detection capability, reduce the false alarm rate, and improve detection accuracy. It enables IoT devices to respond quickly when malicious attacks occur, improves the overall security and stability of the IoT network, and provides a reliable security protection solution for scenarios such as smart cities and industrial Internet.
[0206] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A method for detecting malicious traffic in the Internet of Things, characterized in that: The steps include: S1. Collect malicious traffic data from IoT devices, classify them, and build a malicious traffic sample library. S2. Use text imaging method to convert malicious traffic data into grayscale image, and perform normalization, noise reduction and feature enhancement to obtain processed grayscale image; S3. Based on the processed grayscale image, extract local binary pattern features, global image features, and high-order local automatic correction features, screen highly relevant features, and establish a malicious traffic feature database; S4. Based on the malicious traffic feature database, the space-time capsule Transformer is used to train a malicious traffic detection model, extract the spatial features of the traffic, and combine the space-time attention mechanism to model the traffic time series. Classification training is performed to generate a malicious traffic classification model. S5. Collect real-time traffic data, convert it into grayscale images, extract features, use the trained classification model to extract spatial and temporal features, calculate anomaly scores, and determine traffic categories. S6. Send the detection results to the command and control server, trigger an alarm signal, and send response instructions to the attacked device, including blocking the connection, isolating the device, or adjusting the security policy.
2. The method for detecting malicious traffic in the Internet of Things according to claim 1, wherein: The S2 specifically includes: S21. Read the byte data of the target traffic data file, one byte at a time, and store them into a one-dimensional array in the order in which the data is read; S22. Determine the width of the grayscale image: Where N is the total number of bytes read, W is the width of the grayscale image, Indicates floor operation; S23. Convert the data dimension, convert the one-dimensional array into a two-dimensional array, and calculate the height of the grayscale image: Where H is the height of the grayscale image. If N is not divisible by W, zero padding is used to make the matrix dimension meet W×H. S24. Generate a grayscale image and calculate pixel values: G ij =b k ; Among them, G ij is the pixel value of the i-th row and j-th column in the grayscale image matrix, b k is the kth byte in the traffic data, with pixel values ranging from 0 to 255; S25, normalizing the generated grayscale image, adjusting the image size using a nearest neighbor interpolation algorithm, normalizing the image to a square, and ensuring that the normalized width and height are equal; S26. Store the normalized grayscale image and output it to the malicious traffic feature extraction module.
3. The method for detecting malicious traffic in the Internet of Things according to claim 1, wherein: The S3 specifically includes: S31. Extract the local binary pattern features of the grayscale image and calculate the local binary pattern values of the pixels: LBP ij =∑s(P k -G ij )·2 k ; Among them, LBP ij is the local binary pattern value of the pixel in the i-th row and j-th column in the grayscale image matrix, G ij is the gray value of the pixel, P k are the eight neighboring pixel values of the pixel point, k is the index of the neighboring pixel, ranging from 0 to 7, and s(x) is the binarization function, which is calculated as follows: S32, calculating the global Gist features of the grayscale image, filtering the image using a preset filter group, and performing histogram statistics on the filtered results to form a feature vector; S33, calculating high-order local automatic correction features, performing correlation calculation on pixels of the image at different offset distances, counting the autocorrelation values, and forming a feature vector; S34, fuses local binary pattern features, global Gist features, and high-order local automatic correction features to form a comprehensive feature vector of malicious traffic; S35. Calculate the feature variance of normal traffic and malicious traffic, assuming a feature F k The variances in the normal traffic dataset and the malicious traffic dataset are and Calculate the evaluation value of the feature: Among them, V k is feature F k The evaluation value of is the feature F in the normal traffic data set k The variance of is the feature F in the malicious traffic dataset k variance; S36. Sort the features according to the evaluation values, and select features with higher evaluation values to construct a malicious traffic feature database; S37. Store the filtered feature data and use it for training a malicious traffic detection model.
4. The method for detecting malicious traffic in the Internet of Things according to claim 1, wherein: The S4 specifically includes: S41. Label the normal traffic data and the malicious traffic data to construct a normal traffic dataset and a malicious traffic dataset, wherein the normal traffic dataset includes multiple normal traffic samples, and the malicious traffic dataset includes multiple malicious traffic samples, and assign a corresponding category label to each sample; S42. Use the space-time capsule transformer to extract traffic features. First, the grayscale image of the input traffic sample is transformed through the capsule network to obtain the spatial feature representation of the traffic data. Then, the space-time transformer is used to model the traffic time series information. The spatial and time series features of the traffic sample are combined to generate the final traffic feature representation. S43, feature mapping is performed on the extracted traffic features and input into the classification network. The classification network calculates the classification output result based on the extracted spatial features and temporal features, determines whether the traffic sample is malicious traffic, and outputs the detection result; S44. Calculate the classification error. Let the true category label be y and the output probability of the classification network be The loss function is calculated using cross entropy loss: Among them, L loss function, y represents the true category label of the sample, Represents the sample category probability calculated by the classification network; S45. Optimize the model parameters based on the calculated loss function, set the learning rate to α, and use the gradient descent algorithm to update the model parameters: Among them, W ′ is the updated model parameter, W is the model parameter, is the gradient of the loss function with respect to the weight, and the learning rate α is used to control the update step size; S46: Test the trained model to evaluate the detection accuracy of the model. If the detection accuracy reaches a set threshold, store the trained model parameters. Otherwise, continue training until the model meets the set accuracy requirements. S47. Store the trained model parameters and use them for real-time malicious traffic detection. Use the stored model parameters to detect real-time traffic data, identify whether there is malicious traffic, and output the detection results.
5. The method for detecting malicious traffic in the Internet of Things according to claim 4, wherein: The S42 specifically includes: S421. Perform feature mapping on the grayscale image of the traffic sample. Let the input grayscale image matrix be X. Obtain the initial feature map through multi-layer convolution transformation: U i =σ(W i *X+b i ); Among them, U i is the initial feature map, W i is the convolution kernel weight, * represents the convolution operation, b i is the bias term, σ(·) is the nonlinear activation function; S422, build a capsule network to represent the features at a high level, and set the initial capsule feature to be V i , use affine transformation to calculate the candidate capsule activation vector: Among them, W ij is the affine transformation matrix, b ij is the bias term, is the candidate capsule vector, U i is the input feature map; S423, use the dynamic routing mechanism to calculate the coupling coefficient between capsules, and set the initial weight from capsule i to capsule j to be c ij , the coupling distribution is calculated by softmax normalization: Where t is the number of routing iterations, b ij is the coupling weight from capsule i to capsule j, b ik is the coupling weight from capsule i to capsule k, V j is the target capsule vector, is the coupling weight calculated for the previous round t, is the coupling weight after iteration t+1 rounds, is the candidate affine transformation vector of input capsule i; S424. Perform nonlinear normalization on the updated capsule vector and calculate the final capsule output: Among them, s j is the capsule input after weighted summation, v j For the final capsule output, the squash function is used to ensure the shrinkage and stability of the vector norm; S425, based on the spatiotemporal Transformer, temporal feature modeling is performed, and the input is the capsule vector sequence V at time step t. t , calculate the self-attention weight: Among them, Q t is the query matrix, K t is the bond matrix, d k is the feature dimension, A t is the feature representation of spatiotemporal attention calculation, and softmax(·) is the attention score normalization function; S426, use the multi-head attention mechanism to optimize the temporal features, and set the multi-head attention output M t The calculation is as follows: Among them, H is the number of attention heads, is the projection weight of the query matrix, is the projection weight of the bond matrix, is the projection weight of the value matrix, α h is the attention coefficient, M t is the temporal feature calculated by multi-head attention, is the input key matrix, V t is the input value matrix; S427, fusion capsule network extracted spatial features V j and the temporal features M calculated by the spatiotemporal Transformer t , calculate the final fusion features: Z=LayerNorm(V j +γM t ); Among them, Z is the fused feature vector, V j is the spatial feature extracted by the capsule network, M t is the temporal feature calculated by Transformer, γ is the trainable weight parameter, and LayerNorm performs feature normalization; S428. Perform high-dimensional feature transformation on the fused feature vector Z and calculate the final feature using nonlinear mapping: F=σ(W f Z+b f )+λtanh(W g Z+b g ); Among them, F is the high-dimensional feature after the final conversion, W f is the linear transformation matrix, W g is the nonlinear transformation matrix, b f and b g is the corresponding bias term, σ(·) is the sigmoid activation function, and tanh(·) is the hyperbolic tangent activation function; S429, input feature F for final classification, and use Softmax to calculate the classification probability: Among them, P is the probability distribution after classification, and W T is the weight matrix of the classification layer, b is the bias term of the classification layer, and c is the category index.
6. The method for detecting malicious traffic in the Internet of Things according to claim 1, wherein: The S5 specifically includes: S51. Collect real-time traffic data, divide it according to the set time window, obtain traffic data segments, and perform normalization processing; S52. Use the space-time capsule transformer to extract spatial and temporal features, input the traffic data into the capsule network to calculate the spatial features, and perform time series modeling through the space-time capsule transformer to generate a fusion feature vector; S53. Calculate the anomaly score. Let the mean eigenvector of the normal traffic sample be Calculate the anomaly score S: Among them, S is the anomaly score, Z is the fusion feature vector of the traffic sample to be detected, is the characteristic mean of normal traffic samples; S54. Set an anomaly score threshold T and determine the traffic category based on the anomaly score S. The determination rules are as follows: S55. Output the traffic category detection result. If the traffic is determined to be malicious traffic, send an alarm signal to the command and control server, and execute blocking, isolation or adjust the security policy.
7. An IoT malicious traffic detection device, executing an IoT malicious traffic detection method according to any one of claims 1 to 6, characterized in that: include: The data acquisition module is used to collect network traffic data of IoT devices, convert the traffic data format, and extract traffic feature data; The real-time detection module is used to input the real-time collected traffic data into the trained malicious traffic detection model and analyze the traffic data to detect whether there is malicious traffic; A result sending module, used to send the detection result to the command and control server when malicious traffic is detected; The alarm and response module is used to receive alarm signals returned by the command and control server and send response instructions to the attacked IoT devices to perform operations such as blocking connections, isolating devices, or adjusting security policies.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is enabled to execute the method for detecting malicious traffic in the Internet of Things as described in any one of claims 1 to 6.
Citation Information
Cited By
Terminal equipment control method and system based on Internet of Things
CN120710796A
Iot-based terminal device control method and system
CN120710796B
Honeycomb linkage-oriented attack technology and tactical identification method
CN121309224A
A method for identifying attack techniques and tactics facing a honeycomb linkage
CN121309224B