Lightweight Internet of Things malicious traffic identification method based on cross-view knowledge distillation

By adopting the cross-view knowledge distillation method in the IoT environment, the multi-view teacher model and a single-view student model are constructed, and the problems of insufficient ability to identify complex patterns and high resource occupation in the existing technology are solved, and high-precision malicious traffic recognition and resource optimization are achieved, which is suitable for deployment in resource-constrained IoT environments.

CN119945795AActive Publication Date: 2025-05-06NANJING UNIV OF SCI & TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510190458.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-06
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

Existing IoT traffic recognition technologies are limited by the complexity of feature engineering and the learning ability of models when dealing with complex patterns, and the high demand for computing and storage resources limits their application in resource-constrained IoT environments, and cannot fully capture the behavioral patterns and characteristics of traffic.

Method used

A lightweight method based on cross-view knowledge distillation is adopted, and a multi-view teacher model and a single-view student model are constructed by comprehensively utilizing multi-view traffic data, and a cross-view knowledge distillation strategy is used for training and testing to achieve accurate identification of malicious traffic in the Internet of Things.

Benefits of technology

It improves the accuracy of identifying malicious traffic in the Internet of Things, reduces the complexity of the model and resource utilization, and makes it suitable for deployment in resource-constrained IoT environments, and can fully capture the behavior patterns of traffic and identify abnormal behaviors while ensuring the security of user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945795A_ABST
    Figure CN119945795A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight Internet of Things malicious traffic identification method based on cross-view knowledge distillation, and the method comprises the steps: firstly capturing and preprocessing Internet of Things network traffic, and then extracting features of a flow-level information view and a packet space-time sequence view; a multi-view teacher network is constructed, and the features of all views are deeply coded and fused; according to the method, a single-view lightweight student network is further constructed, and multi-view knowledge is implicitly learned in training through a cross-view knowledge distillation mechanism, so that high-precision malicious traffic classification is realized. The method has the characteristics of light model weight, high precision and low complexity, is particularly suitable for being deployed in an Internet of Things environment with limited resources, and has important significance for maintaining the network security of the Internet of Things.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and in particular to a lightweight Internet of Things malicious traffic identification method based on cross-view knowledge distillation. Background Art

[0002] With the rapid development of IoT technology, the number of IoT devices has increased exponentially, leading to an explosive increase in network traffic. This growth trend has not only promoted the digital transformation of various industries, but also brought unprecedented network security challenges. Cyber ​​attackers use a variety of technologies and means to target the inherent security flaws of IoT devices, such as limited computing and storage capabilities, and potential security vulnerabilities in firmware, launching increasingly complex and covert attacks, posing a severe test to the security of the IoT environment.

[0003] In the existing IoT traffic identification technology, the early simple analysis methods based on ports or loads have gradually become ineffective due to the popularization of dynamic port allocation and data encryption technology. Although traditional machine learning-based methods have improved certain recognition capabilities through automated feature learning, they are still limited by the complexity of feature engineering and the learning ability of the model when dealing with complex patterns.

[0004] In recent years, deep learning-based methods have shown great potential in the field of IoT traffic identification due to their excellent feature extraction and generalization capabilities. However, the high demand for computing and storage resources of these methods limits their practical application in resource-constrained IoT environments. In addition, existing technologies mostly analyze traffic from a single perspective, ignoring the high heterogeneity and diversity of IoT devices, resulting in the inability to fully capture the behavioral patterns and characteristics of traffic.

[0005] In view of this, there is an urgent need for a lightweight method that can adapt to resource-constrained environments and provide high-precision malicious traffic identification. Summary of the invention

[0006] In view of the shortcomings of the prior art, the present invention proposes a lightweight IoT malicious traffic identification method based on cross-view knowledge distillation. The method aims to improve the accuracy of recognition by comprehensively utilizing traffic data from multiple perspectives, while reducing the complexity and resource consumption of the model to adapt to the resource-constrained IoT environment.

[0007] The technical solution to achieve the purpose of the present invention is as follows: In the first aspect, the present invention provides a lightweight IoT malicious traffic identification method based on cross-view distillation, comprising the following steps:

[0008] Step 1: Capture the benign and malicious traffic flowing through the IoT platform, and divert the samples of the input traffic into session traffic based on the five-tuple information, and perform deduplication and grouping operations on it; the five-tuple is the five-tuple of source address, destination address, source port, destination port, and protocol;

[0009] Step 2: For the captured IoT traffic, construct a flow-level information view representation based on the flow duration and the average length of the data packets in the flow; construct a packet spatiotemporal sequence view representation based on the data packet length and the direction of the data packet;

[0010] Step 3: Based on the multi-view traffic representation obtained in step 2, a multi-view teacher model is constructed for training and saving;

[0011] Step 4: Input the spatiotemporal sequence view of the package and construct a single-view student model in combination with a deep separable convolutional network;

[0012] Step 5: Use the cross-view knowledge distillation strategy to train and test the single-view student model to accurately identify malicious traffic in the IoT.

[0013] In a second aspect, the present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the first aspect when executing the program.

[0014] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0015] In a fourth aspect, the present invention provides a computer program product, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.

[0016] Compared with the prior art, the present invention has the following significant advantages: 1) The cross-view knowledge distillation mechanism adopted by the present invention not only improves the recognition accuracy of malicious traffic in the Internet of Things, but also effectively reduces the complexity and resource occupation of the model by optimizing the model structure, so that the model can be efficiently deployed on Internet of Things devices or nodes with limited computing and storage resources; 2) By comprehensively utilizing the features of the two views of flow-level information and data packet sequence, the present invention can comprehensively capture the behavior patterns of Internet of Things traffic and effectively identify various abnormal behaviors including encrypted traffic. At the same time, user data is not touched during the whole process, ensuring the security of user privacy; 3) The packet sequence encoder designed by the present invention adopts the deep separable convolution technology, which replaces the traditional convolution operation and greatly reduces the computational burden of the model. In addition, by applying convolution kernels of different sizes, the encoder can accurately capture the spatiotemporal characteristics of data packets at different time scales, enhancing the model's ability to recognize traffic patterns.

[0017] The present invention is further described in detail below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.

[0019] Figure 1 This is a framework diagram of the lightweight IoT malicious traffic identification method based on cross-view distillation of the present invention. DETAILED DESCRIPTION

[0020] like Figure 1 As shown, the present invention proposes a lightweight IoT malicious traffic identification method based on cross-view distillation, which includes the following steps:

[0021] Step 1: Capture the benign and malicious traffic data flowing through the IoT platform; classify and sessionize the captured traffic data using the five-tuple information (source IP address, destination IP address, source port, destination port, and transport protocol); perform pre-processing operations such as deduplication and grouping on the sessionized traffic data to prepare for subsequent feature extraction;

[0022] Step 2: For the captured IoT traffic, construct a flow-level information view representation based on the flow duration, the average length of the data packets in the flow, etc.; construct a packet spatiotemporal sequence view representation based on the data packet length and the direction of the data packet;

[0023] Here, the extracted flow-level information view features include but are not limited to 31 features such as flow duration and number of packets, and these features are standardized. The specific features are shown in Table 1. For the packet spatiotemporal sequence view, the data packets are marked according to their sending direction. The traffic sent from the IoT device is marked as +1, and the received traffic is marked as -1;

[0024] Table 1. Flow-level information view extraction features and their descriptions

[0025] serial number Feature Description serial number Feature Description 01 Stream Duration 17 Total length of downlink data packets 02 Number of uplink data packets 18 Average downlink packet length 03 Number of downlink data packets 19 Standard deviation of downlink packet length 04 Total number of bidirectional packets 20 Total length of two-way data packets 05 Total average uplink inter-packet delay 21 Average bidirectional packet length 06 Total standard deviation of uplink inter-packet delay 22 Bidirectional packet length standard deviation 07 Total average delay between downlink packets 23 Upstream bytes per second 08 Total standard deviation of downlink inter-packet delay 24 Downlink bytes per second 09 Total mean of two-way packet delay 25 Bidirectional bytes per second 10 Total standard deviation of two-way inter-packet delay 26 Number of bytes in the header of the upstream data packet 11 Uplink packets per second 27 Number of bytes in the header of the downlink data packet 12 Downlink packets per second 28 Number of bytes in bidirectional packet header 13 Number of packets transmitted per second in both directions 29 Uplink data packet header byte ratio 14 Total length of uplink data packets 30 The proportion of downlink data packet header bytes 15 Average uplink packet length 31 Bidirectional data packet header byte ratio 16 Standard deviation of uplink packet length

[0026] Step 3: Use the multi-view features extracted in step 2 to build a multi-view teacher model.

[0027] Here, the multi-view teacher model consists of a multi-view encoding module, a multi-view fusion module, and a classifier; the multi-view encoding module is divided into a stream-level view encoder and a packet sequence view encoder; the stream-level view encoder is used to encode the view features of the stream-level information, and is composed of multiple linear layers combined with residual connections; the packet sequence view encoder is used to encode the view features of the packet spatiotemporal sequence, and is composed of multiple parallel residual convolution blocks of different scales, wherein each residual convolution block contains a 1D convolution layer, a BN layer, a ReLu activation function, and a residual connection; the convolution kernel size of each residual convolution block is different, namely 3, 5, and 7, to capture the spatiotemporal features at different scales;

[0028] Here, the multi-view fusion module is based on a two-stage fusion strategy. The first stage of fusion relies on a multi-head cross-attention mechanism to linearly map the features of different views, generate Query, Key and Value, and perform information fusion through the cross-attention mechanism. The second stage of fusion concatenates the features obtained in the first stage and then inputs them into a multi-layer perceptron for further feature fusion.

[0029] Here, the classifier of the teacher model consists of two fully connected linear layers; the fused multi-view features are input into the classifier and the multi-view teacher model is trained through the cross entropy loss function to optimize the classification performance;

[0030] Step 4: Input the spatiotemporal sequence view representation and construct a single-view student model in combination with a deep separable convolutional network;

[0031] Here, the single-view student model only uses the packet spatiotemporal sequence view as input; the single-view student model is mainly composed of a lightweight packet sequence encoder and a classifier; the lightweight packet sequence encoder is similar to the packet sequence encoder in the multi-view teacher model, and is composed of multiple parallel residual convolution blocks of different scales. Each residual convolution block contains a 1D convolution layer, a BN layer, a ReLU activation function and a residual connection. The 1D convolution block is replaced by a depthwise separable convolution from a standard convolution, and the 1D convolution layer uses a depthwise separable convolution to reduce the computational complexity and the number of parameters; the classifier is composed of two fully connected linear layers, and the output probability distribution is z s ;

[0032] Step 5: Use the cross-view knowledge distillation strategy to train and test the single-view student model to accurately identify malicious traffic in the IoT;

[0033] Here, cross-view knowledge distillation is divided into two parts: soft loss optimization and hard loss optimization. In this way, the student model can implicitly learn the multi-view knowledge of the teacher model and improve the recognition accuracy.

[0034] Soft loss optimization performs knowledge distillation at the response level and feature level respectively. For knowledge distillation at the response level, the output probability distribution of the teacher model and the student model is softened by the temperature parameter T, and the gap between the two is measured by the KL divergence. For knowledge distillation at the feature level, the feature gap is reduced by maximizing the similarity of the output features of the teacher model and the student model.

[0035] Specifically, for knowledge distillation at the response level, assume that the probability distributions of the classifier outputs of the teacher model and the student model are z t and z s , the output probability distribution is softened by introducing the temperature parameter T, that is:

[0036] q t =softmax(z t / T)

[0037] q s =softmax(z s / T)

[0038] Furthermore, the KL divergence is used to measure the soft probability distribution q of the student model. s and the soft probability distribution q of the teacher model t The gap between:

[0039] L KD =D KD (q t ||q s )

[0040] Among them, L KD represents the knowledge distillation loss at the response level, D KD represents the KL divergence function, q t is the soft probability distribution output by the teacher model, q s is the soft probability distribution output by the student model;

[0041] For knowledge distillation at the feature level, assume that the output feature of the multi-view fusion module of the teacher model is f mv , the output of the encoder of the student model is f p ; By maximizing f mv and f p , thereby reducing the gap between the two features, namely:

[0042]

[0043] Thus, the total soft loss function L is obtained soft as follows:

[0044] L soft =L KD +aLf

[0045] Among them, a is a hyperparameter used to balance L KD and L f These two loss terms;

[0046] Furthermore, hard loss optimization is achieved through the cross entropy loss function, namely:

[0047] L hard =L CE

[0048] Among them, L hard represents the hard loss function, L CE represents the cross entropy loss function;

[0049] Therefore, the total loss function L of the single-view student model is std It is expressed as follows:

[0050] L std =λL soft +(1-λ)L hard

[0051] Among them, λ is used to control the relative weight of soft loss and hard loss;

[0052] Furthermore, the trained single-view student model can be deployed in a resource-constrained IoT environment to identify malicious IoT traffic.

[0053] Through the cross-view knowledge distillation strategy, the present invention can not only achieve high-precision identification of malicious traffic in the Internet of Things, but also effectively reduce the complexity and size of the model. It is suitable for deployment in resource-constrained Internet of Things environments and is of great significance for maintaining the security of the Internet of Things network.

[0054] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A lightweight IoT malicious traffic identification method based on cross-view knowledge distillation, characterized in that: The following steps are involved: Step 1: Capture the benign and malicious traffic flowing through the IoT platform, split the samples of the input traffic into session traffic based on the five-tuple information, and perform deduplication and grouping operations on it; The five-tuple is a five-tuple of source address, destination address, source port, destination port, and protocol; Step 2: For the captured IoT traffic, construct a flow-level information view representation based on the flow duration and the average length of the data packets in the flow; construct a packet spatiotemporal sequence view representation based on the data packet length and the direction of the data packet; Step 3: Based on the multi-view traffic representation obtained in step 2, a multi-view teacher model is constructed for training and saving; Step 4: Input the spatiotemporal sequence view of the package and construct a single-view student model in combination with a deep separable convolutional network; Step 5: Use the cross-view knowledge distillation strategy to train and test the single-view student model to accurately identify malicious traffic in the IoT.

2. According to claim 1, the lightweight Internet of Things malicious traffic identification method based on cross-view knowledge distillation is characterized in that: The stream-level information view representation extracted in step 2 includes stream duration, number of uplink data packets, number of downlink data packets, total number of bidirectional data packets, total mean of uplink inter-packet delay, total standard deviation of uplink inter-packet delay, total mean of downlink inter-packet delay, total standard deviation of downlink inter-packet delay, total mean of bidirectional inter-packet delay, total standard deviation of bidirectional inter-packet delay, number of uplink packets per second, number of downlink packets per second, number of bidirectional packets per second, total length of uplink data packets, average length of uplink packet, standard deviation of uplink packet length, total length of downlink data packets, average length of downlink packet, and total length of downlink packet. Length standard deviation, total length of two-way data packets, average length of two-way packet length, standard deviation of two-way packet length, number of bytes transmitted per second in the uplink, number of bytes transmitted per second in the downlink, number of bytes transmitted per second in the two-way, number of bytes transmitted per second in the uplink data packet header, number of bytes transmitted per second in the downlink, number of bytes transmitted per second in the two-way, number of bytes in the uplink data packet header, number of bytes in the downlink data packet header, number of bytes in the two-way data packet header, proportion of bytes in the uplink data packet header, proportion of bytes in the downlink data packet header, proportion of bytes in the two-way data packet header, and these features are standardized; for the packet spatiotemporal sequence view, the direction of the data packet is marked as +1 or -1 according to whether it is sent or received from the IoT device.

3. The lightweight IoT malicious traffic identification method based on cross-view knowledge distillation according to claim 1 is characterized in that: The multi-view teacher model described in step 3 is composed of a multi-view encoding module, a multi-view fusion module and a classifier; The multi-view encoding module is divided into a stream-level view encoder and a packet sequence view encoder; the stream-level view encoder is used to encode the view features of the stream-level information, and is composed of multiple linear layers combined with residual connections; the packet sequence view encoder is used to encode the view features of the packet spatiotemporal sequence, and is composed of multiple parallel residual convolution blocks of different scales, where each residual convolution block contains a 1D convolution layer, a BN layer, a ReLu activation function, and a residual connection; the convolution kernel size of each residual convolution block is different, namely 3, 5, and 7, to capture the spatiotemporal features at different scales; The multi-view fusion module is based on a two-stage fusion strategy. The first stage of fusion is based on a multi-head cross-attention mechanism, which linearly maps the features of different views, generates queries, keys, and values, and performs information fusion through the cross-attention mechanism. The second stage of fusion concatenates the features obtained in the first stage and then inputs them into the multi-layer perceptron for feature fusion; The classifier consists of two fully connected linear layers; the fused multi-view features are input into the classifier to train the multi-view teacher model using the cross entropy loss function.

4. The lightweight IoT malicious traffic identification method based on cross-view knowledge distillation according to claim 1 is characterized in that: The single-view student model described in step 4 only uses the packet spatiotemporal sequence view as input; the single-view student model is mainly composed of a lightweight packet sequence encoder and a classifier; the lightweight packet sequence encoder is composed of multiple parallel residual convolution blocks of different scales, each residual convolution block contains a 1D convolution layer, a BN layer, a ReLU activation function and a residual connection, wherein the 1D convolution layer adopts a depth-wise separable convolution; the classifier is composed of two fully connected linear layers.

5. According to claim 4, the lightweight Internet of Things malicious traffic identification method based on cross-view knowledge distillation is characterized in that: The cross-view knowledge distillation described in step 5 includes two parts: soft loss optimization and hard loss optimization; Soft loss optimization performs knowledge distillation at the response level and feature level respectively. For knowledge distillation at the response level, the output probability distribution of the teacher model and the student model is softened by the temperature parameter T, and the gap between the two is measured by the KL divergence. For knowledge distillation at the feature level, the feature gap is reduced by maximizing the similarity of the output features of the teacher model and the student model. For knowledge distillation at the response level, assume that the probability distributions of the classifier outputs of the teacher model and the student model are z t and z s , the output probability distribution is softened by introducing the temperature parameter T, that is: q t =softmax(z t / T) q s =softmax(z s / T) The soft probability distribution q of the student model is measured by KL divergence s and the soft probability distribution q of the teacher model t The gap between: L KD =K KD (q t ||q s ) Among them, L KD represents the knowledge distillation loss at the response level, D KD represents the KL divergence function, q t is the soft probability distribution output by the teacher model, q s is the soft probability distribution output by the student model; For knowledge distillation at the feature level, assume that the output feature of the multi-view fusion module of the teacher model is f mv , the output of the encoder of the student model is f p ; By maximizing f mv and f p , thereby reducing the gap between the two features, namely: The total soft loss function is as follows: <h2 style=";text-align:left;direction:ltr">L<h2 style=";text-align:left;direction:ltr"> soft <h2 style=";text-align:left;direction:ltr"> =L<h2 style=";text-align:left;direction:ltr"> KD <h2 style=";text-align:left;direction:ltr"> +aL<h2 style=";text-align:left;direction:ltr"> f Among them, L soft represents the total soft loss function, a is a hyperparameter used to balance L KD and L f These two loss terms; Hard loss optimization is achieved through the cross entropy loss function, namely: L hard =L CE Among them, L hard represents the hard loss function, L CE represents the cross entropy loss function; Therefore, the total loss function L of the single-view student model is std It is expressed as follows: THE std =λL soft +(1-λ)L hard Where λ is the relative weight used to control soft loss and hard loss.

6. The lightweight IoT malicious traffic identification method based on cross-view knowledge distillation according to claim 1 is characterized in that: The trained single-view student model is deployed in the IoT environment to achieve real-time identification and classification of malicious IoT traffic.

7. A computer device comprising a memory, a processor and a computer program in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Pollen image classification method based on cross attention distillation Transformer

    CN113887610A

  • Lightweight Internet of Things malicious traffic identification method based on knowledge distillation space-time neural network

    CN116260642A

  • Track target point prediction method based on knowledge distillation

    CN116579423A

  • Lightweight radar target identification method based on spatial-temporal characteristic knowledge distillation

    CN118409289A

  • Joint task nasopharyngeal carcinoma in-situ recurrence prediction method based on attention knowledge migration

    CN118412100A