Bimodal Internet of Things traffic classification method based on fused image and statistical characteristics

By introducing dual-modal information interaction and feature refinement modules in IoT traffic classification, the problem of low accuracy of IoT traffic classification in the existing technology is solved, and higher classification accuracy and generalization capabilities are achieved.

CN120145217APending Publication Date: 2025-06-13CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510077218.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, the accuracy of IoT traffic classification is low, making it difficult to effectively integrate network traffic images and statistical features, resulting in poor classification performance when processing complex and changeable network traffic.

Method used

A dual-modal IoT traffic classification method based on fusion images and statistical features is proposed. By constructing a Bimodal TrafficNet model, the interactive module BCA-Module realizes multi-layer information interaction between network traffic statistical features and images, and optimizes feature extraction of traffic images through the feature refinement module PLIA-Module.

Benefits of technology

It significantly improves the model's perception of network complex traffic patterns, improves classification accuracy and generalization capabilities, and can more effectively handle diversified and complex IoT network traffic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145217A_ABST
    Figure CN120145217A_ABST
Patent Text Reader

Abstract

The invention discloses a bimodal Internet of Things traffic classification method based on fusion images and statistical characteristics, and the method is characterized in that the method comprises the following steps: S1, obtaining original Internet of Things traffic data, and carrying out the preprocessing of the data, and obtaining input data; and S2, inputting the input data into the constructed traffic classification model for training and classification to obtain a classification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network data processing, and particularly to a dual-modal Internet of Things (IoT) traffic classification method based on fused images and statistical features. Background Art

[0002] IoT (Internet of Things) traffic classification refers to analyzing network traffic to identify different types of communication activities to ensure high-quality and reliable service quality (QoS). However, with the rapid increase in IoT devices, network traffic has shown a high degree of diversity and complexity. This not only poses a huge challenge to network resource management but also constitutes a serious threat to network space security.

[0003] IoT is profoundly changing the way we interact with the digital world, and its applications have expanded from daily life in the home environment to complex large-scale systems such as industrial manufacturing. The IoT ecosystem consists of heterogeneous networks and billions of intelligent devices. Its rapid expansion and large-scale high-speed operation have not only promoted the digital transformation of society but also brought great hidden dangers to data protection and information security in IoT. In addition, the diversity of IoT application scenarios and the rapid evolution of communication protocols also pose very demanding requirements for the QoS (Quality of Service) of the network, including low latency, high stability, and fast and flexible service adaptation.

[0004] To address these challenges, accurate network traffic classification has become one of the core means for IoT system optimization. By identifying the types of network traffic, the system can formulate corresponding forwarding priorities based on the characteristics of network traffic, thereby optimizing the scheduling of data packets and resource allocation. For example, in industrial IoT (IIoT), data packets of real-time control signals need to be given priority for transmission; while ordinary sensor data can be appropriately delayed. Therefore, network traffic classification is not only an important prerequisite for ensuring the security of IoT systems but also a key path for improving resource management efficiency and QoS.

[0005] Traditional network traffic classification methods, such as those based on ports and Deep Packet Inspection (DPI), are ineffective in the face of the dynamics and diversity of network traffic. Although machine learning can improve classification performance to a certain extent through feature engineering, its reliance on shallow features makes it less efficient in processing large-scale and complex network traffic.

[0006] In recent years, the rise of deep learning technology has provided a disruptive solution to this problem. Deep learning models can automatically extract high-dimensional features from network traffic data in an end-to-end manner, significantly improving the accuracy of classification. However, most existing methods rely on a single input or a single model. Although this strategy shows good performance in certain scenarios, it is difficult to comprehensively handle complex and changing classification scenarios and network traffic behaviors. Currently, multi-modal deep learning has gradually become a research hotspot in classification. By integrating multi-source information, this technology can capture more comprehensive multi-dimensional features, thus significantly enhancing the robustness and generalization ability of classification models. Although multi-modal deep learning methods have achieved remarkable success in audiovisual speech recognition, emotion recognition, and health analysis, their efficient application in network traffic classification remains an urgent problem to be solved. Specifically, network traffic data usually contains two main information sources: statistical features (such as the number of data packets in a session) and traffic images (such as the visualization of a session). These two types of information are naturally complementary. Statistical features can reflect the global behavior patterns of network traffic, while image features can capture fine-grained local details. How to effectively fuse these two types of information and automatically learn their interaction relationships through deep learning models is the key to improving classification performance. Summary of the Invention

[0007] Aiming at the technical problem of low accuracy of Internet of Things traffic classification in the prior art, the present invention proposes a dual-modal Internet of Things traffic classification method based on the fusion of image and statistical features. By combining two-modal information of network traffic images and network traffic statistical features, the classification accuracy and generalization ability are improved.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] A dual-modal Internet of Things traffic classification method based on the fusion of image and statistical features, comprising the following steps:

[0010] S1: Obtain the original Internet of Things traffic data and perform preprocessing to obtain input data;

[0011] S2: Input the input data into the constructed traffic classification model for training and classification to obtain a classification result.

[0012] Preferably, the S1 includes:

[0013] S1-1: Split the obtained original Internet of Things traffic data into multiple sessions, and then perform cleaning processing to obtain the first data;

[0014] S1-2: Extract statistical features from the first data, then standardize the statistical features, and sort them from large to small based on the standard deviation of the statistical features to obtain a statistical feature input sequence;

[0015] S1-3: Convert the first data into a traffic image.

[0016] Preferably, in the S1-1, the first data acquisition method is:

[0017] First, analyze the traffic using the five-tuple information of the network traffic, and split the data into sessions; then, perform cleaning processing on the generated sessions. The cleaning processing includes deleting duplicate session files and removing empty files, so as to obtain the first data.

[0018] Preferably, in the S1-2, the normalization is:

[0019]

[0020] In formula (1), F S represents the normalized statistical feature, F is the original statistical feature; μ is the average value; σ is the standard deviation; ε is a small constant.

[0021] Preferably, in the S1-3, the conversion method of the traffic image is:

[0022] For each session file, trim it to a normalized 784 bytes, that is, if the packet length exceeds 784 bytes, truncate it; if the packet length is less than 784 bytes, pad it with 0x00; subsequently, the data is converted into a 28×28×1 grayscale image, where each pixel point corresponds to the byte value in the network traffic data, ranging from 0x00 to 0xFF.

[0023] Preferably, the S2 includes:

[0024] S2-1: Build a traffic classification model;

[0025] S2-2: Input the traffic image and statistical features into the traffic classification model, and divide them into the first patch and the second patch respectively;

[0026] S2-3: Extract the first context information from the first patch and the second context information from the second patch respectively;

[0027] S2-4: Exchange the first context information and the second context information to generate the first fusion output corresponding to the traffic image and the second fusion output corresponding to the statistical features;

[0028] S2-5: Refine the features of the first fusion output to obtain the first output result;

[0029] S2-6: Input the first output result into the first classification module to obtain the first classification result, and input the second fusion output into the second classification module to obtain the second classification result.

[0030] Preferably, in the S2-1, the traffic classification model includes a first processing module, a first branch module, a first classification module, a second processing module, a second branch module, a second classification module, and an interaction module;

[0031] The first processing module is used to process the traffic image and divide it into first patches; the first branch module is used to extract first context information from the first patches to obtain first patch tokens and combine them with the first CLS token to obtain a first fusion output; the first classification module is used to obtain a first classification result according to the first fusion output;

[0032] The second processing module is used to process the statistical features and divide them into second patches; the second branch module is used to extract second context information from the second patches to obtain second patch tokens and combine them with the second CLS token to obtain a second fusion output; the second classification module is used to obtain a second classification result according to the second context information;

[0033] The interaction module is used to interact the first context information and the second context information and continuously update the first fusion output and the second fusion output.

[0034] Preferably, in the S2-3, the extraction methods of the first context information and the second context information are as follows:

[0035] R i = T i-1 + MHSA(LN(T i-1 ))), T i = R i + FFN(LN(R i )) (2)

[0036] In formula (2), R i represents the intermediate result after adding a residual connection after the multi-head self-attention module, and T i represents the output of the i-th layer of the Transformer Block; T i-1 represents the output of the (i - 1)-th layer of the Transformer Block; MHSA(LN(T i-1 )) represents the multi-head self-attention mechanism function; FFN(LN(R i )) represents the feed-forward neural network function.

[0037] Preferably, in the S2-4, the traffic image first concatenates the corresponding CLS token with the patch token of the statistical feature, and the formula is:

[0038]

[0039] In formula (3), T'I Represents a concatenated token, Represents the CLS token corresponding to the traffic image; Represents the patch token of the statistical feature; p I (·) is a projection function for dimension alignment;

[0040] The information of the patch token corresponding to the traffic image is fused into the CLS token, and its mathematical representation is as follows:

[0041]

[0042] In formula (4), are learnable parameters, D and h are the embedding dimension and the number of heads respectively; Q represents the query; K represents the key; V represents the value; Represents the CLS token after the projection function for dimension alignment; T' I Represents the concatenated token; A represents the attention weight matrix; M represents the output matrix weighted by the attention mechanism; T represents the transpose matrix;

[0043] After receiving the information of the patch token in the statistical feature, it is then back-projected to the patch token of the traffic image and passed to itself Specifically, the output after the information fusion of the traffic image is defined as follows:

[0044]

[0045] In formula (5), Represents the intermediate result of the projection; g I (·) is the back-projection function; p I (·) is a projection function for dimension alignment; Represents the CLS token corresponding to the traffic image; Represents the patch token corresponding to the traffic image; Represents the first fusion output corresponding to the traffic image.

[0046] Preferably, in S2-5, the method for obtaining the first output result is:

[0047] First, input the first fusion output and divide it into two groups of sub-features in the embedding dimension where B is the batch size, b = 2B, d = D / 2, h = H / P, w = W / P, H is the height, W is the width, and D is the spatial dimension;

[0048] Then, perform pooling operations on T R along the height and width directions respectively:

[0049] T H = AvgPoolH (T R ), T W = Reshape(AvgPool W (T R )) (6)

[0050] In formula (6), and respectively represent the global information aggregating the feature map in the height and width dimensions; AvgPool H represents the pooling operation in the height direction; AvgPool W represents the pooling operation in the width direction; Reshape represents the activation function;

[0051] Subsequently, through concatenation and 1×1 convolution fusion, a unified feature representation is generated:

[0052] T F = f 1×1 (Concat(T H , T W )) (7)

[0053] In formula (7), Concat represents concatenation; represents the unified feature representation;

[0054] T F is divided into two vector-unified feature representations and and the weights T HI and T WI are generated through the activation function:

[0055] T HS , T WS = Split(T F ), T HI = Sigmoid(T HS ), T WI = Reshape(Sigmoid(T WS )) (8)

[0056] In formula (8), T HS represents the unified feature representation in the height direction; T WS represents the unified feature representation in the width direction; Split represents the splitting function; T HI represents the weight in the height direction; T WI represents the weight in the width direction; Sigmoid and Reshape both represent activation functions;

[0057] Finally, T HI , T WIThrough the sub-feature T R After element-wise multiplication and shaping, the position enhancement feature T is obtained EN :

[0058] T EN = Reshape(T R ⊙ T HI ⊙ T WI )(9)

[0059] In formula (9), T EN represents the position enhancement feature; Reshape represents the activation function; ⊙ represents element-wise multiplication;

[0060] The parallel operations of 3×3 convolution and 5×5 convolution are also introduced to capture diverse pattern information in the local area, obtaining T S :

[0061] T S = f 3×3 (T R ) + f 5×5 (T R )(10)

[0062] In formula (10), T S represents the parallel convolution feature; f 3×3 represents 3×3 convolution; f 5×5 represents 5×5 convolution; T R represents the sub-feature of the flow image;

[0063] Perform global average pooling on T EN and adjust the weight distribution through Softmax(·), obtaining

[0064] T N = Softmax(Reshape(AvgPool(T EN )))(11)

[0065] In formula (11), T N represents the global attention feature; AvgPool represents the global average pooling operation; Reshape, Softmax represent the activation functions;

[0066] Meanwhile, calculate the channel weights for T S to obtain the weight information of each channel

[0067] T C = Softmax(Reshape(AvgPool(T S )))(12)

[0068] In formula (12), T C represents the processed pattern information;

[0069] By multiplying the parallel convolution feature T S with the global attention feature T N the first spatial attention map is obtained Performing the position enhancement feature T EN and the channel weight feature T C results in the second attention map

[0070]

[0071] In formula (13), T MA represents the first spatial attention map; T MB represents the second spatial attention map; represents multiplication;

[0072] Adding T MA and T MB and passing through the Reshape operation and then the Sigmoid activation function to generate the final fused feature

[0073] T MSA = Reshape(Sigmoid(T MA + T MB )) (14)

[0074] In formula (14), T MSA represents the fused feature;

[0075] Multiplying T MSA element - by - element with the initial input sub - feature T R to obtain the final output result

[0076] T O = T R ⊙ T MSA (15)

[0077] In formula (15), T O represents the first output result corresponding to the network traffic image.

[0078] In summary, due to the adoption of the above - mentioned technical solutions, compared with the prior art, the present invention has at least the following beneficial effects:

[0079] This paper innovatively introduces an interactive module, namely the BCA-Module, to achieve multi-level information interaction between network traffic statistical features and network traffic images, significantly enhancing the model's perception ability of complex network traffic patterns. In addition, the application of the feature refinement module further optimizes the feature extraction effect of traffic images. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 It is a schematic flow diagram of a dual-modal Internet of Things traffic classification method based on fused images and statistical features according to an exemplary embodiment of the present invention.

[0081] Figure 2 It is a schematic diagram of a traffic classification model according to an exemplary embodiment of the present invention.

[0082] Figure 3 It is a schematic diagram comparing the classification accuracies on different data sets according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] The present invention will be further described in detail below in conjunction with embodiments and specific implementation manners. However, it should not be understood that the scope of the above subject matter of the present invention is limited to the following embodiments. Any technology implemented based on the content of the present invention belongs to the scope of the present invention.

[0084] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present invention.

[0085] This paper proposes a strong interaction classification model that fuses traffic images and statistical features - BimodalTrafficNet. Specifically, this model deeply explores the potential connection between the two by strengthening the hierarchical fusion of statistical features and traffic images to make full use of their advantages in different dimensions. In addition, to further enhance the model's understanding of traffic image information, the PLIA-Module is proposed. To verify the effectiveness and robustness of Bimodal TrafficNet in different scenarios and granularity classification tasks, four publicly available data sets are selected for experiments to ensure the wide applicability and persuasiveness of the results.

[0086] As Figure 1 shown, the present invention provides a dual-modal Internet of Things traffic classification method based on fused images and statistical features, which specifically includes the following steps:

[0087] S1: Obtain the original Internet of Things traffic data and perform preprocessing to obtain the input data.

[0088] In this embodiment, to improve the efficiency of model training and the accuracy of classification results, before converting the data set into an input acceptable to Bimodal TrafficNet, it is necessary to perform systematic preprocessing on the original network traffic data.

[0089] S1-1: Split the obtained original Internet of Things traffic data into multiple sessions, and then perform cleaning processing to obtain the first data.

[0090] In this embodiment, each network traffic object is defined as a session, also known as a bidirectional flow. Specifically, a flow refers to a sequence of data packets with the same five-tuple (transport layer protocol, source IP address, destination IP address, source port, destination port) identification, and the main difference between a session and a flow lies in whether the directional characteristics of the data packets are considered. A session integrates all bidirectional traffic between the source IP and the destination IP into a unit, so it contains richer interaction information than a unidirectional flow. In addition, the research by Wang et al. shows that session features contain more interaction information, which helps to improve the accuracy of network traffic classification. Therefore, in order to more accurately describe the communication characteristics between devices and capture the information of bidirectional interaction, this solution selects sessions as the unit for traffic preprocessing.

[0091] In this embodiment, for the original Internet of Things traffic data stored in the pcap file format, first, the traffic is parsed using the five-tuple information of the network traffic, and the data is split into sessions (which belongs to the existing method, and code is written using Python software, and SplitCap.exe is called to split the traffic into sessions); then, the generated sessions are cleaned, including deleting duplicate session files and removing empty files, to ensure that the samples have effective network traffic feature information.

[0092] S1-2: Extract statistical features from the first data, then standardize the statistical features, and sort them in descending order based on the standard deviation of the statistical features to obtain the input sequence.

[0093] In this embodiment, the statistical features of network traffic play a key role in feature representation. During the process of standardizing the size of traffic images, pruning may cause partial information loss, while statistical features can supplement this missing information from a global perspective. Specifically, there are 26 statistical features, including Num pkts (the number of data packets in a session), Duration window flow (the duration from the first data packet to the last data packet in a session), StDev delta time (the standard deviation of the time intervals between data packet arrivals in a session), etc., which can characterize traffic from a global perspective and enhance the model's understanding of communication behaviors.

[0094] The extraction of statistical features belongs to existing methods. Code can be written in Python software to calculate 26 statistical features. The reference paper is: Wang M, Zheng K, Luo D, et al. An encrypted traffic classification framework based on convolutional neural networks and stacked autoencoders[C] / / 2020 IEEE 6th International conference on computer and communications(ICCC). IEEE, 2020:634-641.

[0095] In this embodiment, after the statistical features of network traffic are extracted, the statistical features are standardized based on the mean and standard deviation of the statistical features (here, the mean and standard deviation are two commonly used indicators in statistics. The feature mean is the sum of all statistical features divided by the number of data, and the standard deviation is obtained based on the mean) to eliminate the dimensional differences between different features. The standardization formula is:

[0096]

[0097] In formula (1), F S represents the standardized statistical feature, F is the original statistical feature value; μ is the mean; σ is the standard deviation; ε is a small constant used to prevent division by zero.

[0098] The standard deviation of statistical features, as a key indicator to measure the importance in data distribution, can reflect the variation range of features in the data. Generally speaking, features with larger standard deviations usually contain more information and make more significant contributions to the classification task. Therefore, after the statistical features are standardized, the statistical features sorted based on the standard deviation size are used as the input sequence of the model to highlight the expression ability of key features in network traffic.

[0099] S1-3: Convert the first data into a traffic image.

[0100] To ensure data privacy and eliminate biases that may be caused by redundant information such as user identities, the IP address and MAC address information in the network traffic session are deleted for anonymization. Specifically, the corresponding address fields are replaced with 0x00 while preserving the structural integrity of the traffic data.

[0101] For each session file, it is trimmed to a normalized 784 bytes. That is, if the packet length exceeds 784 bytes, it is truncated; if the packet length is less than 784 bytes, it is padded with 0x00. Subsequently, the data is converted into a grayscale image of 28×28×1, where each pixel corresponds to the byte value in the network traffic data, ranging from 0x00 to 0xFF.

[0102] In this embodiment, the input data includes traffic images and statistical feature input sequences.

[0103] S2: Input the input data into the constructed traffic classification model for training and classification to obtain classification results.

[0104] In this embodiment, existing research mainly focuses on using original network traffic sessions for classification. Such methods usually require unifying the network traffic size, which will destroy the overall structural information of the network traffic. The statistical features of the traffic can make up for the information loss caused by homogenization. Based on this, this paper proposes a traffic classification model Bimodal TrafficNet that combines bimodal information to achieve network traffic classification.

[0105] As Figure 2 shown, the traffic classification model includes a first processing module, a first branch module, a first classification module, a second processing module, a second branch module, a second classification module, and an interaction module; the first branch module, the second branch module, and the interaction module are combined to obtain a double-branch unit TrafficNet Block.

[0106] The output end of the first processing module is connected to the first input end of the double-branch unit (i.e., connected to the first branch module), the first output end of the double-branch unit (the output end of the first branch module) is connected to the input end of the first classification module, and the output end of the first classification module outputs the first classification result; the second processing module is connected to the second input end of the double-branch unit (i.e., connected to the input end of the second branch module), the second output end of the double-branch unit (the output end of the second branch module) is connected to the input end of the second classification module, and the output end of the second classification module outputs the second classification result; among them, the first branch module is connected to the second branch module through the interaction module.

[0107] The first processing module is used to process the traffic image and divide it into first patches; the first branch module is TrafficNet Block, which is used to extract first context information from the first patches to obtain first patch tokens and combine them with the first CLS token to obtain a first fusion output; the first classification module is an independent multi-layer perceptron (MLP), which is used to obtain a first classification result according to the first fusion output;

[0108] The second processing module is used to process the statistical features and divide them into second patches; the second branch module is TrafficNet Block, which is used to extract second context information from the second patches to obtain second patch tokens and combine them with the second CLS token to obtain a second fusion output; the second classification module is an independent multi-layer perceptron (MLP), which is used to obtain a second classification result according to the second fusion output;

[0109] The interaction module is used to interact the first context information and the second context information, and continuously update and generate the final representation after information fusion, that is, the first fusion output and the second fusion output.

[0110] S2-1: To more effectively extract the local and global features of the IoT network traffic data, the traffic image and statistical features of the network traffic are respectively divided into patches of a fixed size through the PatchEmbedding module. Each patch is mapped to a high-dimensional embedding space to capture a richer information representation.

[0111] Specifically, for the network traffic image x i , with a size of H×W×1, where H is the height, W is the width, and the number of channels is 1. The traffic image x i is divided into non-overlapping first patches of size P×P. Each patch is mapped to a D-dimensional embedding space through a simple convolutional layer and a flattening operation:

[0112]

[0113] In formula (2), represents a real number; H is the height, W is the width, and the number of channels is 1; P represents Patch; N = HW / P 2 , which is the total number of patches; D represents the spatial dimension.

[0114] For the processing of the statistical feature input sequence, a simple and efficient method is adopted. Specifically, each statistical feature is regarded as a patch to obtain a second patch, and the size of each patch in the second patch is 1×1, and it is directly processed (input into the model according to the size of 1×1, and then processed according to the design of the model). Since the dimension of the statistical feature is much smaller than the pixel information of the traffic image, and it mainly carries global statistical information. Therefore, the embedding dimension of the mapping is half of the first patch to balance the weights of the traffic image and the statistical feature in the bimodal fusion.

[0115] In this embodiment, after being divided into patches, in the subsequent transformer, the patches are input one by one as tokens, so they are a bunch of patch tokens; then a CLS token is added in front of these patch tokens (the transformer processing is designed like this), and both branches are processed in this way.

[0116] In addition, considering the high dependence of visual applications on position information, the CLS token is first added in front of the patch tokens for performing the final classification task. In this way, the number of tokens becomes 1+N. Then a position embedding is added to each token to maintain the relative spatial position information. The linear embedding process can be formulated as follows:

[0117] T=[T cls ||T patch +T pos (3)

[0118] In formula (3), T represents the input data; respectively represent the CLS token and the patch tokens, while represents the position embedding of the tokens. These tokens are then passed through the TrafficNetBlock, and finally the network traffic classification task is completed using the CLS token.

[0119] S2-2: Extract the first context information from the first patch and the second context information from the second patch respectively.

[0120] In this embodiment, the first context information is extracted from the first patch through the first branch module to obtain the first patch tokens, and the first fusion output is obtained by combining the first CLS token; the second context information is extracted from the second patch through the second branch module to obtain the second patch tokens, and the second fusion output is obtained by combining the second CLS token; both the first branch module and the second branch module are Transformer Blocks.

[0121] Each Transformer Block consists of a multi-head self-attention mechanism (MHSA) and a feed-forward neural network (FFN). Layer normalization (LN) is applied before each block, and residual connections are applied after each block to alleviate the vanishing gradient problem and accelerate training convergence.

[0122] Specifically, in MHSA, the input is decomposed into queries (Q), keys (K), and values (V), and different representations are obtained through linear transformations. These representations capture the long-range dependencies of the input sequence through the attention mechanism, helping the model understand global context information. Subsequently, the output passing through the MHSA module is fed into the FFN, where the context features from the self-attention layer are further processed, enabling the model to generate more context-aware prediction results. The Transformer Block is formulated as follows:

[0123] R i = T i-1 + MHSA(LN(T i-1 ))), T i = R i + FFN(LN(R i )) (4)

[0124] In formula (4), R i represents the intermediate result after the multi-head self-attention module with residual connections, T i represents the output of the i-th layer of the Transformer Block; T i-1 represents the output of the (i - 1)-th layer of the Transformer Block; MHSA(LN(T i-1 )) represents the multi-head self-attention mechanism function; FFN(LN(R i )) represents the feed-forward neural network function.

[0125] S2-3: In the dual-branch unit structure of the TrafficNet Block, the effective fusion of dual-modal information (the first patch and the second patch) is the key to improving the model performance. To this end, an interaction module BCA-Module is designed, which promotes the in-depth information exchange between the traffic image (I-Branch) and the statistical features (F-Branch) by introducing a cross-modal interaction mechanism.

[0126] The core idea of ​​the interaction module BCA-Module is to use CLS tokens as information agents for each branch, and to participate in cross-modal information interaction with its abstract expression. Specifically, the CLS token of each branch represents the comprehensive information of all patch tokens of the branch. The CLS token is used to interact with the patch tokens of another branch. After cross-modal interaction, each CLS token reversely projects the fused information back to its branch, so that its patch token can also benefit indirectly, thereby enriching the feature expression.

[0127] In this embodiment, the traffic image first converts the corresponding first CLS token Second patch token with statistical features To concatenate, the formula is:

[0128]

[0129] In formula (5), T' I Represents a concatenated token, Indicates the first CLS token corresponding to the traffic image; The second patch token representing the statistical features; p I (·) is a dimension-aligned projection function used to map tokens from different sources into a unified feature space.

[0130] Then, similar to self-attention, this module and T' I Multiple heads are also used to capture contextual information and enhance the representation ability of the model. Specifically, It is the only query because the information of the patch token corresponding to the traffic image has been integrated into the CLS token. Its mathematical representation is as follows:

[0131]

[0132] In formula (6), are learnable parameters, D and h are embedding dimension and number of heads respectively; Q represents query; K represents key; V represents value; Represents the CLS token T' after the dimension-aligned projection function I represents the concatenated token; A represents the attention weight matrix; M represents the output matrix after the attention mechanism weights (that is, the final weighted result); T represents the transposed matrix;

[0133] After receiving the information of the patch token in the statistical feature, it is then back-projected back to the flow image to pass its own patch token Specifically, the output after fusion of traffic image information is The definition is as follows:

[0134]

[0135] In formula (7), represents the projection intermediate result; g I (·) is the back-projection function; p I (·) is the dimension-aligned projection function; represents the CLS token corresponding to the traffic image; represents the patch token corresponding to the traffic image; represents the first fusion output corresponding to the traffic image.

[0136] Similarly, the statistical features will also receive information from the patch tokens in the traffic image. Then, further patch token information interaction is performed at the next Transformer Block to obtain the second fusion output.

[0137] Through this hierarchical cross-modal interaction, the CLS tokens of the two branches can learn richer context information from the other branch and reverse-pass this information to the patch tokens of their respective branches. This interaction mechanism realizes the information sharing between the traffic image and the statistical features at multiple scales, retaining the uniqueness of the modality while enhancing the synergy between the two.

[0138] S2-4: Refine the features of the first fusion output corresponding to the traffic image to obtain the first output result.

[0139] The network traffic image is converted from the original bytes and has highly complex semantic features. Although the traditional attention mechanism can capture the abstract features of the image, it often ignores the pixel-level detail information, and this lack of detail is particularly obvious when analyzing byte-based traffic images. Therefore, this paper proposes a feature refinement module, PLIA-Module, which is located after the interaction module BCA-Module in the image branch. By refining the image feature expression, PLIA-Module can enhance the ability of the traffic image to extract pixel-level semantic information.

[0140] First, the input feature (i.e., the first fusion output ) and is divided into two groups of sub-features in the embedding dimension where B is the batch size, b = 2B, d = D / 2, h = H / P, w = W / P, H is the height, W is the width, and D is the spatial dimension, to reduce the computational amount of a single group and improve the aggregation ability of the local receptive field. Then, perform pooling operations on T R along the height and width directions respectively:

[0141] T H = AvgPool H (T R),T W = Reshape(AvgPool W (T R )) (8)

[0142] In formula (8), and respectively represent the global information aggregating the feature map in the height and width dimensions; AvgPool H represents the pooling operation in the height direction; AvgPool W represents the pooling operation in the width direction; Reshape represents the activation function.

[0143] Subsequently, through concatenation and 1×1 convolution fusion, a unified feature representation is generated.

[0144] T F = f 1×1 (Concat(T H ,T W )) (9)

[0145] In formula (9), Concat represents concatenation; represents the unified feature representation. T F is divided into two vector-unified feature representations and and the weights T HI and T WI are generated through the activation function to shuffle the information in different spaces on the channels, thereby enhancing the information interaction.

[0146] T HS ,T WS = Split(T F ),T HI = Sigmoid(T HS ),T WI = Reshape(Sigmoid(T WS ))(10)

[0147] In formula (10), T HS represents the unified feature representation in the height direction; T WS represents the unified feature representation in the width direction; Split represents the splitting function, using the function torch.split to split T F into two sub-vectors, respectively containing parts of size h and w; T HI represents the weight in the height direction; T WI represents the weight in the width direction; Sigmoid and Reshape both represent the activation function.

[0148] Finally, THI , T WI Through element-wise multiplication with the sub-feature T R and integer shaping, the position-enhanced feature T is obtained EN to strengthen position details and enhance the semantic representation ability of each pixel.

[0149] T EN = Reshape(T R ⊙ T HI ⊙ T WI )(11)

[0150] In formula (11), T EN represents the position-enhanced feature; Reshape represents the activation function; ⊙ represents element-wise multiplication.

[0151] In this embodiment, the feature refinement module also introduces parallel operations of 3×3 convolution and 5×5 convolution to capture diverse pattern information in the local area, obtaining T S .

[0152] T S = f 3×3 (T R ) + f 5×5 (T R )(12)

[0153] In formula (12), T S represents the parallel convolution feature; f 3×3 represents 3×3 convolution; f 5×5 represents 5×5 convolution; T R represents the sub-feature of the flow image.

[0154] To capture the pairwise relationships between different spatial positions, the Dual Stream Fusion Strategy (DSF-Strategy) is proposed to enhance feature interaction. By integrating T EN and T S the final output feature is obtained Specifically, global average pooling is performed on T EN and the weight distribution is adjusted through Softmax(·) to obtain

[0155] T N = Softmax(Reshape(AvgPool(T EN ))(13)

[0156] In formula (13), T NIt represents the global attention feature; AvgPool represents the global average pooling operation; Reshape and Softmax represent activation functions.

[0157] Meanwhile, for T S channel weight calculation is performed to obtain the weight information of each channel

[0158] T C = Softmax(Reshape(AvgPool(T S ))) (14)

[0159] In formula (14), T C represents the processed pattern information.

[0160] Then, by multiplying the parallel convolution feature T S with the global attention feature T N we obtain the first spatial attention map which emphasizes certain important channel features and suppresses unimportant feature components. Similarly, multiplying the position enhancement feature T EN and the channel weight feature T C results in the second attention map which makes it possible to assimilate information at different scales.

[0161]

[0162] In formula (15), T MA represents the first spatial attention map; T MB represents the second spatial attention map; represents multiplication.

[0163] Next, T MA and T MB are added, and after a reshaping operation, the final fused feature is generated through the Sigmoid activation function

[0164] T MSA = Reshape(Sigmoid(T MA + T MB )) (16)

[0165] In formula (16), T MSA represents the fused feature;

[0166] Finally, multiplying T MSA element - by - element with the initial input sub - feature T R we obtain the final output result While retaining the original structure, it enhances the pixel-level feature expression at key positions:

[0167] T O = T R ⊙ T MSA (17)

[0168] In formula (17), T O represents the first output result corresponding to the network traffic image.

[0169] S2-5: Input the first output result into the first classification module to obtain the first classification result, and input the second fusion output into the second classification module to obtain the second classification result; then combine the first classification result and the second classification result to obtain the final classification.

[0170] For example, in the ISCXVPN2016 dataset, the final classification includes Chat, Mail, and FTP file transfer traffic.

[0171] In this embodiment, the present invention verifies the above technical solutions through experiments.

[0172] 1. Dataset

[0173] Four publicly available network traffic datasets are selected: Edge-IIoTset dataset, CICIoT2022dataset, ISCXVPN2016, and USTC-TFC2016 dataset to verify the effectiveness and generalization performance of the model. These network traffic datasets cover three typical scenarios of IIoT, home IoT, and encrypted traffic analysis, as well as two network traffic granularities of service and application, ensuring the performance verification of the proposed model under multi-scenario and multi-granularity conditions. The details of the network traffic datasets and the corresponding tasks are shown in Table 1.

[0174] Table 1 Experimental dataset details

[0175]

[0176] The Edge-IIoTset dataset focuses on the IIoT scenario and covers 20 types of network traffic collected on IoT devices such as ultrasonic sensors and water level detection sensors. We used 19,686 session samples in this dataset to evaluate the network traffic classification and anomaly detection capabilities of Bimodal TrafficNet in industrial scenarios.

[0177] CICIoT2022 is a comprehensive dataset for IoT traffic research, containing normal and attack traffic captured on IoT devices such as WiFi and ZigBee. In the experiment, we selected 10,315 session samples to verify the applicability of our method for traffic classification in the home IoT scenario.

[0178] The ISCXVPN2016 dataset is used to evaluate the network traffic classification performance of Bimodal TrafficNet at the service level and its sensitivity to encrypted features. This network traffic dataset was released by the Canadian Institute of Cybersecurity in 2016 and contains 6 common VPN traffic types. In the experiment, we selected a total of 5,886 session samples.

[0179] The USTC-TFC2016 dataset is a collection of real-world network traffic consisting of ten types of benign application network traffic and ten types of malware network traffic collected by researchers at CTU University. To evaluate the classification performance of the model at the network traffic application level, we extracted 28,962 sessions for the experiment.

[0180] 2. Experimental Environment Setup

[0181] The hardware platform is a server equipped with an Intel Xeon Silver 4310 CPU@2.10GHz×24 and an NVIDIA GeForce RTX 4090, along with a 64-bit Ubuntu 22.04.4 LTS operating system.

[0182] The software platform is based on the Python 3.9 environment and uses PyTorch 2.0.1 as the deep learning framework. To uniformly process the original traffic files, Wireshark software is used to convert them into the standard pcap file format, and the SplitCap 3.0.0 tool is used to achieve rapid traffic segmentation to generate session-based traffic data. In addition, traffic anonymization is completed by the Python-based Scapy 2.6.0 package.

[0183] The model parameters of Bimodal TrafficNet are set as follows: In the data stage, the size of the traffic image is set to 28×28×1, while the size of the statistical features is 1×26. In the I-Branch, the embedding dimension of the patch size is set to 192. Due to the special meaning and significantly smaller dimension of the statistical features compared to the pixel information of the traffic image, the statistical features are directly regarded as a single patch (1×1), and the embedding dimension in the F-Branch is 96. The overall architecture of Bimodal TrafficNet contains L TrafficNet Blocks. N and M Transformer Blocks are designed in the I-Branch and F-Branch respectively, and the specific values of L, M, and N will be discussed in combination with the dataset in the experiments in Section 4.

[0184] In addition, the model proposed in this paper and other models used for comparison are kept consistent in hyperparameter settings. The batch size is 256, and the number of training epochs is set to 100. The optimizer AdamW is selected, and the cosine annealing with warm restart strategy is adopted for learning rate adjustment, with an initial learning rate of 5e-4 and a warm-up period of 10 epochs. The cross-entropy loss is used as the loss function.

[0185] 3. Evaluation Metrics

[0186] To comprehensively evaluate the performance of the proposed model in the traffic classification task, the following commonly used evaluation metrics are adopted: accuracy (ACC), precision (PRE), recall (RC), and F1-score. These metrics measure the classification ability of the model from different perspectives and can effectively reflect the applicability of the model in different scenarios and granularities. The following are the definitions and formulas of each metric:

[0187]

[0188] Among them, TP represents the number of samples correctly classified as the positive class, TN represents the number of samples correctly classified as the negative class, and FP and FN represent the misclassified positive and negative class samples respectively. The F1-score is the harmonic mean of precision and recall, which synthesizes the trade-off between the two and is applicable to scenarios with class imbalance.

[0189] 4. Influence of Model Parameters on Performance

[0190] In this section, this paper studies the influence of four key variables in the model on the classification performance, namely the number of Transformer blocks (M) in the F-Branch, the number of Transformer blocks (N) in the I-Branch, the number of TrafficNet Blocks (L), and the patch size (P) in the I-Branch. Due to the different scenarios and characteristics of the four datasets, the optimal parameter configurations finally selected may vary.

[0191] 4.1. Influence of the Number of Transformer Blocks

[0192] In Bimodal TrafficNet, the settings of N and M directly affect the model's learning ability for traffic images and statistical features. Specifically, when the value of N is large, the model tends to learn richer features from traffic images; while when the value of M is large, the model pays more attention to the global information modeling ability of statistical features. This characteristic stems from the functional differences between the I-Branch and the F-Branch: the I-Branch takes traffic images as input and is good at learning complex detailed representations; while the F-Branch takes statistical features as input and focuses more on global feature extraction. Therefore, the settings of N and M not only determine the learning ability of each branch but also play a key role in their complementarity and the fusion effect of bimodal information.

[0193] As shown in Table 2, on the Edge-IIoTset dataset, the configuration (M, N) = (4, 1) achieved the best performance with an accuracy of 99.62%, indicating that in this scenario, statistical features serve as the main branch to provide key information, while traffic images enhance the detail and robustness of feature representations in the form of an auxiliary branch. On the CICIoT2022 dataset, the configuration (M, N) = (4, 4) performed optimally with an accuracy of 98.28%, which shows that in this scenario, it is necessary to enhance the learning ability of both branches simultaneously and make full use of the local information of traffic images and the global information of statistical features. In addition, on the ISCXVPN2016 and USTC-TFC2016 datasets, the configurations (M, N) = (1, 4) and (M, N) = (2, 4) achieved the best results of 99.33% and 99.36% respectively. This indicates that in the classification scenario of encrypted traffic, the complexity and fine-grained features of traffic images are particularly important. However, for datasets of different scales and granularities, there are also significant differences in their sensitivity to statistical feature modeling.

[0194] Table 2 Comparison of Model Performance under Different Conditions of M and N

[0195]

[0196] 4.2. Influence of the Number of TrafficNet Blocks

[0197] The model depth L determines the feature expression ability of Bimodal TrafficNet. Appropriately increasing the network depth can enhance the model's ability to capture complex features, but it will also bring a significant increase in the number of parameters and computational complexity. When the value of L is too large, the improvement in classification performance tends to saturate and may even decline due to overfitting. Therefore, choosing an appropriate L can achieve a good balance between the model size and classification effect.

[0198] To study the impact of L on classification performance, through the variable control technique, experiments were conducted by adjusting the value of L under the condition of fixing the optimal N and M. The optimal L values on each dataset are different. On the Edge-IIoTset dataset, when L = 2, the classification accuracy of the model reaches 99.64%. This is because the features of this dataset are relatively simple, and a smaller L value is sufficient to capture the main characteristics while avoiding the overfitting risk caused by deep networks. On the other 3 datasets, the optimal L value is 3. In contrast, the features of these datasets are more complex, and appropriately increasing the model depth helps to learn more refined features while effectively controlling the model complexity.

[0199] The experimental results show that the selection of the model depth L should be dynamically adjusted in combination with the dataset features. In scenarios with relatively simple data patterns, a shallower network can meet the classification requirements. For complex classification scenarios, appropriately increasing the network depth can significantly improve the performance, thus achieving the optimal balance between model complexity and classification effect.

[0200] 4.3. Influence of Patch Size

[0201] P determines the granularity of the traffic image segmentation before processing. A smaller patch size (i.e., higher resolution) can capture fine-grained features in the image, but may lead to an increase in computational overhead and excessive attention to details. Larger patches, on the other hand, reduce the computational burden of the model by reducing the resolution, but may lose key details. To study the impact of patch size on classification performance, experiments were conducted on four datasets by adjusting P (candidate values are 1, 2, 4, 7) under the condition of fixing the optimal N, M, and L.

[0202] On the USTC-TFC2016 dataset, the classification effect is the best when P = 2. This indicates that the traffic images of this dataset may contain complex detail information, and smaller patches can help the model capture more fine-grained features, thereby improving the classification performance. On the other 3 datasets, the model performance is the best when P = 4. This shows that the complexity of the image features of these datasets is relatively low, and using medium-sized patches can effectively capture the main features while avoiding redundant information and computational overhead caused by overly fine-grained.

[0203] The experimental results show that the choice of patch size P needs to be balanced according to the characteristics of the dataset. Datasets with higher feature complexity require smaller patches to capture image detail information. This is because the microscopic features in complex traffic images contribute significantly to the classification results, while larger patches may lead to feature loss or blurring. When the patch size is too large (e.g., P = 7), the classification performance of all datasets decreases. This indicates that larger patches cannot fully capture the detailed features in the image, resulting in deteriorated classification effects.

[0204] 5. Classification Performance Analysis

[0205] In this part, the performance of Bimodal TrafficNet in different classification tasks is first visualized through a confusion matrix to comprehensively evaluate the universality of the model in different scenarios and different granularity traffic classification tasks. In addition, through comparative experiments with other classical models, the superior performance of Bimodal TrafficNet in the IoT traffic classification task is verified.

[0206] 5.1. Performance on Different Classification Tasks

[0207] In this part, the classification performance of Bimodal TrafficNet on four datasets is verified. To more intuitively display the classification effect, the confusion matrix corresponding to each dataset is drawn, and the classification performance of the model for different scenarios and different granularity traffic is analyzed in detail based on it. In addition, the classification accuracy of each category of the model on the test dataset is also counted. The diagonal elements in the confusion matrix represent the accurate classification ratio of the model for each category, and the closer to 1, the better the classification effect, while the off-diagonal elements represent the misclassification ratio.

[0208] As Figure 3 shown, the diagonal proportion in the confusion matrix of each dataset is relatively high, indicating that Bimodal TrafficNet has achieved high classification accuracy in different traffic scenarios and granularity tasks, demonstrating its excellent generalization ability and versatility. Specifically, the model has a good classification effect on ISCXVPN2016, and the accuracy of all categories exceeds 90%, and the accuracy of Voip and Chat is 100%. This is because the scenario of this dataset is relatively simple, and the feature differences between different categories are relatively obvious, and the model can easily capture the features and achieve accurate classification. Similarly, for EDGE-IIoTset and USTC-TFC2016, the classification accuracy of the model for most categories is as high as 95%-100%, which indicates that the category distribution of the dataset has obvious feature differences, which enables the model to better capture its patterns, indicating that the model has high applicability in the IoT traffic classification scenario.

[0209] Despite the excellent overall performance, there are still some misclassification phenomena among certain similar traffic types in the model. For example, on the CICIoT2022 dataset, the classification accuracy of the model on the similar categories Interactions_Audio and Interactions_Other has decreased, but the accuracy of the mainstream categories such as Flood is still close to 100%. This is because the characteristics of some traffic types in this dataset are relatively similar, resulting in the model being confused between these categories. This indicates that there is still room for improvement in the model when dealing with traffic types with high similarity.

[0210] The experimental results show that the classification performance of Bimodal TrafficNet on the four datasets demonstrates its good generality and robustness, and it can effectively handle tasks in different scenarios and with different traffic granularities. This verifies the effectiveness of the bimodal fusion mechanism in the model design and proves its application potential in IoT traffic analysis tasks.

[0211] (a) Performance confusion matrices on the Edge-IIoTset dataset, (b) CIC-IoT2022 dataset, (c) ISCXVPN2016 dataset, and (d) USTC-TFC2016 dataset.

[0212] 5.2, Comparison with Other Methods

[0213] To verify the effectiveness of the proposed Bimodal TrafficNet method in this paper, it was compared with existing classical methods. These methods represent different network design and task ideas, including 1D-CNN using CNN as the backbone network, the deep learning framework CAD-Net based on ResNet, the traffic classification framework ViT using Vision Transformer, and the malware classification method MTC-MAE combining MAE and self-supervised framework. The experiments evaluated the performance of each method on different datasets from two dimensions: accuracy and F1 score, and the results are shown in Table 3.

[0214] First, Bimodal TrafficNet performs excellently in different scenarios. Especially on the Edge-IIoTset dataset, Bimodal TrafficNet achieves an accuracy of 99.64% and an F1-score of 98.39%, which are approximately 0.27% and 1.23% higher than 99.37% and 97.16% of CAD-Net respectively. On the CICIoT2022 dataset in the Home IoT scenario, Bimodal TrafficNet obtains an accuracy of 98.28% and an F1-score of 96.16%, significantly outperforming 95.00% and 89.59% of ViT, indicating the high adaptability and precision of the proposed method for IoT scenario traffic classification.

[0215] In addition, Bimodal TrafficNet also demonstrates excellent capabilities in classification tasks at different granularity levels. On the ISCXVPN2016 and USTC-TFC2016 datasets corresponding to service-level and application-level classifications respectively, Bimodal TrafficNet achieves accuracies of 99.33% and 99.69% respectively, which proves that its unique dual-branch structure can effectively capture and fuse bimodal information.

[0216] Generally speaking, compared with traditional model-based methods such as 1D-CNN and CAD-Net, BimodalTrafficNet based on ViT can better mine the detailed information in different traffic data; while compared with methods such as ViT and MTC-MAE, its unique dual-branch structure gives full play to the complementary advantages of traffic statistical features and traffic images through an efficient information interaction mechanism, thus achieving a double improvement in accuracy and generalization.

[0217] Table 3 Performance comparison with other methods

[0218]

[0219] 6. Ablation experiments

[0220] To verify the impact of different modules in the model on the overall performance, a series of ablation experiments were designed. Specifically, by removing the F-Branch, I-Branch, and PLIA-Module respectively to explore the contributions of each component on different datasets. In addition, removing the F-Branch or I-Branch also means removing the BCA-Module in the model.

[0221] As shown in Table 4, without the statistical feature branch, the performance of the four datasets all decreased to varying degrees. This indicates that the F-Branch plays a significant role in traffic behavior analysis and statistical feature capture, especially in scenarios where the feature distribution of the dataset is relatively sparse, and its effectiveness is more obvious. In addition, after removing the I-Branch, the performance of each dataset decreased significantly. Among them, the F1 value of the CICIoT2022 dataset decreased from 96.16% to 58.37%. This shows that the graphical information provided by the I-Branch significantly improves the discrimination ability of the model. In addition, the ablation experiments of the two branches also indirectly verified the irreplaceable role of the dual-modal information interaction mechanism of the BCA-Module in feature extraction and traffic recognition.

[0222] After removing the PLIA-Module, the performance of the model decreased on the CICIDS2017, CICIoT2022, and ISCX2012 datasets, while the impact on USTC-TFC2016 was not obvious. This indicates that the PLIA-Module plays an auxiliary role in refining the feature extraction of the image branch and enhancing the dual-modal information interaction, especially in tasks with complex traffic patterns and high requirements for detail expression, and its contribution is more prominent.

[0223] Table 4 Ablation experiment results of Bimodal TrafficNet

[0224]

[0225] Note: w / o F-Branch: Remove the F-Branch (including the BCA-Module), w / o I-Branch: Remove the I-Branch (including the BCA-Module), and w / o PLIA: Remove the PLIA-Module

[0226] This paper proposes a novel dual-modal IoT traffic classification model, Bimodal TrafficNet, to address the challenges of increasingly complex and diverse network traffic in IoT scenarios. Aiming at the problems of common single information utilization and insufficient feature interaction in existing methods, this paper innovatively introduces the interaction module BCA-Module to achieve multi-layer information interaction between network traffic statistical features and network traffic images, significantly improving the model's perception ability of complex network traffic patterns. In addition, the application of the PLIA-Module further optimizes the feature extraction effect of traffic images. The robustness and effectiveness of BimodalTrafficNet in diverse network traffic scenarios and different classification granularities are verified on 4 public datasets.

[0227] In future research, the focus will be on the lightweight design of the model to better adapt to the resource-constrained characteristics of IoT devices. Specifically, by introducing model compression techniques such as parameter pruning, low-bit quantization, and knowledge distillation, the number of model parameters and computational complexity will be further reduced, thereby reducing the resource consumption of hardware deployment. In addition, we plan to enhance the model's adaptability to unknown traffic types so that it can perceive new traffic patterns in a dynamic environment and improve the detection performance of unknown threats.

[0228] Those of ordinary skill in the art can understand that the above embodiments are specific examples for implementing the present invention, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present invention.

Claims

1. A dual-modal IoT traffic classification method based on fusion images and statistical features, characterized in that: The following steps are involved: S1: Obtain the original IoT traffic data and preprocess it to obtain input data; S2: Input the input data into the constructed traffic classification model for training and classification to obtain the classification results.

2. The dual-modal Internet of Things traffic classification method based on fusion image and statistical features as claimed in claim 1 is characterized in that: The S1 includes: S1-1: Split the acquired original IoT traffic data into multiple sessions, and then perform cleaning processing to obtain first data; S1-2: extracting statistical features from the first data, then standardizing the statistical features, and sorting the statistical features from large to small based on the standard deviation of the statistical features to obtain a statistical feature input sequence; S1-3: Convert the first data into a traffic image.

3. The dual-modal Internet of Things traffic classification method based on fusion image and statistical features as claimed in claim 2 is characterized in that: In S1-1, the first data acquisition method is: First, the traffic is parsed using the five-tuple information of the network traffic, and the data is split into sessions; then, the generated sessions are cleaned, and the cleaning process includes deleting duplicate session files and removing empty files, so as to obtain the first data.

4. The dual-modal Internet of Things traffic classification method based on fusion image and statistical features as claimed in claim 2 is characterized in that: In S1-2, the standardization is: In formula (1), F S Represents the standardized statistical characteristics, F is the original statistical characteristics; μ is the mean; σ is the standard deviation; ε is a small constant.

5. The dual-modal Internet of Things traffic classification method based on fusion image and statistical features as claimed in claim 2 is characterized in that: In S1-3, the conversion method of the traffic image is: For each session file, it is trimmed to a normalized 784 bytes, that is, if the packet length exceeds 784 bytes, it is truncated; if the packet length is less than 784 bytes, it is padded with 0x00; then, the data is converted into a 28×28×1 grayscale image, where each pixel corresponds to the byte value in the network traffic data, ranging from 0x00 to 0xFF.

6. The dual-modal Internet of Things traffic classification method based on fusion image and statistical features as claimed in claim 1, characterized in that: The S2 includes: S2-1: Build a traffic classification model; S2-2: input the traffic image and the statistical features into the traffic classification model and divide them into the first patch and the second patch respectively; S2-3: extracting first context information from the first patch and extracting second context information from the second patch respectively; S2-4: exchanging the first context information and the second context information to generate a first fusion output corresponding to the traffic image and a second fusion output corresponding to the statistical feature; S2-5: Refine the features of the first fusion output to obtain a first output result; S2-6: Input the first output result into the first classification module to obtain a first classification result, and input the second fusion output into the second classification module to obtain a second classification result.

7. The dual-modal Internet of Things traffic classification method based on fusion image and statistical features as claimed in claim 6 is characterized in that: In S2-1, the traffic classification model includes a first processing module, a first branch module, a first classification module, a second processing module, a second branch module, a second classification module, and an interaction module; A first processing module, used for processing the traffic image and dividing it into a first patch; a first branch module, used for extracting first context information from the first patch to obtain a first patch token and combining it with a first CLS token to obtain a first fusion output; A first classification module, used to obtain a first classification result according to the first fusion output; A second processing module, used for processing the statistical features and dividing them into second patches; A second branch module, configured to extract second context information from the second patch to obtain a second patch token, and combine the second CLS token to obtain a second fusion output; A second classification module, used to obtain a second classification result according to the second context information; The interaction module is used for interacting the first context information and the second context information, and continuously updating the first fusion output and the second fusion output.

8. The dual-modal Internet of Things traffic classification method based on fusion image and statistical features as claimed in claim 6, characterized in that: In S2-3, the method for extracting the first context information and the second context information is: R i =T i-1 +MHSA(LN(T i-1 )),T i =R i +FFN(LN(R i )) (2) In formula (2), R i represents the intermediate result of the multi-head self-attention module plus the residual connection, T i represents the output of the i-th layer of TransformerBlock; T i-1 Represents the output of the i-1th layer of Transformer Block; MHSA(LN(T i-1 )) represents the multi-head self-attention mechanism function; FFN(LN(R i )) represents a feedforward neural network function.

9. The dual-modal Internet of Things traffic classification method based on fusion image and statistical features as claimed in claim 6, characterized in that: In S2-4, the traffic image first converts the corresponding CLS token Patch tokens with statistical features To concatenate, the formula is: In formula (3), T' I Represents a concatenated token, Indicates the CLS token corresponding to the traffic image; A patch token representing a statistical feature; p I (·) is the dimension-aligned projection function; The information of the patch token corresponding to the traffic image is fused into the CLS token, which is mathematically represented as follows: In formula (4), are learnable parameters, D and h are embedding dimension and number of heads respectively; Q represents query; K represents key; V represents value; Represents the CLS token after the dimension-aligned projection function; T' I represents the concatenated token; A represents the attention weight matrix; M represents the output matrix after weighting by the attention mechanism; T represents the transposed matrix; After receiving the information of the patch token in the statistical feature, it is then back-projected back to the flow image to pass its own patch token Specifically, the output after fusion of traffic image information is The definition is as follows: In formula (5), Indicates the intermediate result of projection; g I (·) is the back-projection function; p I (·) is the dimension-aligned projection function; Indicates the CLS token corresponding to the traffic image; Indicates the patch token corresponding to the traffic image; Represents the first fusion output corresponding to the traffic image.

10. The dual-modal Internet of Things traffic classification method based on fusion image and statistical features as claimed in claim 6, characterized in that: In S2-5, the method for obtaining the first output result is: First, the first fusion output is input and divided into two groups of sub-features in the embedding dimension Where B is the batch size, b = 2B, d = D / 2, h = H / P, w = W / P, H is the height, W is the width, and D is the spatial dimension; Then, T is measured along the height and width directions respectively. R To perform pooling operation: T H =AvgPool H (T R ),T W =Reshape(AvgPool W (T R )) (6) In formula (6), and Respectively represent the global information of the feature map in the height and width dimensions; AvgPool H Represents the pooling operation in the height direction; AvgPool W Represents the pooling operation in the width direction; Reshape represents the activation function; Subsequently, a unified feature representation is generated through concatenation and 1×1 convolution fusion: T F =f 1×1 (Concat(T H ,T W )) (7) In formula (7), Concat means concatenation; represents unified feature representation; T F Divided into two vectors, unified feature representation and And generate weights T through activation function HI and T WI : T HS ,T WS =Split(T F ),T HI =Sigmoid(T HS ),T WI =Reshape(Sigmoid(T WS )) (8) In formula (8), T HS A unified feature representation representing the height direction; T WS represents the unified feature representation in the width direction; Split represents the splitting function; T HI Indicates the weight in the height direction; T WI Represents the weight in the width direction; Sigmoid and Reshape both represent activation functions; Finally, T HI , T WI By and sub-feature T R After element-by-element multiplication and reshaping, we get the position enhancement feature T EN : T EN =Reshape(T R ⊙T HI ⊙T WI ) (9) In formula (9), T EN Represents position enhancement features; Reshape represents activation function; ⊙ represents element-by-element multiplication; The parallel operation of 3×3 convolution and 5×5 convolution is also introduced to capture the diverse pattern information in the local area, and T S : T S =f 3×3 (T R )+f 5×5 (T R ) (10) In formula (10), T S Represents parallel convolution features; f 3×3 represents 3×3 convolution; f 5×5 represents 5×5 convolution; T R Represents the sub-features of the traffic image; To T EN Perform global average pooling and adjust the weight distribution through Softmax(·), and get T N =Softmax(Reshape(AvgPool(T EN )) (11) In formula (11), T N Represents the global attention feature; AvgPool represents the global average pooling operation; Reshape and Softmax represent activation functions; At the same time, for T S Perform channel weight calculation to obtain the weight information of each channel T C =Softmax(Reshape(AvgPool(T S ))) (12) In formula (12), T C Indicates the processed mode information; By parallelizing the convolutional features T S With the global attention feature T N Multiply them together to get the first spatial attention map The position enhancement feature T EN And channel weight feature T C Proceed to obtain the second attention map In formula (13), T MA represents the first spatial attention map; T MB represents the second spatial attention map; Indicates multiplication; T MA and T MB Add, reshape and use Sigmoid activation function to generate the final fusion feature T MSA =Reshape(Sigmoid(T MA +T MB )) (14) In formula (14), T MSA Indicates fusion features; Will T MSA With the initial input sub-feature T R Multiply element by element to get the final output result T O =T R ⊙T MSA (15) In formula (15), T O Indicates the first output result corresponding to the network traffic image.