Malicious traffic classification method and system of feature fusion network based on gradient sharing
By adopting a model of feature fusion and gradient sharing in the 5G IoT environment, combining the advantages of convolutional neural networks and KANs, the shortcomings of traditional methods in feature extraction and environmental adaptability are solved, and more efficient, accurate and adaptable malicious traffic detection is achieved.
Patent Information
- Application Number
- CN202510347485.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The prior art is difficult to fully capture complex traffic patterns in 5G and IoT environments, and it is poorly adaptable to dynamically changing network environments. Traditional methods cannot take into account the synergy between local and global features when extracting features, resulting in limited detection accuracy.
The 5G IoT malicious traffic classification model based on feature fusion and gradient sharing is adopted. Through the architecture of a combination of convolutional neural network and KAN, deep information learning of local and global features is realized, and the gradient sharing mechanism is used to collaborate and optimize each other during the training process.
It improves the model's learning ability on complex traffic patterns, reduces dependence on specific types of features, enhances the model's adaptability and robustness in different network environments, and significantly improves the accuracy and generalization ability of malicious traffic detection.
Smart Images

Figure CN120075802A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to, but is not limited to, the technical field of 5G Internet of Things and network security, and particularly relates to a malicious traffic classification method and system based on a feature fusion network with gradient sharing. Background Art
[0002] With the rapid development of 5G networks and Internet of Things technologies, network traffic data has become a key basis for ensuring network security and detecting malicious traffic. However, the current network environment is becoming increasingly complex. The low-latency, high-bandwidth, and large-scale connection characteristics of 5G networks, as well as the widespread access of Internet of Things devices, make network traffic highly complex and dynamic. Traditional malicious traffic detection methods face severe challenges. Although some traditional methods have certain effects in specific scenarios, their limitations are becoming increasingly prominent when facing complex 5G and Internet of Things environments, and it is difficult to effectively cope with diverse malicious traffic attacks.
[0003] In the early stage, malicious traffic detection mainly used traditional methods such as rule matching and statistical feature analysis. However, these methods have many problems: (1) Insufficient feature modeling ability. When existing detection methods extract network traffic features, they are often limited to the modeling of a certain type of feature, such as only focusing on local statistical information or global traffic relationships, and it is difficult to simultaneously obtain complementary information of local and global features. In a 5G environment, the high concurrency and low latency characteristics of network traffic make attack patterns more concealed. For example, the traffic fluctuations of a DDoS attack may show large-scale global patterns, and models relying only on local features are difficult to effectively identify; while relying only on global features will ignore the microscopic information at the packet level, resulting in false negatives or false positives when facing complex attack scenarios. (2) Insufficient adaptability in a dynamically changing environment. With the popularization of 5G networks and Internet of Things technologies, the network environment, user behavior, and attack strategies are constantly changing, and network traffic patterns also evolve accordingly. Many traditional methods are based on fixed features and static rules for detection and are difficult to adapt to these changes. In a 5G network, the dynamically switched network topology will affect the attack path, making traditional traffic analysis methods ineffective; malicious attackers will also quickly adjust their strategies to evade detection, and the detection performance of traditional static models will drop significantly when traffic patterns change significantly.
[0004] In view of the above analysis, the technical problems urgently to be solved in the prior art are that existing methods are difficult to comprehensively capture complex traffic patterns in 5G and Internet of Things environments and have poor adaptability to dynamically changing network environments. When traditional methods extract features, they cannot take into account the synergistic effect of local and global features, resulting in limited detection accuracy; at the same time, in the face of constantly changing network environments and attack strategies, traditional detection methods based on fixed features and static rules are difficult to quickly adapt, and the detection performance is unstable. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the present invention provides a 5G Internet of Things malicious traffic classification model and method based on feature fusion and gradient sharing, aiming to solve the deficiencies of traditional malicious traffic detection methods in feature modeling and environmental adaptability. Through feature fusion and gradient sharing technologies, the learning ability of the model for complex traffic patterns is improved, the dependence on specific types of features is reduced, and the adaptability and robustness of the model in different network environments are enhanced.
[0006] The present invention is implemented as follows. A malicious traffic classification method based on a feature fusion network with gradient sharing includes the following steps:
[0007] S1. Data preprocessing: Preprocess the original 5G Internet of Things traffic data, and construct a structured feature channel and an image feature channel respectively. For the structured feature channel, use Wireshark to parse the PCAP file in the 5G traffic data into a vector format, extract communication features, and perform standardization processing; for the image feature channel, map the byte distribution of the 5G traffic into an image format, reorganize the original byte stream in the PCAP file into a fixed-length one-dimensional vector, and after binarization, select the first 784 bits and map them into a grayscale image.
[0008] S2. Local feature extraction: Use a local feature extraction module based on a convolutional neural network to process the data in the image feature channel, and automatically learn the local spatial characteristics in the traffic data by stacking convolutional layers and pooling layers, and extract microscopic pattern features such as packet length, time interval, and protocol identifier.
[0009] S3. Global feature extraction: Construct a global feature extraction module through KAN to dynamically model the data in the structured feature channel, capture the time series distribution and global dependence relationship of the traffic, and construct the overall behavior characteristics of the traffic from a macroscopic level. When constructing the global feature, adopt the method of gradually inputting samples in chronological order to simulate the real network environment.
[0010] S4. Feature fusion: In the feature fusion module, fuse the features extracted from the structured feature channel and the image feature channel. For the structured feature channel, extract both packet-level features such as source IP, destination IP, protocol type, and packet length, and data stream-level features such as flow duration, total number of packets, and total number of bytes, and combine the two to construct a comprehensive feature representation for each packet; for the image feature channel, convert the byte stream into a grayscale image and then extract local features by a convolutional neural network. Finally, merge the features of the two channels into a unified high-dimensional feature vector to simultaneously obtain the global characteristics and local details of the traffic data.
[0011] S5, Gradient Sharing Initialization: Before model training, initialize the parameters related to the gradient sharing mechanism, including determining the formulas for calculating the gradient weights of the convolutional neural network and KAN channels, initializing the KAN and convolutional neural network models, and the optimizer.
[0012] S6, Gradient Sharing Training: During the training process, for each batch of training data, input the structured features and image features into the KAN and convolutional neural network models respectively, calculate their respective outputs and losses. Calculate the gradients of the convolutional neural network and the KAN channels themselves, as well as the gradient guidance between them, according to the losses. Calculate the final updated gradients for each channel by dynamically adjusting the weights.
[0013] S7, Model Training and Optimization: Through multiple rounds of training, using the gradient sharing mechanism, enable the KAN and convolutional neural network models to cooperate and co-optimize with each other during the feature learning process, continuously adjust the model parameters, improve the model's ability to identify malicious traffic, reduce the conflict between local and global feature learning, and enhance the model's sensitivity to diverse malicious traffic patterns.
[0014] S8, Model Testing: Test the trained malicious traffic classification model on the test set, and evaluate the model performance through metrics such as accuracy, precision, recall, and F1 score to verify the classification accuracy of the model for malicious traffic in different 5G Internet of Things traffic scenarios.
[0015] Furthermore, S1 specifically includes: Process the collected original 5G Internet of Things traffic dataset, convert it into a format suitable for model input, and construct the data for the structured feature channel and the image feature channel. For the construction of the structured feature channel data: First, use the Wireshark tool to parse the PCAP files in the 5G traffic data into vector format, and extract rich communication features such as source IP, destination IP, protocol type, packet length, flow duration, total number of packets, and total number of bytes during the communication process. These features cover information at the packet level and the data flow level, comprehensively reflecting the static attributes and dynamic behaviors of network communication. Subsequently, perform standardization processing on all the extracted features. Assume X is the current value, X max is the maximum value of this feature, X min is the minimum value of this feature, and the normalization formula is:
[0016]
[0017] Unify features with different magnitudes and distributions into the same scale range, avoid inconsistent influence weights of different features on model learning, and ensure that the model can learn feature information more effectively.
[0018] For the construction of image feature channel data, the present invention converts the byte distribution of 5G traffic into an image form. The specific operation is to reorganize the original byte stream in the PCAP file into a one-dimensional vector of a fixed length, and instead of normalizing it, directly perform binary processing on it to obtain a series of vectors consisting only of 0s and 1s. Select the first 784 bits from them and map them into a 28×28 grayscale image, where each pixel value corresponds to a byte value in the vector, so that the subsequent convolutional neural network module can extract the local spatial characteristics of the traffic from it.
[0019] Further, S2 specifically includes: using a local feature extraction module based on a convolutional neural network to process the data in the image feature channel. The convolutional neural network module is composed of two convolutional layers and a pooling layer stacked. In the convolutional layer, convolutional kernels of different sizes and strides slide and convolve on the image to automatically learn local spatial characteristics in the traffic data, such as microscopic pattern features like packet length, time interval, and protocol identifier. In the pooling layer, max pooling or average pooling operations are used to downsample the feature map output by the convolutional layer, reducing the data volume while retaining key features, reducing the model's computational complexity, and improving the model's training efficiency and generalization ability.
[0020] Further, S3 specifically includes: constructing a global feature extraction module with KAN to perform dynamic modeling on the data in the structured feature channel. The KAN module receives the preprocessed structured feature vector and dynamically captures the time series distribution and global dependency relationship of the traffic through its unique network structure. When constructing the global feature, instead of inputting all the data at once, the samples are input step by step in chronological order to simulate the dynamic change process of the data in the real network environment, constructing the overall behavior characteristics of the traffic from a macroscopic level, capturing macroscopic information such as the time series characteristics of the traffic and the network topology structure, and forming a complement to the local features extracted by the convolutional neural network.
[0021] Further, S4 specifically includes: In the feature fusion module, fuse the features extracted from the structured feature channel and the image feature channel. For structured features, at the data packet level, extract features such as source IP, destination IP, protocol type, and data packet length, which provide static information of traffic samples; at the data flow level, group traffic samples with the same five-tuple (source IP, destination IP, source port, destination port, protocol) as a data flow, and extract features such as flow duration, total number of data packets, and total number of bytes. These features reflect the spatio-temporal correlation between data packets and the overall behavior characteristics of the data flow from a dynamic perspective. Then, add the features at the data flow level after the features of the corresponding data packets to construct a comprehensive feature representation for each data packet, so that it contains both its own attributes and the context information of the data flow to which it belongs. For image features, after the feature extraction is completed by the convolutional neural network module, the features of the two channels are concatenated into a unified high-dimensional feature vector in a certain order, so that the model can simultaneously obtain the global characteristics and local details of traffic data, and improve the ability to identify malicious traffic.
[0022] Further, S5 specifically includes: Before model training, initialize the parameters related to the gradient sharing mechanism. Determine the formulas for calculating the gradient weights of the image feature channel and the structured feature channel. Assume ω Local is the weight of the image feature channel, ω Global is the weight of the structured feature channel, L Local is the loss of the image feature channel, L Global is the loss of the structured feature channel. Then the weight formulas for the two channels are respectively:
[0023]
[0024] Further, S6 specifically includes: During the training process, for each batch of training data, input the structured features and image features into the KAN and the convolutional neural network model respectively. The KAN model outputs the prediction results for the global features, and the convolutional neural network model outputs the prediction results for the local features. Calculate the loss values between them and the true labels respectively to obtain the gradients G Local,self and G Global,self of the convolutional neural network and the KAN channel itself. Then, according to the gradient sharing mechanism, calculate the gradient guidance G Global→Local of the structured feature channel for the image feature channel and the gradient guidance G Local→Global of the image feature channel for the structured feature channel. Through the previously calculated weights ω Local and ω Global , the gradient calculation formulas for the two channels are as follows:
[0025] G Local =ω Local GLocal,self +(1 - ω Local )G Global,self (4)
[0026] G Global =ω Local G Local,self +(1 - ω Local )G Global,self (5)
[0027] Finally, update the parameters of the image feature channel and the structured feature channel models respectively according to the updated gradients, so that the two models cooperate with each other and optimize synergistically during the training process, improving the model's ability to identify malicious traffic.
[0028] Furthermore, S7 specifically includes: through multiple rounds of training, using the gradient sharing mechanism, enabling KAN and the convolutional neural network model to continuously cooperate with each other during the feature learning process. In each round of training, the two models continuously adjust their own parameters according to the shared gradient information to optimize the learning effect of malicious traffic features. As the number of training rounds increases, the model gradually reduces the conflict between local and global feature learning and enhances the sensitivity to diverse malicious traffic patterns. During the training process, observe the changes in indicators such as the accuracy rate and loss value of the model on the validation set. When the performance indicators of the model on the validation set tend to be stable and there is no obvious improvement, it is considered that the model training reaches a good state, completing the model training and optimization process, and obtaining a trained malicious traffic classification model.
[0029] Furthermore, S8 specifically includes: testing the trained malicious traffic classification model on the test set. In the testing stage, process the 5G Internet of Things traffic data in the test set into structured and image-based features suitable for model input according to the preprocessing method in the training stage, and input them into the trained model. The model infers the test data according to the learned feature patterns and outputs the prediction results of whether each traffic sample is malicious traffic.
[0030] Another object of the present invention is to provide a malicious traffic classification system based on a gradient-sharing feature fusion network for a malicious traffic classification method based on a gradient-sharing feature fusion network, including:
[0031] A data preprocessing module for processing 5G traffic data into structured and visualized features suitable for model input;
[0032] A local feature extraction module for extracting local microscopic features of traffic data based on a convolutional neural network;
[0033] A global feature extraction module for constructing global behavior features of traffic using KAN;
[0034] Feature Fusion Module: used to fuse the features extracted from the structured feature channel and the image feature channel;
[0035] Gradient Sharing Module: used to enable two modules to complement each other and optimize collaboratively in terms of feature representation, improving the training stability and convergence speed of the model.
[0036] Combining the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by the present invention are as follows:
[0037] First, the present invention proposes a 5G Internet of Things malicious traffic classification model and method based on feature fusion and gradient sharing. The Convolutional Neural Network (CNN) has a powerful local feature extraction ability, while the Kolmogorov–Arnold Networks (KAN) is good at modeling global features. The present invention combines the two. Through the feature fusion module, it can simultaneously learn the deep information of local and global features, effectively capturing the complex patterns of network traffic.
[0038] At the same time, to improve the collaboration and robustness of the model, the present invention introduces a gradient sharing mechanism. This mechanism allows multiple sub-models to dynamically and real-time share gradients and jointly optimize the loss function, enabling the sub-models to guide and correct errors in each other during the training process. This two-way learning mechanism overcomes the one-way learning limitation of the traditional knowledge distillation method where the "teacher model guides the student model", ensuring that the refined modeling of local features and the macroscopic abstraction of global features complement each other. It not only reduces the conflict between local and global feature learning but also improves the sensitivity of the model to diverse malicious traffic patterns.
[0039] Generally speaking, the present invention combines the advantages of feature fusion and gradient sharing, providing a more efficient, accurate, and adaptable 5G Internet of Things malicious traffic classification model and method. Through feature fusion and gradient sharing technologies, it improves the model's learning ability for complex traffic patterns, reduces the dependence on specific types of features, and enhances the adaptability and robustness of the model in different network environments. This method can effectively cope with the diversity of malicious traffic patterns and the dynamic changes of network environments, and has broad application prospects and important commercial value in the field of network security.
[0040] The present invention proposes a gradient sharing and feature fusion mechanism, and constructs a malicious traffic classification model based on a convolutional neural network and KAN for 5G Internet of Things malicious traffic classification. During the model construction process, structured feature channels and image feature channels are constructed through data preprocessing, and are respectively used by KAN and the convolutional neural network for feature extraction. The convolutional neural network utilizes its powerful local feature extraction ability to obtain micro-pattern features at the packet level; KAN, relying on its advantage in global feature modeling, captures the time series distribution and global dependency relationship of traffic. In the feature fusion module, the features extracted by both are fused, enabling the model to simultaneously learn the deep information of local and global features. During the training process, the gradient sharing mechanism allows multiple sub-models to dynamically and real-time share gradients and jointly optimize the loss function. In the initial stage of training, although there is no fixed master-slave relationship like traditional knowledge distillation, the two modules each learn features based on their own advantages. As the training progresses, gradient sharing is achieved by dynamically adjusting the weights, and the two modules guide and correct errors from each other. In this way, the present invention can effectively improve the accuracy of 5G Internet of Things malicious traffic classification. On multiple publicly available 5G-IoT traffic data sets, the F1 value of the model of the present invention has increased by 7% compared with existing methods, and it can accurately identify malicious traffic.
[0041] The technical effects and advantages of the technical solution of the present invention are as follows: First, the method of the present invention effectively improves the feature modeling ability. Traditional malicious traffic detection methods often only focus on a certain type of feature, while the present invention, through feature fusion and gradient sharing, takes into account the complementary information of local and global features at the same time, and comprehensively captures the complexity of network traffic. Second, in the dynamically changing 5G and Internet of Things environment, the present invention demonstrates good adaptability. The gradient sharing mechanism enables the model to quickly learn new traffic patterns, and can still maintain stable detection performance when the network environment, user behavior, and attack strategies change. Compared with traditional methods, the present invention has significantly improved in core indicators such as detection accuracy, precision, and recall, and has stronger generalization ability.
[0042] Second, as the creative auxiliary evidence of the claims of the present invention, it is also reflected in the following important aspects:
[0043] (1) The expected benefits and commercial value after the transformation of the technical solution of the present invention are as follows: The expected benefit is to improve the network security protection level of the 5G Internet of Things, provide an efficient and accurate malicious traffic detection model for the network security protection system, and reduce the losses caused by network attacks. It can be widely applied to fields such as telecommunications operators, Internet of Things device manufacturers, and enterprise network security protection, support traffic detection in multiple network environments, and has broad market prospects and commercial value.
[0044] (2) The technical solution of the present invention solves the technical problems that people have been eager to solve but have never succeeded in: solving the problem of difficult malicious traffic detection in the 5G Internet of Things environment due to complex and changeable traffic patterns, difficult to comprehensively capture features, and poor adaptability of traditional methods. Through gradient sharing and feature fusion, the present invention enables the convolutional neural network and KAN to cooperate with each other during the training process, overcomes the limitations of a single model to a certain extent, reduces the dependence on specific types of features, improves the detection accuracy and the robustness of the model, and can effectively cope with diverse malicious traffic attacks.
[0045] (3) The technical solution of the present invention overcomes the limitations of traditional methods: traditional malicious traffic detection methods cannot take into account the synergistic effect of local and global features during feature extraction, and have poor adaptability to dynamic network environments. Through the proposed gradient sharing and feature fusion mechanism, the present invention enables the convolutional neural network and KAN to share gradient information during the training process, optimize the feature learning process, reduce the conflict between local and global feature learning, enhance the sensitivity of the model to diverse malicious traffic patterns, and thus improve the accuracy and generalization ability of malicious traffic classification.
[0046] Third, by introducing a 5G Internet of Things malicious traffic classification method based on feature fusion and gradient sharing, the present invention solves a series of technical problems in the prior art and achieves significant technological progress. The following details these aspects.
[0047] Technical problems solved:
[0048] 1. Limitations of traditional methods: Existing 5G Internet of Things malicious traffic detection methods have serious deficiencies in feature modeling, usually only focusing on the modeling of a certain type of feature, and it is difficult to capture the complementary information of local and global features at the same time. Moreover, these methods mostly perform detection based on fixed features and static rules. When the network environment, user behavior, and attack strategies change, the model is difficult to quickly adapt to new traffic patterns. Through the feature fusion module, the present invention fuses the powerful local feature extraction ability of CNN and the advantage of KAN in global feature modeling, and simultaneously learns the deep information of local and global features; with the help of the gradient sharing mechanism, it allows multiple sub-models to dynamically and real-time share gradients and jointly optimize the loss function, enabling the model to quickly adjust the learning strategy in the dynamic 5G and Internet of Things environment and improve the adaptability of detection.
[0049] 2. Poor generalization ability: Traditional malicious traffic detection methods are often designed for specific network environments or datasets and lack broad applicability. Facing the complex and diverse traffic patterns in the 5G Internet of Things environment, including encrypted traffic, industrial Internet of Things scenario traffic, etc., traditional methods are difficult to effectively cope with. In the present invention, during the feature extraction and training process, various types of feature information are fused, and gradient sharing is used to promote collaborative learning among sub-models, enabling the model to learn more general traffic feature representations, thereby improving the generalization ability on different network environments and traffic datasets.
[0050] 3. Low accuracy: Due to the complexity and dynamism of 5G Internet of Things traffic, traditional detection methods are prone to false negatives or false positives, resulting in low detection accuracy. Some methods may rely too much on certain types of features and cannot accurately identify malicious traffic when facing complex attack scenarios. In the present invention, more comprehensive traffic features are obtained through feature fusion, combined with the gradient sharing mechanism to enhance the model's sensitivity to diverse malicious traffic patterns, and the model parameters are continuously optimized during the training process, thereby improving the overall detection accuracy and F1 value.
[0051] Obtained significant technological progress:
[0052] 1. Reduced the required abnormal label samples: The present invention uses gradient sharing and feature fusion for model training. In the initial stage of training, each sub-model can perform feature learning based on its own advantages. As the training progresses, collaborative optimization is achieved through gradient sharing. This method enables the model to effectively learn the feature patterns of malicious traffic with only partial normal traffic samples and a small number of malicious traffic samples, reducing the dependence on large-scale malicious traffic annotation data.
[0053] 2. Improved generalization ability: The method of the present invention is applicable to the detection of 5G Internet of Things traffic in various network environments, including network traffic, encrypted traffic, industrial Internet of Things scenarios, and 5G traffic, etc. Through feature fusion and gradient sharing, the model can learn the common and different features of different types of traffic, so that when facing new network environments and traffic patterns, it can quickly adapt and accurately detect malicious traffic, improving the applicability of the model in different environments.
[0054] 3. Improved accuracy and efficiency: Through the gradient sharing mechanism, the convolutional neural network and the KAN two sub-models guide and correct errors with each other during the training process, optimizing the feature learning process. This collaborative training method not only improves the detection accuracy but also reduces the training time and computational overhead compared with traditional single-model training. Experiments on multiple public datasets show that the malicious traffic detection model is significantly superior to existing methods in core indicators such as detection accuracy and false positive rate, and the classification accuracy has increased by up to 7%, and the training convergence speed is also faster.
[0055] The method provided by the present invention has achieved remarkable progress technically, effectively solved multiple defects of the prior art, and provided an innovative solution in the field of 5G Internet of Things malicious traffic detection. These technical improvements not only improve the accuracy of detection, but also reduce the dependence on large-scale labeled data, which is of great significance for the development of 5G Internet of Things network security protection.
[0056] Fourth, the model and technical design adopted by the 5G Internet of Things malicious traffic classification method based on feature fusion and gradient sharing of the present invention address several key technical problems in the prior art and achieve remarkable technical progress.
[0057] 1. Feature fusion and gradient sharing mechanism: The present invention adopts an architecture combining a convolutional neural network and KAN. Through the feature fusion module, it utilizes the powerful local feature extraction ability of the convolutional neural network and the advantage of KAN in modeling global features to simultaneously learn the deep information of local and global features. During the training process, a gradient sharing mechanism is introduced. In the initial stage, each module conducts feature learning based on its own advantages. As the training progresses, different modules dynamically adjust their parameters by sharing gradient information, and dynamically interact and share the gradient information of the local feature extraction module and the global feature modeling module to enhance the recognition ability of complex malicious traffic patterns.
[0058] 2. Model collaborative optimization: During the gradient sharing process, the convolutional neural network and KAN models collaborate and optimize each other. When each sub-model performs backpropagation, it not only updates its own parameters but also receives gradient signals from other modules, achieving collaborative optimization of multiple modules during the feature learning process. This mechanism ensures complementarity between the refined modeling of local features and the macroscopic abstraction of global features, enabling more comprehensive learning of malicious traffic patterns and improving the detection performance.
[0059] 3. Reducing the dependence on specific features and large-scale labeled data: Through feature fusion, the model can comprehensively utilize various types of feature information, reducing the dependence on a single type of feature. At the same time, the gradient sharing mechanism enables the model to better learn the features of different traffic patterns during training. Even when only using some normal traffic samples and a small number of malicious traffic samples for training, it can effectively learn the feature patterns of malicious traffic, reducing the dependence on large-scale malicious traffic labeled data.
[0060] Technical problems solved:
[0061] 1. Limitations of a single model: Traditional 5G Internet of Things malicious traffic detection methods usually rely on a single model, making it difficult to simultaneously consider the extraction and analysis of local and global features, resulting in limited detection capabilities. The present invention enables the convolutional neural network and KAN to cooperate with each other through feature fusion and gradient sharing, leveraging their respective advantages to enhance the detection capabilities.
[0062] 2. Insufficient generalization ability: Previous methods were usually designed for specific network environments or datasets and it was difficult to adapt to the complex and diverse traffic patterns in the 5G Internet of Things environment, including encrypted traffic, industrial Internet of Things scenario traffic, etc. Through feature fusion and gradient sharing, the present invention enables the model to learn more general traffic feature representations and improves the adaptability to different network environments and traffic datasets.
[0063] 3. Accuracy of anomaly detection: When existing methods detect malicious traffic in the 5G Internet of Things, due to the inability to comprehensively capture complex traffic patterns, problems such as high false alarm rates or missed detections are likely to occur. The present invention obtains more comprehensive traffic features through feature fusion, combines the gradient sharing mechanism to enhance the sensitivity of the model to diverse malicious traffic patterns, and continuously optimizes the model parameters during the training process to improve the accuracy and stability of detection.
[0064] Technical progress obtained:
[0065] 1. Improved the accuracy rate of malicious traffic detection in the 5G Internet of Things: Through feature fusion and gradient sharing of the convolutional neural network and KAN, the present invention effectively improves the accuracy of malicious traffic detection. On multiple publicly available 5G-IoT traffic datasets, the F1 value of the model has increased by 7% compared with existing methods.
[0066] 2. Reduced the number of labeled samples required for model training: The present invention mainly relies on feature fusion and gradient sharing for malicious traffic feature learning, greatly reducing the need for large-scale malicious traffic annotation data. It can achieve efficient model training when only using some normal traffic samples and a small number of malicious traffic samples.
[0067] 3. Enhanced the adaptability and generalization ability of the system: The feature fusion and gradient sharing mechanism enables the present invention to adapt to 5G Internet of Things traffic in different network environments, including network traffic, encrypted traffic, industrial Internet of Things scenarios, and 5G traffic, etc., improving the detection ability for unknown malicious traffic patterns.
[0068] The present invention not only solves the existing problems technically, but also demonstrates significant technical advantages in practical applications, providing an efficient and flexible new method for malicious traffic detection in the 5G Internet of Things. Description of the drawings
[0069] Figure 1 is the flowchart of the malicious traffic classification method of the feature fusion network based on gradient sharing provided by the embodiment of the present invention;
[0070] Figure 2 is the flowchart of feature fusion provided by the embodiment of the present invention;
[0071] Figure 3 It is a structural diagram of a malicious traffic classification system based on a gradient-sharing feature fusion network provided in an embodiment of the present invention.
[0072] Figure 4 This is the result of comparing other methods on the 5G-NIDD dataset provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0073] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0074] In view of the problems existing in the prior art in the detection of malicious traffic in 5G Internet of Things, the present invention provides a 5G Internet of Things malicious traffic classification method based on feature fusion and gradient sharing. The method integrates the powerful local feature extraction capability of convolutional neural networks and the modeling advantage of KAN for global features. With the help of the information interaction mechanism of gradient sharing, the ability to identify malicious traffic in 5G Internet of Things is effectively improved, while reducing the dependence on large-scale annotated data, and enhancing the generalization and adaptability of the model in complex network environments.
[0075] First, the raw data of 5G IoT traffic is preprocessed to construct structured feature channels and image feature channels respectively. In the construction of structured feature channels, the Wireshark tool is used to parse the PCAP file into a vector format, extract the communication features and standardize them; in the construction of image feature channels, the traffic byte distribution is mapped into an image format, and the raw byte stream is processed into a grayscale image of a specific specification.
[0076] Next, the local feature extraction module based on convolutional neural network is used to process the image feature channel data. The micro-pattern features such as data packet length and time interval are automatically learned by stacking convolutional layers and pooling layers. With the help of KAN, a global feature extraction module is constructed to dynamically model the structured feature channel data and capture the time series distribution and global dependency of the traffic.
[0077] Under the gradient sharing mechanism, the convolutional neural network and KAN modules of the present invention are trained collaboratively. Each module independently learns the traffic data features based on its own advantages, and at the same time transfers knowledge to each other by sharing gradient information, achieving collaborative optimization and improving the detection performance of malicious traffic. When one of the modules has an advantage in learning a specific malicious traffic pattern, the other module can obtain relevant information with the help of gradient sharing, thereby improving the overall detection accuracy. In addition, the gradient sharing mechanism enables the two modules to supervise each other during the training process, further optimizing the model's ability to detect malicious traffic.
[0078] likeFigure 1 As shown in the figure, an embodiment of the present invention provides a malicious traffic classification method based on a feature fusion network with gradient sharing, including the following steps:
[0079] S1, Data preprocessing: Preprocess the original 5G Internet of Things traffic data, and construct a structured feature channel and an image feature channel respectively. For the structured feature channel, use Wireshark to parse the PCAP file in the 5G traffic data into a vector format, extract communication features, and perform standardization processing; for the image feature channel, map the byte distribution of the 5G traffic into an image format, reorganize the original byte stream in the PCAP file into a fixed-length one-dimensional vector, and select the first 784 bits after binarization and map them into a grayscale image.
[0080] S2, Local feature extraction: Use a local feature extraction module based on a convolutional neural network to process the data in the image feature channel, and automatically learn the local spatial characteristics in the traffic data by stacking convolutional layers and pooling layers, and extract microscopic pattern features such as packet length, time interval, and protocol identifier.
[0081] S3, Global feature extraction: Construct a global feature extraction module through KAN to perform dynamic modeling on the data in the structured feature channel, capture the time series distribution and global dependency relationship of the traffic, and construct the overall behavior characteristics of the traffic from a macroscopic level. When constructing global features, adopt the method of gradually inputting samples in chronological order to simulate the real network environment.
[0082] S4, Feature fusion: In the feature fusion module, fuse the features extracted from the structured feature channel and the image feature channel. For the structured feature channel, extract both packet-level features, such as source IP, destination IP, protocol type, and packet length, and data stream-level features, such as flow duration, total number of packets, and total number of bytes, and combine the two to construct a comprehensive feature representation for each packet; for the image feature channel, convert the byte stream into a grayscale image and then extract local features by a convolutional neural network. Finally, merge the features of the two channels into a unified high-dimensional feature vector to simultaneously obtain the global characteristics and local details of the traffic data.
[0083] S5, Gradient sharing initialization: Before model training, initialize the parameters related to the gradient sharing mechanism, including determining the formula for calculating the gradient weights of the convolutional neural network and KAN channels, initializing the KAN and convolutional neural network models, and the optimizer.
[0084] S6, Gradient Sharing Training: During the training process, for each batch of training data, the structured features and image features are respectively input into KAN and the convolutional neural network model, and their respective outputs and losses are calculated. Based on the losses, the gradients of the convolutional neural network and KAN channels themselves, as well as the gradient guidance between them, are calculated. The final updated gradients of each channel are calculated by dynamically adjusting the weights.
[0085] S7, Model Training and Optimization: Through multiple rounds of training and using the gradient sharing mechanism, KAN and the convolutional neural network model cooperate with each other and co-optimize during the feature learning process, continuously adjusting the model parameters to improve the model's ability to identify malicious traffic, reduce the conflict between local and global feature learning, and enhance the model's sensitivity to diverse malicious traffic patterns.
[0086] S8, Model Testing: The trained malicious traffic classification model is tested on the test set, and the model performance is evaluated through indicators such as accuracy, precision, recall, and F1 score to verify the classification accuracy of the model for malicious traffic in different 5G Internet of Things traffic scenarios.
[0087] Furthermore, S1 specifically includes: Processing the collected original 5G Internet of Things traffic dataset, converting it into a format suitable for model input, and constructing structured feature channel and image feature channel data. For the construction of structured feature channel data: First, the PCAP files in the 5G traffic data are parsed into vector format using the Wireshark tool, and rich communication features such as source IP, destination IP, protocol type, packet length, flow duration, total number of packets, and total number of bytes during the communication process are extracted. These features cover information at the packet level and data flow level, comprehensively reflecting the static attributes and dynamic behaviors of network communication. Subsequently, all the extracted features are standardized. Assuming X is the current value, X max is the maximum value of this feature, and X min is the minimum value of this feature. The normalization formula is:
[0088]
[0089] Unify features of different magnitudes and distributions into the same scale range, avoid inconsistent influence weights of different features on model learning, and ensure that the model can learn feature information more effectively.
[0090] For the construction of image feature channel data, the present invention transforms the byte distribution of 5G traffic into an image form. The specific operation is to reorganize the original byte stream in the PCAP file into a one-dimensional vector of a fixed length, and instead of normalizing it, directly perform binary processing on it to obtain a series of vectors consisting only of 0s and 1s. Select the first 784 bits from them and map them into a 28×28 grayscale image, where each pixel value corresponds to a byte value in the vector, so that the subsequent convolutional neural network module can extract the local spatial characteristics of the traffic from it.
[0091] Further, S2 specifically includes: using a local feature extraction module based on a convolutional neural network to process the data in the image feature channel. The convolutional neural network module is composed of two convolutional layers and a pooling layer stacked. In the convolutional layer, convolutional kernels of different sizes and strides slide and convolve on the image to automatically learn local spatial characteristics in the traffic data, such as microscopic pattern features like packet length, time interval, and protocol identifier. In the pooling layer, max pooling or average pooling operations are used to downsample the feature map output by the convolutional layer, reducing the data volume while retaining key features, reducing the model's computational complexity, and improving the model's training efficiency and generalization ability.
[0092] Further, S3 specifically includes: constructing a global feature extraction module with KAN to dynamically model the data in the structured feature channel. The KAN module receives the preprocessed structured feature vectors and dynamically captures the time series distribution and global dependencies of the traffic through its unique network structure. When constructing global features, instead of inputting all the data at once, the samples are input step by step in chronological order to simulate the dynamic change process of data in a real network environment, constructing the overall behavior characteristics of the traffic from a macroscopic level, capturing macroscopic information such as the time series characteristics of the traffic and the network topology structure, which complements the local features extracted by the convolutional neural network.
[0093] Further, S4 specifically includes: In the feature fusion module, the features extracted from the structured feature channel and the image feature channel are fused. For the structured features, at the data packet level, features such as source IP, destination IP, protocol type, and data packet length are extracted, which provide static information of the traffic samples; at the data flow level, traffic samples with the same five-tuple (source IP, destination IP, source port, destination port, protocol) are grouped as a data flow, and features such as flow duration, total number of data packets, and total number of bytes are extracted. These features reflect the spatio-temporal correlation between data packets and the overall behavior characteristics of the data flow from a dynamic perspective. Then, the features at the data flow level are added after the features of the belonging data packet to construct a comprehensive feature representation for each data packet, so that it contains both its own attributes and the context information of the belonging data flow. For the image features, after the feature extraction is completed by the convolutional neural network module, the features of the two channels are concatenated into a unified high-dimensional feature vector in a certain order, so that the model can simultaneously obtain the global characteristics and local details of the traffic data and improve the recognition ability of malicious traffic.
[0094] Further, S5 specifically includes: Before model training, initialize the parameters related to the gradient sharing mechanism. Determine the formulas for calculating the gradient weights of the image feature channel and the structured feature channel. Assume ω Local is the weight of the image feature channel, ω Global is the weight of the structured feature channel, L Local is the loss of the image feature channel, L Global is the loss of the structured feature channel. Then the weight formulas for the two channels are respectively:
[0095]
[0096] Further, S6 specifically includes: During the training process, for each batch of training data, input the structured features and image features into the KAN and the convolutional neural network model respectively. The KAN model outputs the prediction results for the global features, and the convolutional neural network model outputs the prediction results for the local features. Calculate the loss values between them and the true labels respectively to obtain the gradients G Local,self and G Global,self of the convolutional neural network and the KAN channel itself. Then, according to the gradient sharing mechanism, calculate the gradient guidance G Global→Local of the structured feature channel to the image feature channel and the gradient guidance G Local→Global of the image feature channel to the structured feature channel. Through the previously calculated weights ω Local and ω Global , the gradient calculation formulas for the two channels are as follows:
[0097] G Local =ω Local GLocal,self +(1 - ω Local )G Global,self (4)
[0098] G Global =ω Local G Local,self +(1 - ω Local )G Global,self (5)
[0099] Finally, update the parameters of the image feature channel and the structured feature channel models according to the updated gradients respectively, so that the two models cooperate with each other and optimize synergistically during the training process, improving the model's ability to identify malicious traffic.
[0100] Furthermore, S7 specifically includes: through multiple rounds of training, using the gradient sharing mechanism, enabling KAN and the convolutional neural network model to continuously cooperate with each other during the feature learning process. In each round of training, the two models continuously adjust their own parameters according to the shared gradient information to optimize the learning effect of malicious traffic features. As the number of training rounds increases, the model gradually reduces the conflict between local and global feature learning and enhances the sensitivity to diverse malicious traffic patterns. During the training process, observe the changes in indicators such as the accuracy rate and loss value of the model on the validation set. When the performance indicators of the model on the validation set tend to be stable and there is no obvious improvement, it is considered that the model training reaches a good state, completing the model training and optimization process, and obtaining a trained malicious traffic classification model.
[0101] Furthermore, S8 specifically includes: testing the trained malicious traffic classification model on the test set. In the testing stage, construct the structured feature channel and image feature channel data for the 5G Internet of Things traffic data in the test set according to the preprocessing method in the training stage, and input them into the trained model. The model infers the test data according to the learned feature patterns and outputs the prediction results of whether each traffic sample is malicious traffic.
[0102] The present invention can effectively make up for the deficiencies of traditional methods in 5G Internet of Things malicious traffic detection, such as incomplete feature extraction and poor adaptability to complex traffic patterns. At the same time, it enhances the ability to identify unknown malicious traffic patterns through gradient sharing and feature fusion, reduces the dependence on large-scale labeled data, and provides an innovative solution for efficient and accurate malicious traffic detection in the 5G Internet of Things environment.
[0103] The 5G Internet of Things malicious traffic classification method based on feature fusion and gradient sharing provided by the embodiments of the present invention involves preprocessing 5G traffic data, training local and global feature extraction modules, then feature fusion and gradient sharing training, and finally testing on the test set. The signal and data processing processes in each step will be explained in detail below:
[0104] S101, Data Conversion and Partitioning: First, construct a structured feature channel and an image feature channel for 5G traffic data respectively. For the structured feature channel, use Wireshark to parse the PCAP files in the 5G traffic data into vector format, extract communication features such as source IP, destination IP, protocol type, packet length, etc., and perform standardization processing; for the image feature channel, map the byte distribution of 5G traffic into image format, reorganize the original byte stream in the PCAP file into a fixed-length one-dimensional vector, and after binaryization, select the first 784 bits and map them into a 28×28 grayscale image. Then, divide the processed dataset into three parts: a normal sample training set, a mixed sample training set, and a mixed sample test set to ensure that the model performance can be effectively evaluated during the training and testing phases.
[0105] S102, Training Normal Local Patterns: Use the local feature extraction module based on CNN to conduct training on the normal sample training set. This module processes the grayscale images in the image feature channel by stacking convolutional layers and pooling layers, automatically learns microscopic pattern features such as packet length and time interval, so as to accurately grasp the local feature patterns of normal traffic.
[0106] S103, Unsupervised Collaborative Learning: On the mixed sample training set, the global feature extraction module based on KAN and the local feature extraction module based on CNN jointly model the 5G traffic data by processing data from different feature channels. The former dynamically models the structured features to capture the time series distribution and global dependencies of the traffic; the latter continuously extracts local microscopic features. The two modules learn in parallel, respectively capturing normal and abnormal patterns to better understand the potential complex relationships in the data.
[0107] S104, Initial Unidirectional Guidance: At the initial stage of training the mixed sample training set, the local feature extraction module based on CNN plays a guiding role, providing the local feature information it has learned to the global feature extraction module based on KAN to help the latter better carry out learning and enable it to initially understand the local feature manifestations in normal and abnormal traffic patterns.
[0108] S105, Gradient Sharing: As the training progresses, when the performance of the global feature extraction module based on KAN gradually approaches 70% of that of the local feature extraction module based on CNN, the two modules enter the collaborative optimization stage. In this stage, the two modules cooperate with each other, exchange information, and jointly adjust parameters to further improve the learning efficiency and accuracy of the model.
[0109] S106, Collaborative Optimization: In the collaborative optimization stage, the local feature extraction module based on the convolutional neural network and the global feature extraction module based on KAN not only share the features they have learned, but also further adjust the parameters by sharing gradients to enhance each other's learning effects. This process can enhance the generalization ability of the model, enabling it to adapt to more complex and variable 5G traffic data.
[0110] S107, Model Consistency: After multiple rounds of collaborative optimization, the performance of the local feature extraction module based on the convolutional neural network and the global feature extraction module based on KAN gradually stabilizes and complements each other, enabling them to identify malicious traffic patterns in a wider range of and varying 5G traffic data. At this point, the model basically reduces its dependence on a large amount of labeled data and can effectively process 5G traffic data with different features and types.
[0111] S108, Final Prediction: After completing multiple rounds of collaborative optimization, apply the trained model to the data in the mixed sample test set. Through this test set, the model can predict new 5G traffic samples to determine whether they belong to the malicious traffic category, thus providing accurate results for subsequent malicious traffic detection.
[0112] Throughout the process, each step involves the processing and transformation of data to meet the requirements of subsequent steps. At the same time, by combining graph convolutional networks, graph attention networks, and collaborative learning, this method can more comprehensively extract abnormal patterns in log data and improve the accuracy of abnormal traffic detection.
[0113] Furthermore, S101 specifically includes: converting the original log data set into a directed graph and dividing it into a normal sample training set, a mixed sample training set, and a mixed sample test set. Among them, the normal sample training set only contains normal log data, the mixed sample training set contains a certain proportion of abnormal logs, and the mixed sample test set serves as the final evaluation data set. In data preprocessing, perform deduplication, format normalization, timestamp conversion, and feature vectorization on the log data to ensure that the input data is suitable for graph neural network modeling. The directed graph after converting the log data can be represented as a heterogeneous graph G=(V, E), where V is the node set and E is the edge set.
[0114] Furthermore, S102 specifically includes: using a graph convolutional network for pre-training on the normal sample training set to learn the structural patterns and feature relationships between normal logs. Assume H (l) represents the node representation matrix of the l-th layer, A is the adjacency matrix of the directed graph, w (l) is the weight matrix of the l-th layer, and σ is the activation function. The graph convolutional network represents the nodes through the following formula:
[0115] H (l+1) = σ(AH (l) W (l) ) (6)
[0116] Further, S103 specifically includes: On the mixed sample training set, the graph attention network and the graph convolutional network perform unsupervised learning simultaneously. The two models jointly learn the normal patterns and abnormal patterns in the mixed samples through the graph structure. At this stage, the graph attention network calculates the attention weights between adjacent nodes, performs weighted aggregation on the neighbors of each node, and learns the interdependent relationships between nodes. Suppose there is a graph containing N nodes, and the feature vector of each node is represented as h = {h 1 , h 2 ,..., h N}, where represents the feature vector of node i, and F is the feature dimension. The graph attention network calculates the attention coefficient e ij between adjacent nodes i and j, and the formula is as follows:
[0117] e ij = LeakyReLU(a T [Wh i || Wh j ) (7)
[0118] Further, S104 specifically includes: In the initial stage of the mixed sample training set, the graph convolutional network acts as a teacher network to guide the training of the graph attention network. At this stage, the learning samples of the graph attention network depend on the output of the graph convolutional network, and the prediction result of the graph convolutional network is used as the true label of the graph attention network. During the training process, the loss functions of the graph attention network and the graph convolutional network are combined, and the training of the graph attention network is accelerated through collaborative learning, enabling it to better learn abnormal patterns. Let y be the predicted value of the graph convolutional network, be the predicted value of the graph attention network, and σ be the sigmod function. The training loss function is as follows:
[0119]
[0120] Further, S105 specifically includes: When the prediction accuracy of the graph attention network reaches 70% of the prediction accuracy of the graph convolutional network, the two models start to enter the collaborative learning stage. At this stage, the graph attention network and the graph convolutional network guide each other and alternately act as the teacher and student roles. The two models share their respective feature information and are optimized through gradient sharing. At this time, the graph attention network can be optimized based on the output of the graph convolutional network, and at the same time, the graph convolutional network will also refer to the prediction result of the graph attention network for adjustment to improve the overall detection performance.
[0121] Furthermore, S106 specifically includes: in the collaborative learning process, the graph attention network and the graph convolution network transfer learned knowledge and shared gradients to each other by sharing and fusing feature information. Through this knowledge sharing mechanism, the two models can continuously improve each other, enhance the generalization ability and adaptability of the models, and thus better adapt to a variety of complex and dynamic log data.
[0122] Furthermore, S107 specifically includes: through multiple rounds of collaborative learning, the two models gradually converge and can identify abnormal patterns in a wider range of log data. At this point, the two models have eliminated the dependence on large-scale labeled data to a certain extent, and can continue to effectively identify abnormal patterns under different log formats and changes, improving the robustness of the model. By iterating the loss value and gradient after each iteration, the accuracy of the entire model is improved, and ultimately a good prediction can be made for each sample.
[0123] Further, S108 specifically includes: Finally, the trained model is predicted on the mixed sample test set. In the test phase, the model infers based on the sample features in the test set and outputs a prediction result of whether each sample is abnormal. The test results will be evaluated based on indicators such as accuracy, recall, and F1 score to verify the actual performance and effectiveness of the model.
[0124] like Figure 3 As shown, the abnormal traffic classification model based on feature fusion and gradient sharing provided by the embodiment of the present invention includes:
[0125] Data preprocessing module, used to process 5G traffic data into structured and graphical features suitable for model input;
[0126] The local feature extraction module extracts local microscopic features of traffic data based on convolutional neural networks;
[0127] The global feature extraction module uses KAN to construct the global behavior characteristics of traffic;
[0128] Feature fusion module: used to fuse the features extracted by the structured feature channel and the image feature channel;
[0129] Gradient sharing module: used to achieve mutual complementation and collaborative optimization of the two modules in feature expression, thereby improving the training stability and convergence speed of the model.
[0130] The method of the present invention provides an effective solution for 5G Internet of Things malicious traffic detection by fusing different types of features and using a gradient sharing mechanism.
[0131] Example 1: Malicious traffic classification detection in 5G IoT environment
[0132] 1. Data preprocessing: For 5G traffic data, construct a structured feature channel and an image feature channel. When constructing the structured feature channel, use Wireshark to parse the PCAP files in the 5G traffic data into vector format, extract communication features such as source IP, destination IP, protocol type, packet length, etc., and then perform standardization processing on these features to make the influence of different features more balanced in model learning. In terms of constructing the image feature channel, convert the byte distribution of 5G traffic into an image form, organize the original byte stream in the PCAP file into a one-dimensional vector of fixed length, after binary processing, select the first 784 bits and map them into a 28×28 grayscale image to facilitate the subsequent CNN module to extract local features.
[0133] 2. Training of local and global feature extraction modules: On the prepared dataset, train the local feature extraction module based on convolutional neural network and the global feature extraction module based on KAN respectively. The convolutional neural network module processes the grayscale image in the image feature channel by stacking convolutional layers and pooling layers, and automatically learns microscopic pattern features such as packet length and time interval. The KAN module dynamically models the vector data in the structured feature channel, captures the time series distribution and global dependencies of the traffic, and constructs the overall behavior features of the traffic. During the training process, continuously adjust the model parameters to enable the two modules to better extract the features they are responsible for.
[0134] 3. Feature fusion and gradient sharing training: In the feature fusion module, fuse the features extracted from the structured feature channel and the image feature channel. For the structured features, integrate the features at the packet level and the data stream level to construct a comprehensive feature representation for each packet. Then merge the features of the two channels into a unified high-dimensional feature vector. During training, start the gradient sharing mechanism. In the initial stage of training, each module learns features based on its own advantages. As the training progresses, different modules dynamically adjust their own parameters by sharing gradient information to achieve collaborative optimization and improve the model's learning ability for malicious traffic patterns.
[0135] 4. Model evaluation and detection: After multiple rounds of training, obtain the trained model. Use this model to detect the 5G IoT traffic data in the test set. The model infers the traffic samples based on the learned feature patterns to determine whether each sample is malicious traffic. Use indicators such as accuracy, precision, recall, and F1 score to evaluate the model performance. Continuously optimize the model according to the evaluation results to improve its accuracy and robustness in detecting malicious traffic in the 5G IoT environment.
[0136] This embodiment demonstrates how to efficiently perform malicious traffic classification and detection in the 5G Internet of Things (IoT) environment by leveraging feature fusion and gradient sharing techniques. Through these techniques, the model can fully learn the local and global features of traffic data, adapt to the complex and changing traffic patterns in the 5G IoT environment, and improve the detection accuracy and the generalization ability of the model.
[0137] A 5G IoT malicious traffic detection method based on feature fusion and gradient sharing, characterized in that a gradient sharing mechanism is introduced during the model training process to improve the detection performance. The method includes the following steps:
[0138] Step 1: Load multiple publicly available 5G-IoT traffic datasets, such as TON IoT, ISCXTor2016, and 5G-NIDD, and preprocess the original traffic data.
[0139] Step 2: Construct a structured feature channel and an image feature channel for the 5G traffic data respectively. For the structured feature channel, use Wireshark to parse the PCAP files in the 5G traffic data into vector format, extract communication features, and perform normalization processing; for the image feature channel, map the byte distribution of the 5G traffic into image format, reorganize the original byte stream in the PCAP file into a fixed-length one-dimensional vector, and after binarization, select the first 784 bits and map them into a grayscale image.
[0140] Step 3: Divide the processed dataset into three parts: a normal sample training set, a mixed sample training set, and a mixed sample test set to ensure that the model has good generalization ability during training and testing.
[0141] Step 4: Design a malicious traffic detection model based on feature fusion and gradient sharing, using a convolutional neural network and KAN as the core modules, and introduce an attention mechanism during the feature fusion process to enhance the expression ability of malicious traffic features.
[0142] Step 5: Initialize the convolutional neural network and the KAN model respectively. On the normal sample training set, the convolutional neural network learns the local microscopic features in the traffic data by stacking convolutional layers and pooling layers, and KAN learns the global features of the traffic through dynamic modeling to ensure that both models can learn normal traffic patterns.
[0143] Step 6: Perform model training on the mixed sample training set, introduce the gradient sharing mechanism, so that the convolutional neural network and KAN can dynamically and real-time share gradients and jointly optimize the loss function during the training process, realize the mutual guidance of the two modules, and optimize the detection performance.
[0144] Step 7: During the training process, dynamically adjust the weights of the convolutional neural network and KAN gradient sharing according to the model training situation, so that the model can be better co-optimized at different training stages.
[0145] Step 8: After the training is completed, conduct model evaluation on the mixed sample test set, calculate performance metrics such as accuracy, precision, recall, and F1-score, and compare with other 5G Internet of Things malicious traffic detection methods.
[0146] Step 9: Use the adversarial sample generation method to test the robustness of the model, analyze the performance of the model when facing adversarial attacks, and optimize the model structure to enhance its detection ability for unknown malicious traffic patterns.
[0147] In the final experimental results, the proposed malicious traffic detection model based on feature fusion and gradient sharing performs excellently on multiple public datasets, achieving an accuracy of 98.2% and an F1-score of 98.9% on the 5G-NIDD dataset, and has better detection performance compared to other methods, as shown in Table 1.
[0148]
[0149] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable logic devices such as field programmable gate arrays, or can be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software, such as firmware.
[0150] The above is only the specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any modification, equivalent replacement, and improvement made within the spirit and principle of the present invention by those skilled in the art within the technical scope disclosed by the present invention shall be covered by the protection scope of the present invention.
Claims
1. A malicious traffic classification method based on gradient sharing feature fusion network, characterized in that: The following steps are involved: S1, data preprocessing: preprocess the original data of 5G IoT traffic to construct structured feature channels and image feature channels respectively; For the structured feature channel, Wireshark is used to parse the PCAP file in the 5G traffic data into a vector format, extract the communication features, and perform standardization processing. For the image feature channel, the byte distribution of the 5G traffic is mapped into an image format, and the original byte stream in the PCAP file is reorganized into a one-dimensional vector of a fixed length. After binarization, the first 784 bits are selected and mapped into a grayscale image. S2, local feature extraction: using the local feature extraction module based on convolutional neural network to process the data in the image feature channel, automatically learn the local spatial characteristics in the traffic data by stacking convolutional layers and pooling layers, and extract micro-pattern features, including packet length, time interval and protocol identification; S3, global feature extraction: A global feature extraction module is constructed through KAN to dynamically model the data in the structured feature channel, capture the time series distribution and global dependency of traffic, and construct the overall behavior characteristics of traffic from a macro level; When constructing global features, samples are input step by step in time order to simulate the real network environment; S4, feature fusion: In the feature fusion module, the features extracted from the structured feature channel and the image feature channel are fused; The structured feature channel extracts packet-level features, including source IP, destination IP, protocol type, and packet length, and simultaneously extracts data stream-level features, including flow duration, total number of packets, and total number of bytes. The packet-level features and data stream-level features are combined to construct a comprehensive feature representation for each packet. For the image feature channel, the byte stream is converted into a grayscale image and then a convolutional neural network is used to extract local features. Finally, the features of the two channels are merged into a unified high-dimensional feature vector to simultaneously obtain the global characteristics and local details of the traffic data. S5, gradient sharing initialization: before model training, initialize the parameters related to the gradient sharing mechanism, including determining the formula for calculating the gradient weights of the convolutional neural network and KAN channels, and initializing the KAN and convolutional neural network models and optimizers; S6, gradient sharing training: During the training process, for each batch of training data, the structured features and image features are input into the KAN and convolutional neural network models respectively, and their respective outputs and losses are calculated; the gradients of the convolutional neural network and KAN channels themselves are calculated based on the losses, as well as the gradient guidance between each other; the final update gradient of each channel is calculated by dynamically adjusting the weights; S7, model training and optimization: Through multiple rounds of training and using the gradient sharing mechanism, the KAN and convolutional neural network models can collaborate and optimize with each other in the feature learning process, continuously adjust model parameters, improve the model's ability to identify malicious traffic, reduce the conflict between local and global feature learning, and enhance the model's sensitivity to diverse malicious traffic patterns; S8, model testing: The trained malicious traffic classification model is tested on the test set, and the model performance is evaluated by accuracy, precision, recall and F1 score indicators to verify the accuracy of the model in classifying malicious traffic in different 5G IoT traffic scenarios.
2. The malicious traffic classification method based on gradient sharing feature fusion network as claimed in claim 1 is characterized in that: In step S1: For the construction of structured feature channel data: first, the PCAP file in the 5G traffic data is parsed into a vector format using the Wireshark tool to extract the communication features such as source IP, destination IP, protocol type, packet length, flow duration, total number of packets, and total number of bytes in the communication process. These communication features cover information at the packet level and data flow level, and fully reflect the static properties and dynamic behaviors of network communication; then, all the extracted features are standardized, assuming that X is the current value, X max is the maximum value of this feature, X min is the minimum value of this feature, and the normalization formula is: Unify features of different magnitudes and distributions into the same scale range to avoid inconsistent weights of different features on model learning, ensuring that the model can learn feature information more effectively; For the construction of image feature channel data, the byte distribution of 5G traffic is converted into image form; the specific operation is: the original byte stream in the PCAP file is reorganized into a one-dimensional vector of fixed length, and it is directly binarized without normalization to obtain a series of vectors consisting of only 0 and 1; the first 784 bits are selected and mapped into a 28×28 grayscale image, where each pixel value corresponds to a byte value in the vector, so that the subsequent convolutional neural network module can extract the local spatial characteristics of the traffic from it.
3. The malicious traffic classification method based on gradient sharing feature fusion network as described in claim 1 is characterized in that: In the step S2: the convolutional neural network module is composed of two convolutional layers and a pooling layer stacked together; in the convolutional layer, sliding convolution is performed on the image through convolution kernels of different sizes and steps to automatically learn local spatial characteristics in traffic data, including micro-pattern features such as data packet length, time interval and protocol identification; in the pooling layer, maximum pooling or average pooling operation is used to downsample the feature map output by the convolutional layer, thereby reducing the amount of data while retaining key features, reducing the model calculation complexity, and improving the model training efficiency and generalization ability.
4. The malicious traffic classification method based on gradient sharing feature fusion network as claimed in claim 1 is characterized in that: In step S3: the KAN module receives the preprocessed structured feature vector, and dynamically captures the time series distribution and global dependency of the traffic through its unique network structure; when constructing the global features, the sample is input step by step in time sequence to simulate the dynamic change process of data in a real network environment, and the overall behavior characteristics of the traffic are constructed from a macro level to capture macro information, including the time series characteristics of the traffic and the network topology, which complements the local features extracted by the convolutional neural network.
5. The malicious traffic classification method based on gradient sharing feature fusion network as claimed in claim 1 is characterized in that: The step S4 specifically includes: in a feature fusion module, fusing the features extracted by the structured feature channel and the image feature channel; For structured features, at the packet level, the source IP, destination IP, protocol type, and packet length features are extracted. These features provide static information of traffic samples. At the data flow level, traffic samples with the same five-tuple are grouped as a data flow. The five-tuple includes source IP, destination IP, source port, destination port, and protocol. The flow duration, total number of packets, and total number of bytes are extracted. These features reflect the spatiotemporal association between packets and the overall behavioral characteristics of the data flow from a dynamic perspective. Then, the features at the data flow level are added to the features of the data packets to which they belong, and a comprehensive feature representation is constructed for each data packet, so that it contains both its own attributes and the contextual information of the data flow to which it belongs. For image features, after the convolutional neural network module completes feature extraction, the features of the two channels are spliced into a unified high-dimensional feature vector in a certain order, so that the model can simultaneously obtain the global characteristics and local details of the traffic data, and improve the ability to identify malicious traffic.
6. The malicious traffic classification method based on gradient sharing feature fusion network as claimed in claim 1 is characterized in that: The step S5 specifically includes: determining a formula for calculating the gradient weights of the image feature channel and the structural feature channel, assuming that ω Local is the weight of the image feature channel, ω Global is the weight of the structured feature channel, L Local is the loss of the image feature channel, L Global is the loss of the structured feature channel, and the weight formulas of the two channels are:
7. The malicious traffic classification method based on gradient sharing feature fusion network as claimed in claim 6 is characterized in that: The step S6 specifically includes: during the training process, for each batch of training data, the structured features and image features are respectively input into the KAN and convolutional neural network models; the KAN model outputs the prediction results of the global features, and the convolutional neural network model outputs the prediction results of the local features, and the loss values between them and the true labels are respectively calculated to obtain the gradient G of the convolutional neural network and the KAN channel itself. Local,self and G Global,self ; Then, according to the gradient sharing mechanism, the gradient guidance G of the structured feature channel to the image feature channel is calculated Global→Local And the gradient guidance G of the image feature channel to the structured feature channel Local→Global ; The weight ω obtained by calculation Local and ω Global , the gradient calculation formula of the two channels is as follows: G Local =ω Local G Local,self +(1-ω Local )G Global,self (4) G Global =ω Local G Local,self +(1-ω Local )G Global,self (5) Finally, the parameters of the image feature channel and structured feature channel models are updated respectively according to the updated gradient, so that the two models can cooperate and optimize with each other during the training process, thereby improving the model's ability to identify malicious traffic.
8. The malicious traffic classification method based on gradient sharing feature fusion network as claimed in claim 1 is characterized in that: The step S7 specifically includes: in each round of training, the two models continuously adjust their own parameters according to the shared gradient information to optimize the learning effect of malicious traffic features; as the number of training rounds increases, the model gradually reduces the conflict between local and global feature learning and enhances the sensitivity to diversified malicious traffic patterns; during the training process, the changes in the accuracy and loss value indicators of the model on the verification set are observed. When the performance indicators of the model on the verification set tend to be stable and no longer have a significant improvement, the model training and optimization process is completed to obtain a trained malicious traffic classification model.
9. The malicious traffic classification method based on gradient sharing feature fusion network as claimed in claim 1 is characterized in that: The step S8 specifically includes: in the testing phase, the 5G IoT traffic data in the test set is preprocessed in the training phase to construct structured feature channels and image feature channel data, respectively, and input into the trained model; the model infers the test data according to the learned feature pattern, and outputs a prediction result of whether each traffic sample is malicious traffic.
10. A malicious traffic classification system based on gradient sharing feature fusion network according to any one of the methods of claims 1 to 9, characterized in that: include: Data preprocessing module, used to process 5G traffic data into structured and graphical features suitable for model input; The local feature extraction module extracts local microscopic features of traffic data based on convolutional neural networks; The global feature extraction module uses KAN to construct the global behavior characteristics of traffic; Feature fusion module: used to fuse the features extracted by the structured feature channel and the image feature channel; Gradient sharing module: used to achieve mutual complementation and collaborative optimization of the two modules in feature expression, thereby improving the training stability and convergence speed of the model.
Citation Information
Patent Citations
Malicious traffic detection method based on multilevel feature fusion
CN118157929A
Industrial internet multi-instance sequential network flow data processing method
CN118631501A
Malicious traffic classification method and device based on big data, equipment and medium
CN118631510A
Malicious traffic classification and identification method, system and device based on graph convolutional neural network, and medium
CN119172143A
Efficient skin disease segmentation method based on multi-scale and mixed attention mechanism
CN119228824A