Efficient intelligent network attack classification tool based on federated learning framework
Through the federated learning framework, combining multiple classifiers and D-S evidence theory optimization algorithms, the adaptability and insufficient computing resources of the network attack classification method in the face of unknown attacks and changing environments is solved, and efficient and accurate network attack detection and adaptive classification are achieved.
Patent Information
- Application Number
- CN202510462099.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
The existing cyber attack classification methods are poor in adaptability when facing unknown attack vectors, changing network environments and new attack types, and are difficult to effectively identify and classify, and lack computing resources and real-time performance.
An efficient and intelligent network attack classification tool based on the federated learning framework is adopted, combining multi-task logistic regression classifier, random forest classifier, XGBoost classifier and long-term short-term memory neural network classifier, and a D-S evidence theory optimization algorithm is used to train models and update parameters to achieve data privacy protection and efficient classification.
It improves the accuracy and robustness of cyber attack detection, is adaptable, can quickly respond to new attacks, ensure data privacy, reduce computing complexity, and adapt to different network environments and attack scenarios.
Smart Images

Figure CN120296476A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and more particularly to an efficient intelligent network attack classification tool based on a federated learning framework. Background Art
[0002] The main drawback of current network attack classification methods lies in their poor adaptability. Especially when facing unknown attack vectors, ever-changing network environments, and newly emerging attack types, existing methods cannot flexibly and effectively adjust, resulting in insufficient recognition ability for attack scenarios.
[0003] For example: Existing machine learning classifiers are usually optimized according to specific data categories, but this approach has significant limitations in dealing with complex, multi-dimensional data environments. For example, traditional methods such as support vector machines (SVM), decision trees, and k-nearest neighbors (KNN) are difficult to adapt to new attack types or changing patterns when facing highly uncertain and fluctuating network attacks; especially in network attack detection tasks, there is often a contradiction between the classification accuracy and speed of these models, making their recognition ability for new attacks insufficient.
[0004] Deep learning methods, such as convolutional neural networks (CNN) and recurrent neural networks (RNN), can capture complex patterns and temporal dependencies in data, so they have significant advantages in certain tasks; however, these methods usually require a large amount of computing resources and a long training time, which makes it difficult for them to be deployed in real-time network environments with limited resources and are vulnerable to overfitting, especially when the amount of data is insufficient.
[0005] Hybrid methods aim to integrate multiple paradigms to improve detection and can enhance the accuracy of the system; but in practical applications, these hybrid methods often face problems such as high computational complexity and poor real-time performance, especially when facing large-scale, dynamically changing network traffic.
[0006] Traditional statistical methods, such as K-Means, although perform well in identifying operation biases and abnormal behaviors, they usually rely on assuming a specific form of data distribution and are difficult to handle highly non-linear and high-dimensional data.
[0007] In addition, the dependence of many detection methods on computationally intensive algorithms makes them face challenges of insufficient computing resources and slow processing speed in practical applications, especially when the attack pattern is complex and involves multiple data dimensions.
[0008] Therefore, how to provide a more robust and efficient network attack classification tool that can enhance adaptability, flexibly and effectively adjust, and improve the recognition ability of attack scenarios when facing unknown attack vectors, ever-changing network environments, and newly emerging attack types is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0009] In view of this, the present invention provides an efficient and intelligent network attack classification tool and system based on a federated learning framework to solve some of the technical problems mentioned in the background art.
[0010] To achieve the above object, the present invention adopts the following technical solutions:
[0011] An efficient and intelligent network attack classification tool based on a federated learning framework includes a central server and multiple edge nodes. Each edge node deploys a local model, and the central server deploys a central model. Both the local model and the central model include a multi-task logistic regression classifier, a random forest classifier, an XGBoost classifier, a long short-term memory neural network classifier, and a D-S evidence theory optimization algorithm;
[0012] The local model collects local network traffic data, combines the D-S evidence theory optimization algorithm and four types of complementary classifiers for model training, and performs abnormal data detection and classification, and transmits the key update parameters related to the abnormality of the encrypted abnormal data to the central server;
[0013] The central model receives the encrypted update parameters and aggregates them. After updating the global model, it performs attack type classification based on the encrypted abnormal data combined with the D-S evidence theory optimization algorithm, and redistributes the classification results and the updated global model to each edge node.
[0014] Preferably, the edge nodes include a user side, a server side, a database side, a network side, and an application side.
[0015] Preferably, the specific content of combining the D-S evidence theory optimization algorithm and four types of complementary classifiers for model training is as follows:
[0016] Input the network traffic data set, and the D-S evidence theory optimization algorithm assigns initial confidence levels to the four classifiers and enters the four classifiers in parallel for classification tasks. Among them, the local model task is abnormal detection, and the central model task is attack classification;
[0017] After the classification results are output, evidence combination is performed according to the Dempster combination rule to integrate the classification predictions from the multi-task logistic regression classifier, the random forest classifier, the XGBoost classifier, and the long short-term memory neural network classifier, and at the same time consider the consistency and conflict between the classification predictions;
[0018] Aggregate the confidence values on the frame of discernment through the combined basic probability assignment, derive the confidence function of the decision, and finally select the hypothesis with the highest confidence value to make the final classification decision based on the threshold criterion.
[0019] Preferably, the specific content of the D-S evidence theory optimization algorithm is as follows:
[0020] Improve the dynamic update of the confidence quality and adjust the degree of trust in the outputs of each classifier according to the historical performance of the classifier;
[0021] Introduce context-aware evidence combination and allocate confidence quality for different attack patterns;
[0022] Weighted conflict resolution: assign a weight to each classifier, prioritize according to the reliability of the classifier for a specific anomaly, and calculate the confidence quality through weighted combination;
[0023] Integrate time-evidence aggregation, dynamically update the confidence quality based on the time-series trend, according to historical evidence and newly emerging data, to adapt to new attack patterns.
[0024] Preferably, the specific method for improving the dynamic update of the confidence quality is as follows:
[0025] Update the confidence of each hypothesis A according to the historical performance of the classifier:
[0026]
[0027] where ω classifier (A) represents the reliability of the classifier predicting hypothesis A over time, hypothesis A is a possible attack type, B is a subset of A, and m is the confidence quality;
[0028] For any initial confidence quality distribution, through the sequence of update rules m k will converge:
[0029]
[0030] where m k (A) is the k-th confidence quality of hypothesis A, m * (A) is the updated confidence quality, and the stability of the confidence quality assignment is guaranteed by the gradual decrease of entropy. With each iteration, the entropy will gradually decrease, thus guiding the system towards the final convergence state:
[0031]
[0032] where H(m) is the entropy, Θ is the frame of discernment, representing all possible attack types.
[0033] Preferably, the weight ω assigned to each classifier i is:
[0034]
[0035] where ω i represents the relative importance of the i-th classifier C i among all classifiers. The numerator represents the accuracy of the i-th classifier, and the denominator represents the sum of the accuracies of all classifiers from the 1st classifier to the nth classifier;
[0036] The final confidence quality is calculated through weighted combination:
[0037]
[0038] Preferably, after integrating the time evidence aggregation, the dynamically updated confidence quality is:
[0039] m t m(A)=(1 - α)·m t-Δt (A)+α·m new (A) where α is a balancing factor used to control the ratio between historical evidence and new evidence, and t represents the time instant.
[0040] Preferably, the global model parameters aggregated by the central server are:
[0041]
[0042] where represents the data contribution weight of each node;
[0043]
[0044] where L i is the loss function minimized in the local training process, and θ′ i is the optimized parameter transferred from the local model to the central model.
[0045] Preferably, before model training, four types of complementary classifiers selected for the local model and the central model, namely, the multi-task logistic regression classifier, the random forest classifier, the XGBoost classifier, and the long short-term memory neural network classifier, are respectively optimized and designed.
[0046] Preferably, the optimization of the multi-task logistic regression classifier is specifically as follows: capturing common attack patterns by sharing the parameter matrix while retaining task-specific parameters; adopting adversarial training and incorporating clean samples and adversarial samples into the loss function; introducing local adaptive kernel logistic regression and using the kernel method to capture complex attack patterns in the high-dimensional space;
[0047] The optimization of the random forest classifier is specifically as follows: Using recursive feature elimination, iteratively removing unimportant features and retraining the model to determine the most influential features; integrating an adaptive feature weighting mechanism to dynamically adjust the feature importance scores according to the dataset features; implementing a stability-aware decision threshold to dynamically adjust the classification threshold based on the confidence distribution of the predictions.
[0048] The optimization of the XGBoost classifier is specifically as follows: Using Bayesian optimization to adjust hyperparameters. During the training process, according to the sum of the gradients and Hessian matrices in the loss function, assign missing values to the split side that makes the model fit better.
[0049] The optimization of the long short-term memory neural network classifier is specifically as follows: Integrating an attention mechanism to focus on the relevant time steps for attack detection; adding peephole connections to allow direct access to the neuron states; using a mixed activation function to more effectively model complex nonlinear relationships, accurately capture the temporal dynamics of the data, and enhance the classification ability for evolving attack patterns.
[0050] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses and provides an efficient and intelligent network attack classification tool based on a federated learning framework, having the following beneficial effects:
[0051] (1) High accuracy and robustness: The present invention fuses the outputs of multiple classifiers through an optimized Dempster-Shafer evidence theory algorithm, significantly improving the accuracy of attack detection. Compared with a single classifier, the fusion method can effectively overcome the conflicts between classifiers and improve the detection ability for complex attack patterns (such as unknown attacks).
[0052] (2) Strong adaptability: The present invention can dynamically adjust the confidence quality, adaptively select a suitable classifier according to different attack scenarios. The adaptive mechanism enables the framework to maintain a high detection accuracy and robustness when facing various known and unknown attacks. Especially when facing new attacks, it can continuously optimize its decision-making process based on historical data to ensure a quick adaptation to unknown attacks.
[0053] (3) Data privacy protection: Using federated learning technology enables the present invention to still perform effective model training and updating without exposing users' sensitive data. This decentralized training method ensures the protection of the data privacy of each node and avoids the security risks brought by centralized data storage.
[0054] (4) Flexible deployment and scalability: The system framework design proposed by the present invention is lightweight, can adapt to different network environments, and supports distributed deployment. It can not only handle network systems of different scales but also flexibly handle different types of data and attack scenarios.
[0055] (5) Computational efficiency and real-time performance: While ensuring high accuracy, the present invention effectively reduces the computational complexity by combining the advantages of multiple classifiers. Especially for real-time detection requirements, it can quickly respond to abnormal activities in the real environment, ensuring timely discovery and response to various attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the provided drawings.
[0057] Figure 1 Schematic diagram of an efficient and intelligent network attack classification tool based on a federated learning framework provided by the present invention;
[0058] Figure 2 Flowchart of an efficient and intelligent network attack classification tool based on a federated learning framework provided by the present invention;
[0059] Figure 3 Internal flowchart of the local model and the central model provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0061] The embodiments of the present invention disclose an efficient and intelligent network attack classification tool based on a federated learning framework, as shown in Figure 1 and Figure 2 , which includes a central server and multiple edge nodes. Each edge node deploys a local model, and the central server deploys a central model. Both the local model and the central model include a multi-task logistic regression classifier, a random forest classifier, an XGBoost classifier, a long short-term memory neural network classifier, and a D-S evidence theory optimization algorithm;
[0062] The local model collects local network traffic data, combines the D-S evidence theory optimization algorithm and four types of complementary classifiers for model training, and performs abnormal data detection and classification, and transmits the key update parameters related to the abnormality of the encrypted abnormal data to the central server;
[0063] The central model receives and aggregates the encrypted update parameters. After updating the global model, it classifies the attack types based on the encrypted abnormal data combined with the D-S evidence theory optimization algorithm, and redistributes the classification results and the updated global model to each edge node.
[0064] To further implement the above technical solution, the edge nodes include the user side, the server side, the database side, the network side, and the application side.
[0065] In practical applications, the local network traffic data includes a network attack data set, which contains 49 features covering various statistical information of network traffic, including: basic traffic features such as packet length and time interval; content features such as TCP flags and protocol types; time features such as the duration of each flow and the number of packets; and a network intrusion detection data set, which contains more than 80 features, including: traffic features such as traffic size, flow rate, and the number of bytes per connection; time features such as the start and end times of the connection and the connection duration; and protocol features such as the protocol types used in the connection (e.g., TCP, UDP, etc.).
[0066] To further implement the above technical solution, as Figure 3 , the specific content of model training by combining the D-S evidence theory optimization algorithm and four types of complementary classifiers is:
[0067] Input the network traffic data set. The D-S evidence theory optimization algorithm assigns initial confidence levels to the four classifiers and they enter the four classifiers in parallel for classification tasks. Among them, the local model task is anomaly detection, and the central model task is attack classification;
[0068] After the classification results are output, evidence combination is performed according to the Dempster combination rule to integrate the classification predictions from the multi-task logistic regression classifier, the random forest classifier, the XGBoost classifier, and the long short-term memory neural network classifier, while considering the consistency and conflict between the classification predictions;
[0069] The aggregated confidence values on the frame of discernment are represented by the combined basic probability assignment, the confidence function for decision-making is derived, and finally the hypothesis with the highest confidence value is selected for the final classification decision based on the threshold criterion.
[0070] Traditional D-S evidence theory is applied to the sensor fusion of homogeneous data sources. However, in a network attack environment, the data sources are usually heterogeneous and contain various types of data. This characteristic makes the D-S evidence theory particularly important in a dynamic and uncertain environment because this data may be noisy, missing, or affected by attacks.
[0071] In this embodiment, the frame of discernment Θ represents all possible attack types, Θ = {θ1, θ2,..., θ n}. The BPA is called the belief quality \(m\), which assigns a probability to each subset of \(\Theta\) representing the degree of trust in that subset. In this embodiment, the BPA needs to satisfy the following conditions:
[0072] \(m: 2^{\Theta}\to[0,1]\) Θ \(\to[0,1]\)
[0073]
[0074] where \(m(A)\) is the belief quality assigned to subset \(A\), and \(\varphi\) is the empty set;
[0075] The Dempster combination rule is commutative and associative:
[0076]
[0077] where \(\oplus\) represents the Dempster's combination operation;
[0078] The basic probability assignments \(m_1\) and \(m_2\) come from two independent evidence sources, specifically classifiers. The Dempster rule is used to fuse them into a single BPA:
[0079]
[0080] where the conflict coefficient \(K\) i is a normalization factor for explaining the conflict between evidence sources;
[0081] \(K\) 12 \(=\sum_{B\cap C = \varphi}m_1(B)\cdot m_2(C)\) B∩C=φ \(m_1(B)\cdot m_2(C)\)
[0082] \(\frac{1}{1 - K}\) i By redistributing the conflicting belief values to ensure that the combined belief function remains a valid BPA, and then gradually combining \(m_3\) with \(m_4\), the combined BPA and the conflict coefficient from four sources are given.
[0083] To further implement the above technical solution and improve the robustness and accuracy of detection, the present invention proposes a D-S evidence theory optimization algorithm and improves the D-S evidence fusion process through the following key enhancements:
[0084] Improve the dynamic update of belief quality and adjust the degree of trust in the outputs of each classifier according to the historical performance of the classifier;
[0085] Introduce context-aware evidence combination and allocate belief quality for different attack patterns;
[0086] Weighted conflict resolution, assign a weight to each classifier, prioritize according to the reliability of the classifier for a specific anomaly, and calculate the belief quality through weighted combination;
[0087] Integrate temporal evidence aggregation, based on time series trends, and dynamically update the confidence quality according to historical evidence and emerging data to adapt to new attack patterns.
[0088] To further implement the above technical solution, the specific method for improving the dynamic update of confidence quality is as follows:
[0089] Update the confidence of each hypothesis A according to the historical performance of the classifier:
[0090]
[0091] where ω classifier (A) represents the reliability of the classifier predicting hypothesis A over time, hypothesis A is a possible type of attack, B is a subset of A, and m is the confidence quality;
[0092] For any initial confidence quality distribution, through a sequence of update rules m k will converge:
[0093]
[0094] where m k (A) is the k-th confidence quality of hypothesis A, m * (A) is the updated confidence quality, and the stability of the confidence quality assignment is guaranteed by the gradual decrease of entropy. With each iteration, the entropy will gradually decrease, thus guiding the system towards the final convergence state:
[0095]
[0096] where H(m) is the entropy, Θ is the frame of discernment, representing all possible types of attacks.
[0097] To further implement the above technical solution, the weight ω assigned to each classifier i is:
[0098]
[0099] where ω i represents the relative importance of the i-th classifier C i among all classifiers. The numerator represents the accuracy of the i-th classifier, and the denominator represents the sum of the accuracies of all classifiers from the 1st classifier to the n-th classifier;
[0100] The final confidence quality is calculated through weighted combination:
[0101]
[0102] To further implement the above technical solution, after integrating the time evidence aggregation, the dynamically updated confidence quality is as follows:
[0103] m t (A) = (1 - α)·m t-Δt (A) + α·m new (A)
[0104] where α is a balancing factor used to control the ratio between historical evidence and new evidence, and t represents the time instant.
[0105] To further implement the above technical solution, in this embodiment, the classifier in the local model, i.e., the local classifier, is designed to enhance the sensitivity to abnormal patterns. Specifically, the feature weights are adjusted to amplify the signals related to anomalies while reducing the influence of the features corresponding to typical operating behaviors; during the training phase, each model updates its model using its local dataset, which contains normal and abnormal records;
[0106]
[0107] where L i is the loss function minimized during the local training process, and θ′ i are the optimized parameters transmitted from the local model to the central model;
[0108] To improve efficiency, each meter adopts an adaptive learning rate η i , to ensure stable convergence during the training process:
[0109]
[0110] Once the training is completed, the updated model parameters are encrypted and sent to the central server. To further protect privacy, only the important updates related to the detected anomalies are selectively transmitted. These updates contain the extracted attack-related information while filtering out the normal traffic data, reducing the risk of sensitive data exposure;
[0111] Once an anomaly is detected, the central server takes over the classification of specific types of attacks. The central server receives the encrypted model parameters from the participating local models, aggregates them, and updates the global model M g , and the aggregated global model parameters θ′ g are calculated as:
[0112]
[0113] where represents the data contribution weight of each node;
[0114] Then, the global model is fine-tuned to accurately classify the detected anomalies. The classification process includes adjusting the feature weights to emphasize the attack-specific features; leveraging the feature importance score φ k and the anomaly score S i to further improve the classification accuracy:
[0115]
[0116] The updated model is used to identify the attack type:
[0117] C k = classify(θ″ g , D abn )
[0118] Each detected anomaly is classified into a specific attack type, and the classification result is securely transmitted back to the local. The updated global model M″ g is reallocated to the local to enhance their continuous learning and adaptation. Each local node downloads the model and updates it in the local system to improve the attack detection accuracy.
[0119] To further implement the above technical solution, before model training, four types of complementary classifiers selected for the local model and the central model, namely the multi-task logistic regression classifier, the random forest classifier, the XGBoost classifier, and the long short-term memory neural network classifier, are respectively optimized and designed.
[0120] To further implement the above technical solution, the optimization of the MTLR multi-task logistic regression classifier is specifically as follows: By sharing the parameter matrix W shared to capture common attack patterns while retaining the task-specific parameters to improve the performance across attack types and reduce model redundancy; adopting adversarial training, incorporating clean samples and adversarial samples into the loss function to enhance the model's robustness; introducing the local adaptive kernel logistic regression LAK-LR, using the kernel method to capture complex attack patterns in the high-dimensional space and improving the accuracy without sacrificing efficiency;
[0121] Specifically:
[0122] In MTLR, all tasks share a parameter matrix W shared , and each task also has a specific task matrix whose objective function is to jointly minimize for all tasks:
[0123]
[0124] where T represents the number of tasks, N iis the number of samples for task i, and λ is the regularization parameter used to control the sparsity of task-specific weights;
[0125] To enhance the robustness of the model, adversarial samples are generated by dynamically adjusting the perturbation direction, which is adjusted based on the model confidence and sample distribution. The update rule for adversarial samples is:
[0126]
[0127] where α is the step size, represents the gradient of the loss function with respect to the adversarial sample, γ is a hyperparameter that adjusts the perturbation direction to be based on the sample mean x mean ;
[0128] The modified objective function introduces adversarial sample training and dynamically adjusts the weight factor β:
[0129]
[0130] where β is dynamically adjusted during training;
[0131]
[0132] where k controls the rate of change and t0 is the transition point;
[0133] To address the limitations of traditional logistic regression in modeling non-linear relationships, LAK-LR based on kernel methods is introduced, which can capture complex attack patterns in high-dimensional space and improve accuracy without sacrificing computational efficiency;
[0134]
[0135] where K j (x, x′) represents different kernel functions, and θ j (x) is the adaptive weight;
[0136]
[0137] where μ j is the mean of the selected sample subset, and σ j controls the influence range of the kernel function;
[0138] The final classification decision is:
[0139]
[0140] where α i is the learning coefficient of the training samples, and K(x, x i ) calculates the similarity between the input and the training sample x i .
[0141] The optimization of the random forest classifier RF is specifically as follows: Using recursive feature elimination (RFE), iteratively remove unimportant features and retrain the model to determine the most influential features, reduce dimensionality, reduce noise, and enhance interpretability; Integrate an adaptive feature weighting mechanism to dynamically adjust the feature importance scores according to the characteristics of the dataset, prevent overfitting, and adapt to different environments; Implement a stability-aware decision threshold (SADT) to dynamically adjust the classification threshold based on the confidence distribution of the predictions and reduce the misclassification rate in boundary attack scenarios.
[0142] Specifically:
[0143] Initialize RFE by including all features in the feature set F = {(f1, f2,..., f n )}, where n is the total number of features;
[0144] Introduce an adaptive weighting function to calculate the importance score of features by combining historical feature correlations and the characteristics of the current dataset:
[0145]
[0146] where ω k is the dynamic weight during iteration;
[0147] Adjust based on the model confidence. Perf k (f i ) represents the performance score of feature f i in the t-th iteration, and N is the total number of iterations; In each iteration, the weight ω k is updated in the following way to prefer features with high stability and large influence:
[0148]
[0149] where σ(f i ) represents the standard deviation of feature f i in different iterations, and ∈ is a small positive number used to avoid division by zero;
[0150] Remove features with importance scores lower than a certain threshold θ, and update the feature set, where θ is an adaptive threshold dynamically adjusted based on the dataset complexity; When the classification performance is stable, or the number of selected features reaches the set minimum value |F′| min , stop the iteration.
[0151] To further improve robustness, introduce a stability-aware decision threshold (SADT) to dynamically adjust the classification boundary to reduce the misclassification problem of adversarial samples. The traditional random forest classifier uses a fixed classification threshold, while the present invention adaptively adjusts the decision boundary according to the confidence distribution:
[0152]
[0153] Among them, is the average confidence score of the classifier, is the standard deviation of the confidence, and λ is an adjustable parameter used to control the adjustment amplitude;
[0154] The final classification decision is based on the following formula:
[0155]
[0156] Among them, is the prediction probability of the random forest; through this method, misclassification can be reduced in boundary cases (such as weak attack signals in anomaly detection), and the overall detection accuracy can be improved.
[0157] The optimization of the XGBoost classifier is specifically as follows: Using Bayesian optimization to adjust hyperparameters such as tree depth, learning rate, and regularization term. During the training process, according to the sum of the gradient and Hessian matrix in the loss function, the missing values are assigned to the splitting side that makes the model fit better, enhancing the robustness and adaptability of the model to incomplete or noisy data, and being suitable for real-time applications;
[0158] Specifically:
[0159] Bayesian optimization automatically adjusts the hyperparameters of XGB to achieve optimal performance, and the optimization objective is:
[0160]
[0161] Among them, f(θ) represents the value of the loss function under the hyperparameter θ;
[0162] In each iteration, the update of the hyperparameters follows the following criterion:
[0163]
[0164] Among them, μ t (θ) and σ t (θ) represent the mean and standard deviation of the current hyperparameters respectively, estimated by Gaussian process regression, used to balance exploration and exploitation;
[0165] The optimization objective of XGB includes error loss and regularization term to prevent the model from overfitting;
[0166] The objective function is:
[0167]
[0168] Among them, is the prediction error, Ω(f k) is the regularization term of the k-th decision tree, where γ and λ are hyperparameters that control the complexity of the tree and the L2 regularization degree of the leaf nodes respectively;
[0169] In actual data collection, some data may have missing values. Therefore, an optimization method for missing values based on gradient information is adopted. First, when splitting the decision tree, XGB calculates the gradient and Hessian of each feature (used to measure the second-order information of the loss function):
[0170]
[0171] where G and H are the cumulative values of the gradient and Hessian respectively, and I left and I right represent the sets of sample indices divided into the left subtree and the right subtree respectively;
[0172] Then, XGB calculates the gain to evaluate the best split point:
[0173]
[0174] By automatically selecting whether the missing value should be assigned to the left subtree or the right subtree, the optimality of data division is ensured;
[0175] The optimization of the long short-term memory neural network classifier is specifically as follows: integrating the attention mechanism to focus on the relevant time steps for attack detection; adding peephole connections to allow direct access to the neuron state, thereby improving the management of the state during the entire sequence processing; using a mixed activation function to more effectively model complex nonlinear relationships, accurately capture the time dynamics of the data, and enhance the classification ability for evolving attack patterns;
[0176] Specifically:
[0177] Introduce the attention mechanism to enable the model to focus on more relevant time steps to improve the attack detection ability;
[0178] The calculation formula for the attention weight is:
[0179]
[0180] where α t is the attention weight at the t-th time step, e t is the attention score assigned to the time step t, and the previous hidden state h t-1 and the current input x t are considered during the calculation. v, W h , W x are trainable weight parameters, and b a is the bias term;
[0181] Context vector and updated hidden state calculation:
[0182]
[0183] Among them, c t is the context vector, which integrates important information at different time steps, and h t is the updated hidden state, and o t is the weight of the output gate to control the information flow;
[0184] Adding Peephole connections allows the gated unit to directly access the neuron state (Cell State), thereby improving the time series modeling ability;
[0185] Gated calculation with Peephole connections:
[0186] f t = σ(W f [h t-1 , x t , c t-1 +b f )
[0187] i t = σ(W i [h t-1 , x t , c t-1 +b i )
[0188] o t = σ(W o [h t-1 , x t , c t +b o )
[0189] Among them, f t is the forget gate, which controls the degree of forgetting of past information; i t is the input gate, which determines the influence of the current input on the neuron state; o t is the output gate, which determines which information is output to the hidden state; c t-1 and c t are the neuron states at the previous time step and the current time step respectively;
[0190] After introducing Peephole connections, LSTM can capture the long-term dependencies of time series data more accurately, improving the stability and robustness of the attack detection model;
[0191] To enhance the modeling ability of LSTM for non-linear relationships, a mixed activation function is adopted in the input modulation gate, replacing the traditional tanh with ReLU to enhance the model's representation ability:
[0192]
[0193] This hybrid activation strategy can improve the ability of LSTM to capture complex attack patterns, making its performance more stable in the attack classification task.
[0194] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.
[0195] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An efficient intelligent network attack classification tool based on the federated learning framework, characterized in that, It includes a central server and multiple edge nodes. Each edge node deploys a local model, and the central server deploys a central model. Both the local model and the central model include a multi-task logistic regression classifier, a random forest classifier, an XGBoost classifier, a long short-term memory neural network classifier, and a D-S evidence theory optimization algorithm; The local model collects local network traffic data, combines the D-S evidence theory optimization algorithm and four types of complementary classifiers for model training, and conducts anomaly data detection and classification, and transmits the key update parameters related to the anomaly of the encrypted anomaly data to the central server; The central model receives and aggregates the encrypted update parameters. After updating the global model, it conducts attack type classification based on the encrypted anomaly data combined with the D-S evidence theory optimization algorithm, and redistributes the classification results and the updated global model to each edge node.
2. The efficient and intelligent network attack classification tool based on the federated learning framework according to claim 1, characterized in that, The edge nodes include the user side, the server side, the database side, the network side, and the application side.
3. An efficient and intelligent network attack classification tool based on the federated learning framework according to claim 1, characterized in that, The specific content of combining the D-S evidence theory optimization algorithm and four types of complementary classifiers for model training is as follows: Input the network traffic data set. The D-S evidence theory optimization algorithm assigns initial confidence levels to the four classifiers and they enter the four classifiers in parallel for classification tasks. Among them, the local model task is anomaly detection, and the central model task is attack classification; After the classification results are output, evidence combination is performed according to the Dempster combination rule to integrate the classification predictions from the multi-task logistic regression classifier, random forest classifier, XGBoost classifier, and long short-term memory neural network classifier, and at the same time consider the consistency and conflict between the classification predictions; Represent the aggregated confidence values on the frame of discernment through the combined basic probability assignment, deduce the confidence function of the decision, and finally select the hypothesis with the highest confidence value and make the final classification decision based on the threshold criterion.
4. An efficient intelligent network attack classification tool based on the federated learning framework according to claim 1, characterized in that, The specific content of the D-S evidence theory optimization algorithm is as follows: Improve the dynamic update of confidence quality and adjust the degree of trust in the outputs of each classifier according to the historical performance of the classifier; Introduce context-aware evidence combination and allocate confidence quality for different attack patterns; Weighted conflict resolution, assign a weight to each classifier, prioritize according to the reliability of the classifier for a specific anomaly, and calculate the confidence quality through weighted combination; Integrate time-evidence aggregation, and dynamically update the confidence quality based on the time series trend, according to historical evidence and newly emerging data to adapt to new attack patterns.
5. An efficient and intelligent network attack classification tool based on the federated learning framework according to claim 4, characterized in that The specific method for improving the dynamic update of confidence quality is as follows: Update the confidence of each hypothesis A according to the historical performance of the classifier: where ω classifier (A) represents the reliability of the classifier's prediction of hypothesis A over time, where hypothesis A is a possible attack type, B is a subset of A, and m is the confidence quality; For any initial confidence mass distribution, through a sequence of update rules, m k will converge: where m k (A) is the k-th confidence mass for hypothesis A, m * (A) is the updated confidence mass, and the stability of the confidence mass assignment is ensured by a gradual decrease in entropy, which decreases with each iteration, guiding the system towards a final convergent state: Among them, H(m) is entropy, Θ is the frame of discernment, representing all possible attack types.
6. The efficient and intelligent network attack classification tool based on the federated learning framework according to claim 4, characterized in that, The weight ω assigned to each classifier i is as follows: Among them, ω i represents the relative importance of the i-th classifier C i among all classifiers. The numerator represents the accuracy of the i-th classifier, and the denominator represents the sum of the accuracies of all classifiers from the 1st classifier to the nth classifier; The final confidence quality is calculated through weighted combination:
7. An efficient intelligent network attack classification tool based on the federated learning framework according to claim 1, characterized in that, After integrating time-evidence aggregation, the dynamically updated confidence quality is: m t (A) = (1 - α)·m t-Δt (A) + α·m new (A) Among them, α is a balance factor used to control the ratio between historical evidence and new evidence, and t represents the time.
8. An efficient and intelligent network attack classification tool based on the federated learning framework according to claim 1, characterized in that, The global model parameters aggregated by the central server are: Among them, represents the data contribution weight of each node; where L i is the loss function minimized in the local training process, and θ′ i is the optimized parameter transferred from the local model to the central model.
9. An efficient intelligent network attack classification tool based on the federated learning framework according to claim 1, characterized in that, Before model training, four types of complementary classifiers selected for the local model and the central model, namely, the multi-task logistic regression classifier, the random forest classifier, the XGBoost classifier, and the long short-term memory neural network classifier, are optimized and designed respectively.
10. An efficient and intelligent network attack classification tool based on the federated learning framework according to claim 1, characterized in that, The optimization of the multi-task logistic regression classifier is specifically as follows: capture common attack patterns by sharing parameter matrices while retaining task-specific parameters; adopt adversarial training and incorporate clean samples and adversarial samples into the loss function; introduce local adaptive kernel logistic regression and use the kernel method to capture complex attack patterns in high-dimensional space; The optimization of the random forest classifier is specifically as follows: use recursive feature elimination to iteratively remove unimportant features and retrain the model to determine the most influential features; integrate an adaptive feature weighting mechanism to dynamically adjust feature importance scores according to the characteristics of the dataset; implement a stability-aware decision threshold to dynamically adjust the classification threshold based on the confidence distribution of predictions; The optimization of the XGBoost classifier is specifically as follows: use Bayesian optimization to adjust hyperparameters. During the training process, according to the sum of the gradient and the Hessian matrix in the loss function, assign missing values to the split side that makes the model fit better; The optimization of the long short-term memory neural network classifier is specifically as follows: integrate an attention mechanism to focus on relevant time steps for attack detection; Add peephole connections to allow direct access to neuron states; use a mixed activation function to more effectively model complex nonlinear relationships, accurately capture the temporal dynamics of data, and enhance the classification ability for evolving attack patterns.
Citation Information
Cited By
Privacy protection type distributed optimization method and system based on optimal control
CN121503734A
A privacy-preserving distributed optimization method and system based on optimal control
CN121503734B
A federated learning method and system fusing attack tracing and data integrity auditing
CN122640242A